LCSMC-Net: Lightweight CAN Intrusion Detection via Separable Multiscale Convolution and Attention
Abstract
1. Introduction
- A Separable Multiscale Convolution Lite (SMC-Lite) module that employs dual-scale depthwise convolution () with elementwise averaging fusion, specifically calibrated to capture CAN attack temporal signatures—short-term bursts and medium-term periodic anomalies—within a minimal parameter budget.
- A Lightweight Channel-Temporal Attention (LCTA) mechanism that decomposes attention into independent channel recalibration and conditional temporal weighting with adaptive pruning, achieving complexity suitable for real-time embedded inference.
- A compact 6-dimensional CAN-optimized feature representation that encodes protocol semantics (ID priority, timestamp interval, data entropy, anomaly score), enabling cross-protocol generalization between CAN 2.0B and CAN-FD with only accuracy degradation.
- An end-to-end compression pipeline integrating Bayesian hyperparameter optimization (TPE) and knowledge distillation, yielding a final model of 9401 parameters (<10 KB) with detection accuracy and inference latency on ARM Cortex-M4.
2. Background: CAN Protocol and Attack Taxonomy
2.1. Controller Area Network Protocol Fundamentals
2.2. CAN Intrusion Attack Taxonomy
- Denial-of-Service (DoS): This attack involves saturating the CAN bus with high-priority messages (e.g., ID 0x000) to consume bandwidth and preclude legitimate communications [3]. Signatures include bus utilization approaching and the disruption of periodic message transmission.
- Fuzzing: This reconnaissance technique entails injecting messages with randomized identifiers and payloads to probe ECU vulnerability [11]. Key indicators include previously unseen identifiers and elevated entropy within message content.
- Spoofing: Capitalizing on the lack of authentication, adversaries masquerade as legitimate ECUs to inject fabricated sensor readings (e.g., RPM or Gear status) [5]. Application-layer variants, such as RPM Spoofing (ID 0x316) and Gear Spoofing (ID 0x43F), can precipitate erroneous ECU decisions.
- Replay: Adversaries capture legitimate message sequences and retransmit them to execute unauthorized actions, exploiting the absence of freshness mechanisms (e.g., timestamps or nonces) [12].
2.3. Embedded Intrusion Detection Challenges
3. Related Work
3.1. CAN Bus Security and Intrusion Detection Systems
3.2. Lightweight Neural Network Architectures
3.3. Hyperparameter Optimization for Resource-Constrained Models
3.4. Knowledge Distillation for Model Compression
3.5. Attention Mechanisms for Efficient Feature Selection
4. Proposed LCSMC-Net Framework
4.1. Overview and Design Philosophy
- 6-Dimensional CAN-Optimized Feature Engineering, which transforms raw CAN messages into a compact protocol-aware representation encoding ID priority, temporal patterns, payload statistics, and anomaly indicators;
- Input Embedding Layer, which projects the 6-dimensional features into a higher-dimensional latent space via convolution to enable richer feature interaction;
- Three Hierarchical Feature Extraction Stages, each consisting of an SMC-Lite block (capturing multi-temporal attack patterns via dual-scale depthwise convolution) followed by an LCTA mechanism (selectively emphasizing discriminative features and anomalous time steps), with progressive channel expansion and spatial downsampling;
- Lightweight Classification Head, which applies Global Average Pooling followed by a fully connected layer to produce the final attack category prediction with minimal parameters.
- Domain-Specific Optimization: CAN intrusion detection operates on short, discrete, event-triggered message sequences (typically time steps) with protocol-specific semantics. The brevity and discrete nature of CAN windows impose strict constraints on receptive field design and demand compact, targeted feature extraction rather than broad-spectrum signal decomposition.
- Aggressive Parameter Efficiency: Automotive ECUs allocate fewer than for IDS models. This extreme constraint necessitates that every architectural component—convolution, fusion, and attention—be optimized for minimal parameter footprint while maintaining detection accuracy. Design choices such as elementwise averaging fusion and decomposed attention with adaptive pruning are directly motivated by this budget.
- Multitemporal Feature Extraction: CAN attacks manifest at different temporal scales: DoS flooding produces short-term bursts within 1–3 consecutive frames, while spoofing attacks introduce medium-term periodic deviations across 3–5 message cycles. The architecture must capture both granularities simultaneously while recognizing that these scales reflect overlapping views of the same anomaly rather than independent signal components.
4.2. Separable Multiscale Convolution Lite (SMC-Lite) Blocks
4.2.1. Motivation and Design Rationale
4.2.2. Mathematical Formulation of SMC-Lite
4.2.3. Complexity Analysis
4.3. Lightweight Channel-Temporal Attention (LCTA) Mechanism
4.3.1. Motivation
- Decomposed Channel-Temporal Attention: Rather than a unified global attention mechanism, LCTA decomposes the computation into two sequential, independent stages: Channel Attention identifies which features are discriminative for the current input (e.g., prioritizing Data_Entropy for Fuzzing detection vs. ID_Priority for DoS detection), while Temporal Attention identifies when anomalies occur within the message window. This decomposition reflects the structure of CAN intrusion detection, where feature relevance and temporal localization are independent decision axes that benefit from separate modeling.
- SE-Style Channel Recalibration: The 6D CAN-optimized features have explicit, heterogeneous protocol semantics—each feature channel carries distinct physical meaning (e.g., ID_Low encodes ECU identity, Data_Entropy encodes payload randomness). A Squeeze-and-Excitation mechanism [18] with reduction ratio is well-suited to learn input-dependent channel weights that selectively emphasize the most relevant features for each traffic class, achieving effective recalibration with minimal parameters.
- Adaptive Temporal Pruning: CAN sequences undergo progressive downsampling through the hierarchical architecture (). When , only two time steps remain, offering negligible temporal structure for meaningful attention computation. LCTA automatically disables its temporal branch in this regime, eliminating unnecessary parameters and computation. This adaptive behavior is a direct consequence of designing for short CAN sequences, where aggressive spatial compression renders temporal attention redundant in deeper layers.
4.3.2. Mathematical Formulation of LCTA
4.4. Hierarchical Architecture with Adaptive Attention
4.4.1. Three-Stage Feature Extraction Pipeline
4.4.2. Classification Head
4.5. 6-Dimensional CAN-Optimized Feature Engineering
Rationale and Formulation
- F1: ID_Low (): This feature extracts the lower 8 bits of the identifier. Since CAN IDs are often allocated in blocks to specific ECUs, the lower bits represent the identity signature of the transmitting node. Unseen values in indicate potential unauthorized message injection or masquerading attempts.
- F2: ID_Priority (): This feature extracts the upper 3 bits, which control the CSMA/CR arbitration process. In Denial-of-Service (DoS) attacks, adversaries typically inject messages with ID 0x000 (highest priority) to monopolize the bus. captures this arbitration abuse, enabling detection of bus contention anomalies.
- F3: Data_Mean (): Calculated as the arithmetic mean of the payload bytes . Legitimate sensor data typically shows physical continuity. In contrast, spoofing attacks may freeze the payload to a constant value (resulting in zero variance in over time) or inject abrupt value jumps, which this feature can detect statistically.
- F4: Data_Entropy (): This feature uses Shannon entropy to quantify randomness within a message. Normal vehicle signals are highly structured and have low entropy. In contrast, Fuzzing attacks, which inject randomized payloads to crash ECUs, produce maximum entropy. effectively discriminates such high-entropy injection attacks.
- F5: Time_Pattern (): This represents the inter-arrival time () between consecutive messages. Under normal conditions, varies around a fixed period (e.g., 10 ms). DoS attacks cause due to bus flooding, while replay attacks often break the timing regularity. This feature converts temporal anomalies into input values for the neural network.
- F6: Anomaly_Score (): To capture historical context, we define a composite Z-score metric based on a sliding window W. Let and be the rolling mean and standard deviation of the sequence of values. is computed as . This feature functions as a soft outlier detector, highlighting sporadic anomalies that may appear subtle individually but are statistically significant compared to recent history.
4.6. Hyperparameter Optimization via Bayesian Search
4.7. Knowledge Distillation for Ultra-Constrained Deployment
4.8. Hardware Requirements and Deployment Conditions
4.8.1. Target Platform Specifications
4.8.2. Resource Consumption Analysis
4.8.3. Deployment Feasibility
5. Experimental Evaluation
5.1. Dataset Description and Preprocessing
- Normal: Legitimate ECU communications during standard driving scenarios (urban, highway, parking).
- DoS (Denial-of-Service): High-priority message flooding at ID 0x000, injected at 3000 frames per second (fps).
- Fuzzing: Random ID (0x000–0x7FF) and payload injection to probe ECU vulnerabilities.
- Impersonation (Spoofing): Fabricated RPM (ID 0x316) and Gear (ID 0x43F) signals mimicking legitimate sensor data.
- Normal: Baseline CAN-FD traffic from powertrain ECUs.
- Flooding: Bus saturation attacks at (lower intensity than Dataset 1 due to CAN-FD arbitration mechanisms).
- Fuzzing: Extended payload randomization exploiting the 64-byte frame capacity.
- Malfunction: Simulated sensor failures (e.g., stuck-at-zero faults) injected into wheel speed and throttle position signals.
Evaluation Metric Rationale
5.2. Comprehensive Performance Evaluation on Dataset 1
5.2.1. Overall Classification Performance and Robustness
5.2.2. Multi-Dimensional SOTA Comparison
5.2.3. Feature Representation Analysis (t-SNE Visualization)
5.2.4. Computational Efficiency and Resource Trade-Off
5.2.5. Ablation Study: Validating Architectural Innovations
5.2.6. Cross-Validation Analysis
5.2.7. Training Dynamics and Stability
5.3. Cross-Protocol Generalization on Dataset 2 (CAN-FD)
5.3.1. Generalization Performance and Protocol Agnosticism
5.3.2. Temporal Dynamics via Advanced Recurrence Plot Analysis
- Normal Traffic (Figure 11a): The recurrence plot generated by LCSMC-Net exhibits a highly ordered, lattice-like topology characterized by continuous diagonal lines. This regular pattern corresponds to the deterministic timing behavior of legitimate CAN frames, where ECUs transmit messages at fixed intervals (e.g., 10 ms or 20 ms cycles), reflecting the temporal stability of normal in-vehicle network operation.
- Flooding Traffic (Figure 11b): The recurrence plot reveals a transition to dense block structures. These “recurrence blocks” indicate a state of high temporal self-similarity, resulting from the attacker’s continuous injection of identical high-priority messages. This observation suggests that the Time_Pattern feature effectively captures the bus saturation state induced by flooding attacks.
- Fuzzing Traffic (Figure 11c): The recurrence plot exhibits a stochastic and fragmented texture, characterized by short, broken diagonals and isolated points. This pattern is consistent with the entropy-maximizing nature of Fuzzing attacks, where randomized IDs and payloads disrupt the temporal correlations typical of in-vehicle networks.
- Malfunction Traffic (Figure 11d): The recurrence plot retains a quasi-periodic structure similar to Normal traffic but exhibits subtle disruptions in diagonal laminarity (i.e., gaps in the diagonal lines). LCSMC-Net’s ability to distinguish this class from Normal traffic demonstrates the sensitivity of its SMC-Lite blocks to microscale temporal anomalies that do not fundamentally alter the global periodicity.
5.3.3. Granular Signal-Level Anomaly Detection
- Arbitration Monopoly (Figure 12a): During normal operation (), the ID pattern extracted by LCSMC-Net (green points) fluctuates widely, reflecting the diverse communication among multiple ECUs. Upon attack onset (), the pattern converges to a static minimum value (red points, normalized ). This corresponds to the injection of ID 0x000, demonstrating the model’s ability to detect arbitration abuse via the ID_Priority feature.
- Pattern Rigidity (Figure 12b): The Data Length Code (DLC) transitions from stochastic fluctuations to a rigid, repetitive sequence during the attack phase. This loss of entropy is a hallmark of automated injection tools, which LCSMC-Net effectively captures via the Anomaly_Score feature.
- Payload Distribution Shift (Figure 12c): The payload mean (dots) and its variance (blue shaded band) exhibit a sudden stabilization upon attack onset. The contraction of the confidence interval during the attack phase indicates a significant reduction in data diversity, demonstrating the effectiveness of statistical features such as Data_Mean in detecting anomalies even on variable-length CAN-FD frames.
- Instantaneous Response (Figure 12d): LCSMC-Netś aggregated anomaly score (blue line) exhibits a rapid surge, crossing the detection threshold (red dashed line, ) precisely at , coinciding with the attack onset. This near-instantaneous detection, occurring within three samples of the sliding window stride, demonstrates that LCSMC-Net can identify Flooding attacks at their inception, enabling countermeasures to be triggered before safety-critical actuators are compromised.
5.3.4. Temporal Localization via Grad-CAM
- Flooding: The final layer exhibits distinct vertical activation bands. This contiguous high-attention region corresponds to the duration of the high-frequency message burst.
- Malfunction: In contrast to Flooding, the Malfunction class exhibits sharp, discrete activations at specific time steps. This is consistent with the nature of sensor faults, which often manifest as sudden value jumps.
- Normal: For Normal traffic, the activation landscape is comparatively uniform and low-intensity.
5.4. Interpretability and Behavioral Analysis
5.4.1. Feature Importance and Decision Logic
5.4.2. Local Decision Boundary Analysis via LIME
5.4.3. Internal Attention Dynamics via Channel Heatmaps
- Flooding: In Layers 2 and 3, specific channels exhibit intense activation (white/bright yellow bands), while the majority are suppressed. This suggests LCSMC-Net has effectively “locked onto” the deterministic signature of high-priority message injection.
- Fuzzing: The attention map remains relatively dispersed even in deeper layers, reflecting the stochastic nature of Fuzzing attacks.
- Malfunction: The heatmap exhibits a hybrid pattern more structured than Fuzzing yet distinct from Normal.
5.5. Knowledge Distillation and Efficiency Optimization
5.5.1. Distillation Framework
5.5.2. Hyperparameter Tuning Results
- Distillation Temperature (): The lower bound is set by the need to soften teacher outputs sufficiently to reveal inter-class relationships (as discussed in Figure 17a, produces overly sharp distributions), while the upper bound prevents gradient vanishing caused by excessive entropy maximization ().
- Learning Rate (): Spans the typical operational regime for Adam optimizer on embedded models, with the understanding that higher temperatures require smaller learning rates to compensate for gradient scaling (see nonlinear coupling in Figure 17a).
- Soft Label Weight (): Ensures balance between teacher supervision and ground-truth guidance. Values cause the student to over-rely on hard labels (losing distillation benefits), while propagates teacher errors and destabilizes training.
- Batch Size (): Constrained by Car-Hacking dataset size ( samples) and GPU memory limitations.
- Weight Decay (): Standard range for L2 regularization; excessive values () suppress feature learning, while insufficient regularization () fails to prevent overfitting.
5.5.3. Comprehensive Performance Benchmarking
- Parameter reduction: compression ( parameters).
- Memory footprint: reduction ().
- Inference speedup: acceleration ().
- Accuracy cost: degradation ().
5.5.4. Granular Classwise Detection Analysis
5.5.5. Ablation Study on Distillation Mechanism
- Optimal Entropy Regime: The optimal performance at is associated with the lowest Soft Loss (). This suggests that the student has effectively assimilated the teacher’s probability distribution. Deviating from this optimal point, either by sharpening targets () or excessively smoothing them (), increases the KL divergence, indicating a loss of information.
- Hybrid Supervision Necessity: For , the optimal results are achieved with a teacher-dominant mix (). However, increasing to reduces performance (), suggesting that ground truth supervision (Hard Loss) remains essential for correcting teacher biases.
- Dark Knowledge Transfer: The significant reduction in Soft Loss (from an initial value of to ) quantitatively suggests that the student is not merely memorizing labels but is learning the relational structure among classes.
6. Discussion
6.1. Key Findings and Implications
6.2. Generalization and Overfitting Analysis
- Minimal train-test gap (): Without regularization, deep models on the Car-Hacking dataset typically exhibit 2–5% overfitting gaps (as observed in preliminary unregularized trials during hyperparameter tuning).
- Low cross-validation variance (): The tight standard deviation across 5 folds indicates that the model learns generalizable patterns rather than memorizing fold-specific noise. For comparison, removing Dropout increases CV variance to based on ablation trials.
- Smooth training dynamics: LCSMC-Net’s validation loss curve exhibits monotonic convergence (Figure 10), in stark contrast to the volatile oscillatory behavior of Baseline CNN and ShuffleNetV2 (which use less aggressive regularization). These oscillations are symptomatic of overfitting to batch-specific noise.
6.3. Scalability Analysis
6.4. Advanced Attack Robustness
6.5. Performance Estimation Methodology
6.6. Baseline Selection Rationale
7. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| CAN | Controller Area Network |
| CAN-FD | CAN with Flexible Data-Rate |
| ECU | Electronic Control Unit |
| IDS | Intrusion Detection System |
| DoS | Denial-of-Service |
| MAC | Message Authentication Code |
| ADAS | Advanced Driver Assistance Systems |
| LCSMC-Net | Lightweight CAN Intrusion Detection with Separable Multiscale Convolution |
| SMC-Lite | Separable Multiscale Convolution Lite |
| LCTA | Lightweight Channel-Temporal Attention |
| TPE | Tree-structured Parzen Estimator |
| KD | Knowledge Distillation |
| FLOPs | Floating Point Operations |
| CNN | Convolutional Neural Network |
| RNN | Recurrent Neural Network |
| LSTM | Long Short-Term Memory |
| DLC | Data Length Code |
| CRC | Cyclic Redundancy Check |
| MHSA | Multihead Self-Attention |
| SE | Squeeze-and-Excitation |
| ECA | Efficient Channel Attention |
| CBAM | Convolutional Block Attention Module |
| GAP | Global Average Pooling |
| SHAP | SHapley Additive exPlanations |
| LIME | Local Interpretable Model-agnostic Explanations |
References
- Charette, R. This car runs on code. IEEE Spectr. 2009, 46, 3. [Google Scholar]
- Robert Bosch GmbH. CAN Specification Version 2.0; Robert Bosch GmbH: Gerlingen-Schillerhohe, Germany, 1991. [Google Scholar]
- Koscher, K.; Czeskis, A.; Roesner, F.; Patel, S.; Kohno, T.; Checkoway, S.; McCoy, D.; Kantor, B.; Anderson, D.; Shacham, H.; et al. Experimental Security Analysis of a Modern Automobile. In Proceedings of the IEEE Symposium on Security and Privacy, Oakland, CA, USA, 16–19 May 2010; pp. 447–462. [Google Scholar]
- Miller, C.; Valasek, C. Remote Exploitation of an Unaltered Passenger Vehicle. In Proceedings of the Black Hat USA, Las Vegas, NV, USA, 1–6 August 2015; pp. 1–91. [Google Scholar]
- Choi, W.; Jo, H.J.; Woo, S.; Chun, J.Y.; Park, J.; Lee, D.H. Identifying ECUs Using Inimitable Characteristics of Signals in Controller Area Networks. IEEE Trans. Veh. Technol. 2018, 67, 4757–4770. [Google Scholar] [CrossRef] [Scilit]
- Checkoway, S.; McCoy, D.; Kantor, B.; Anderson, D.; Shacham, H.; Savage, S. Comprehensive Experimental Analyses of Automotive Attack Surfaces. In Proceedings of the USENIX Security Symposium, San Francisco, CA, USA, 8–12 August 2011; pp. 77–92. [Google Scholar]
- Wolf, M.; Weimerskirch, A.; Paar, C. Security in Automotive Bus Systems. In Proceedings of the Workshop on Embedded Security in Cars (ESCAR), Bochum, Germany, 17–19 February 2004; pp. 1–13. [Google Scholar]
- Groza, B.; Murvay, S.; van Herrewege, A.; Verbauwhede, I. LiBrA-CAN: Lightweight Broadcast Authentication for Controller Area Networks. ACM Trans. Embed. Comput. Syst. 2017, 16, 90. [Google Scholar] [CrossRef] [Scilit]
- Mundhenk, P.; Steinhorst, S.; Lukasiewycz, M.; Fahmy, S.A.; Chakraborty, S. Lightweight Authentication for Secure Automotive Networks. In Proceedings of the Design, Automation and Test in Europe Conference and Exhibition (DATE), Grenoble, France, 9–13 March 2015; pp. 285–288. [Google Scholar]
- Wu, W.; Li, R.; Xie, G.; An, J.; Bai, Y.; Zhou, J.; Li, K. A Survey of Intrusion Detection for In-Vehicle Networks. IEEE Trans. Intell. Transp. Syst. 2020, 21, 919–933. [Google Scholar] [CrossRef] [Scilit]
- Kang, M.J.; Kang, J.W. Intrusion Detection System Using Deep Neural Network for In-Vehicle Network Security. PLoS ONE 2016, 11, e0155781. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Song, H.M.; Woo, J.; Kim, H.K. In-Vehicle Network Intrusion Detection Using Deep Convolutional Neural Network. Veh. Commun. 2020, 21, 100198. [Google Scholar] [CrossRef] [Scilit]
- Zhang, J.; Zhang, X.; Liu, Z.; Fu, F.; Jiao, Y.; Xu, F. A network intrusion detection model based on BiLSTM with multi-head attention mechanism. Electronics 2023, 12, 4170. [Google Scholar] [CrossRef] [Scilit]
- Cheng, P.; Han, M.; Liu, G. DESC-IDS: Towards an efficient real-time automotive intrusion detection system based on deep evolving stream clustering. Future Gener. Comput. Syst. 2023, 140, 266–281. [Google Scholar] [CrossRef] [Scilit]
- Han, S.; Mao, H.; Dally, W.J. Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding. In Proceedings of the International Conference on Learning Representations (ICLR), San Juan, Puerto Rico, 2–4 May 2016. [Google Scholar]
- Howard, A.G.; Zhu, M.; Chen, B.; Kalenichenko, D.; Wang, W.; Wey, T.; Andreetto, M.; Hartwig, A. MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications. arXiv 2017, arXiv:1704.04861. [Google Scholar] [CrossRef] [Scilit]
- Tan, M.; Le, Q.V. EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks. In Proceedings of the International Conference on Machine Learning (ICML), Long Beach, CA, USA, 9–15 June 2019; pp. 6105–6114. [Google Scholar]
- Hu, J.; Shen, L.; Sun, G. Squeeze-and-Excitation Networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA, 18–22 June 2018; pp. 7132–7141. [Google Scholar]
- Wang, Q.; Wu, B.; Zhu, P.; Li, P.; Zuo, W.; Hu, Q. ECA-Net: Efficient Channel Attention for Deep Convolutional Neural Networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 13–19 June 2020; pp. 11534–11542. [Google Scholar]
- Young, C.; Olufowobi, H.; Bloom, G.; Zambreno, J. Automotive intrusion detection based on constant can message frequencies across vehicle driving modes. In Proceedings of the ACM Workshop on Automotive Cybersecurity, Dallas, TX, USA, 27 March 2019; pp. 9–14. [Google Scholar]
- Hanselmann, M.; Strauss, T.; Dormann, K.; Ulmer, H. CANet: An Unsupervised Intrusion Detection System for High Dimensional CAN Bus Data. IEEE Access 2020, 8, 58194–58205. [Google Scholar] [CrossRef] [Scilit]
- Müter, M.; Asaj, N. Entropy-based anomaly detection for in-vehicle networks. In Proceedings of the 2011 IEEE Intelligent Vehicles Symposium (IV), Baden, Germany, 5–9 June 2011; pp. 1110–1115. [Google Scholar]
- Marchetti, M.; Stabili, D.; Guido, A.; Colajanni, M. Evaluation of anomaly detection for in-vehicle networks through information-theoretic algorithms. In Proceedings of the 2016 IEEE 2nd International Forum on Research and Technologies for Society and Industry Leveraging a Better Tomorrow (RTSI), Bologna, Italy, 7–9 September 2016; pp. 1–6. [Google Scholar]
- Song, H.M.; Kim, H.R.; Kim, H.K. Intrusion detection system based on the analysis of time intervals of CAN messages for in-vehicle network. In Proceedings of the 2016 International Conference on Information Networking (ICOIN), Kota Kinabalu, Malaysia, 13–15 January 2016; pp. 63–68. [Google Scholar]
- Li, H.; Kadav, A.; Durdanovic, I.; Samet, H.; Graf, H.P. Pruning Filters for Efficient ConvNets. In Proceedings of the International Conference on Learning Representations (ICLR), Toulon, France, 24–26 April 2017. [Google Scholar]
- Hinton, G.; Vinyals, O.; Dean, J. Distilling the knowledge in a neural network. arXiv 2015, arXiv:1503.02531. [Google Scholar] [CrossRef] [Scilit]
- Liu, H.; Simonyan, K.; Yang, Y. Darts: Differentiable architecture search. arXiv 2018, arXiv:1806.09055. [Google Scholar]
- Sandler, M.; Howard, A.; Zhu, M.; Zhmoginov, A.; Chen, L.C. Mobilenetv2: Inverted residuals and linear bottlenecks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 18–22 June 2018; pp. 4510–4520. [Google Scholar]
- Zhang, X.; Zhou, X.; Lin, M.; Sun, J. Shufflenet: An extremely efficient convolutional neural network for mobile devices. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 18–22 June 2018; pp. 6848–6856. [Google Scholar]
- Han, K.; Wang, Y.; Tian, Q.; Guo, J.; Xu, C.; Xu, C. Ghostnet: More features from cheap operations. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 13–19 June 2020; pp. 1580–1589. [Google Scholar]
- Akiba, T.; Sano, S.; Yanase, T.; Ohta, T.; Koyama, M. Optuna: A next-generation hyperparameter optimization framework. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, Anchorage, AK, USA, 4–8 August 2019; pp. 2623–2631. [Google Scholar]
- Park, W.; Kim, D.; Lu, Y.; Cho, M. Relational knowledge distillation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA, 16–20 June 2019; pp. 3967–3976. [Google Scholar]
- Chen, D.; Mei, J.P.; Wang, C.; Feng, Y.; Chen, C. Online knowledge distillation with diverse peers. In Proceedings of the AAAI conference on artificial intelligence, New York, NY, USA, 7–12 February 2020; Volume 34, pp. 3430–3437. [Google Scholar]
- Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, Ł.; Polosukhin, I. Attention is all you need. In Proceedings of the Neural Information Processing Systems, Long Beach, CA, USA, 4–9 December 2017; Volume 30, pp. 1–11. [Google Scholar]
- Hou, Q.; Zhou, D.; Feng, J. Coordinate attention for efficient mobile network design. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Online, 19–25 June 2021; pp. 13713–13722. [Google Scholar]
- Woo, S.; Park, J.; Lee, J.Y.; Kweon, I.S. Cbam: Convolutional block attention module. In Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany, 8–14 September 2018; pp. 3–19. [Google Scholar]
- Chen, Y.; Dai, X.; Liu, M.; Chen, D.; Yuan, L.; Liu, Z. Dynamic convolution: Attention over convolution kernels. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 13–19 June 2020; pp. 11030–11039. [Google Scholar]
- Yan, S.; Shao, H.; Wang, J.; Zheng, X.; Liu, B. LiConvFormer: A lightweight fault diagnosis framework using separable multiscale convolution and broadcast self-attention. Expert Syst. Appl. 2024, 237, 121338. [Google Scholar] [CrossRef] [Scilit]
- ISO 11898-1:2015; Road Vehicles—Controller Area Network (CAN)—Part 1: Data Link Layer and Physical Signalling. International Organization for Standardization: Geneva, Switzerland, 2015.
- ISO/SAE 21434:2021; Road Vehicles—Cybersecurity Engineering. International Organization for Standardization: Geneva, Switzerland, 2021.
- ISO 26262:2018; Road Vehicles—Functional Safety. International Organization for Standardization: Geneva, Switzerland, 2018.

















| Scale | Coverage | Target Attack Pattern | Selected |
|---|---|---|---|
| 37.5% | DoS attacks: inter-arrival time collapses (2–3 frames) | Yes | |
| 62.5% | Spoofing: ECU periodicity disrupted (3–5 messages) | Yes | |
| 87.5% | Near-global operation, loses localization | No | |
| >100% | Exceeds window length, physically meaningless | No |
| Stage | Standard Conv () | SMC-Lite | Reduction |
|---|---|---|---|
| Stage 1 () | 288 | 224 | 22.2% |
| Stage 2 () | 1152 | 576 | 50.0% |
| Stage 3 () | 2304 | 1024 | 55.6% |
| Total | 3744 | 1824 | 51.3% |
| Stage | Layer | Configuration | Output | Key Design |
|---|---|---|---|---|
| Input | Feature Engineering | 6-D CAN features | Protocol-specific | |
| Stage 1 | SMC-Lite1 | C: , | Dual-scale fusion | |
| LCTA1 | , Temporal: ON | Full attention | ||
| MaxPool | , stride 2 | Downsample | ||
| Stage 2 | SMC-Lite2 | C: , | Dual-scale fusion | |
| LCTA2 | , Temporal: OFF | Adaptive pruning | ||
| MaxPool | , stride 2 | Downsample | ||
| Stage 3 | SMC-Lite3 | C: , | Dual-scale fusion | |
| LCTA3 | , Temporal: OFF | Adaptive pruning | ||
| Classifier | GAP | Global Average Pooling | 32 | Spatial reduction |
| Dense1 | , ReLU, Dropout(0.3) | 76 | Feature expansion | |
| Dense2 | , ReLU, Dropout(0.2) | 38 | Bottleneck | |
| Output | , Softmax | 5 | 5-class prediction | |
| Total Parameters | 9401 | |||
| Dim | Feature | Protocol Semantics | Attack Detection Role |
|---|---|---|---|
| 1 | ID_Low | ECU identity signature (lower 8 bits) | Spoofing: unauthorized ID injection |
| 2 | ID_Priority | CAN arbitration priority (upper 3 bits) | DoS: ID = 0x000 for bus monopolization |
| 3 | Data_Mean | Payload byte average | Spoofing: abnormal sensor values |
| 4 | Data_Entropy | Payload randomness (Shannon entropy) | Fuzzing: high entropy from random injection |
| 5 | Time_Pattern | Inter-message timing interval | DoS: ; Spoofing: periodicity disruption |
| 6 | Anomaly_Score | Z-score from sliding window | All attacks: composite outlier indicator |
| Specification | Minimum | Recommended | Representative ECU |
|---|---|---|---|
| Flash Memory | 64 KB | 256 KB | Infineon AURIX TC3xx |
| RAM | 16 KB | 64 KB | NXP S32K144 |
| Clock Frequency | 64 MHz | 160 MHz | STM32F407 |
| Architecture | ARM Cortex-M4 | Cortex-M7 | TI TMS570 |
| FPU | Optional (INT8) | Required (FP32) | — |
| Resource | Teacher (FP32) | Teacher (INT8) | Student (FP32) | Student (INT8) |
|---|---|---|---|---|
| Model Size | 36.7 KB | 9.2 KB | 14.3 KB | 3.6 KB |
| Latency @64 MHz | 16.52 ms | 12.8 ms | 7.22 ms | 5.6 ms |
| Peak RAM | ∼8 KB | ∼6 KB | ∼4 KB | ∼3 KB |
| FLOPs | 2.84M | 2.84M | 1.11M | 1.11M |
| Design Goal | Target Threshold | Metric | LCSMC-Net Result |
|---|---|---|---|
| High accuracy | Detection rate >95% | Recall | 99.91% |
| Low false alarms | False alarm rate <1% | 1-Precision | 0.08% |
| Real-time response | Inference <10 ms | Latency | 7.22 ms |
| Model | Type | Params (k) | FLOPs (M) | Acc (%) | F1 | Latency (ms) |
|---|---|---|---|---|---|---|
| Baseline CNN [11] | 1D-CNN | 45.0 | 0.85 | 98.70 | 0.9850 | 4.50 |
| LSTM [13] | RNN | 150.0 | 2.10 | 98.40 | 0.9812 | 18.20 |
| ResNet-18 (1D) | Deep CNN | 3800.0 | 12.50 | 99.10 | 0.9905 | 35.60 |
| MobileNetV2 [28] | Lightweight | 2200.0 | 5.60 | 98.85 | 0.9878 | 12.40 |
| ShuffleNetV2 [29] | Lightweight | 850.0 | 3.40 | 97.20 | 0.9715 | 9.80 |
| EfficientNet-B0 [17] | NAS | 4000.0 | 40.00 | 99.45 | 0.9932 | 45.10 |
| LiConvFormer [38] | Transformer | 15,000.0 | 78.50 | 99.50 | 0.9948 | 120.50 |
| LCSMC-Net (Ours) | Hybrid | 9.4 | 2.84 | 99.89 | 0.9985 | 16.52 |
| Model | Resilience (Protocol) | Adversarial Resistance | Interpretability | Deployment Complexity | Overall Score |
|---|---|---|---|---|---|
| LCSMC-Net | 0.55% drop | Multi-constraint | Protocol | Plug-and-play | 8.9/10 |
| (CAN→CAN-FD) | (6D space) | semantics | (Raw frames) | ||
| LiConvFormer | 1.20% drop | Black-box | Opaque attn. | High (Pretraining) | 5.2/10 |
| (54% worse) | (FGSM/PGD vuln.) | weights | (Weeks data) | ||
| ResNet-18 | 1.50% drop | CAM-based | Learned filters | Medium (Manual) | 6.0/10 |
| (64% worse) | (Gradient attack) | (Indirect) | (Feature eng.) | ||
| EfficientNet-B0 | 2.20% drop | Black-box | Opaque repr. | High (NAS + Transfer) | 4.5/10 |
| (75% worse) | (FGSM/PGD vuln.) | (No CAN map) | (GPU cluster) |
| Variant | Description | Acc. (%) | Acc. | Params |
|---|---|---|---|---|
| Full LCSMC-Net | All components | 99.89 | – | 9401 |
| w/o Multiscale | Single-scale conv ( only) | 98.76 | 8857 | |
| w/o LCTA | No channel-temporal attention | 98.92 | 7529 | |
| w/o Adaptive Pruning | Temporal att. always active | 99.71 | 9433 |
| Fold | Accuracy (%) | F1-Score |
|---|---|---|
| 1 | 99.43 | 0.9943 |
| 2 | 99.41 | 0.9941 |
| 3 | 99.79 | 0.9979 |
| 4 | 99.35 | 0.9935 |
| 5 | 99.70 | 0.9970 |
| Mean ± Std | 99.54 ± 0.17 | 0.9954 ± 0.0017 |
| Model | Complexity | Performance (%) | Efficiency | ||||
|---|---|---|---|---|---|---|---|
| Param | FLOP | Ratio | INT8 | Fidelity | Lat. | Spd. | |
| Teacher | 9401 | 2.84 M | 1.00× | 98.17 | – | 16.52 | 1.00× |
| Student (Scratch) | 3674 | 1.11 M | 2.56× | 95.20 | 96.97 | 7.22 | 2.29× |
| Student (KD) | 3674 | 1.11 M | 2.56× | 97.70 | 99.52 | 7.22 | 2.29× |
| Traffic Class | Attack Characteristics | Accuracy | Security Implication |
|---|---|---|---|
| DoS Attack | High-Freq. Bus Flooding | 98.30% | Critical Defense |
| Fuzzing | Random ID/Payload Injection | 97.80% | High Entropy Detection |
| Impersonation | Spoofed Periodic Signals | 96.90% | Contextual Anomaly |
| Normal | Legitimate ECU Comm. | 98.10% | Low False Alarm Rate |
| Overall | – | 97.70% | Robust Generalization |
| Hyperparameters | Convergence Metrics | Performance | |||
|---|---|---|---|---|---|
| Temp () | Alpha () | Soft Loss | Hard Loss | Acc (%) | Stability |
| 1.0 | 0.7 | 0.25 | 1.58 | 96.45 | Moderate |
| 2.0 | 0.7 | 0.18 | 1.59 | 97.35 | Good |
| 3.0 | 0.7 | 0.19 | 1.65 | 97.42 | Good |
| 4.0 | 0.7 | 0.15 | 1.60 | 97.70 | Optimal |
| 5.0 | 0.7 | 0.16 | 1.62 | 97.12 | Over-smoothed |
| 3.0 | 0.3 | 0.22 | 1.35 | 96.88 | Slow Conv. |
| 3.0 | 0.9 | 0.12 | 2.10 | 97.35 | Unstable |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Hou, M.; Lan, B.; Tang, C.; Huang, J. LCSMC-Net: Lightweight CAN Intrusion Detection via Separable Multiscale Convolution and Attention. Sensors 2026, 26, 1399. https://doi.org/10.3390/s26041399
Hou M, Lan B, Tang C, Huang J. LCSMC-Net: Lightweight CAN Intrusion Detection via Separable Multiscale Convolution and Attention. Sensors. 2026; 26(4):1399. https://doi.org/10.3390/s26041399
Chicago/Turabian StyleHou, Mengdi, Bitie Lan, Chenghua Tang, and Jianbo Huang. 2026. "LCSMC-Net: Lightweight CAN Intrusion Detection via Separable Multiscale Convolution and Attention" Sensors 26, no. 4: 1399. https://doi.org/10.3390/s26041399
APA StyleHou, M., Lan, B., Tang, C., & Huang, J. (2026). LCSMC-Net: Lightweight CAN Intrusion Detection via Separable Multiscale Convolution and Attention. Sensors, 26(4), 1399. https://doi.org/10.3390/s26041399

