This is an early access version, the complete PDF, HTML, and XML versions will be available soon.
Open AccessArticle
Compressing Transformer-Based Sensor Fusion Models for Autonomous Driving Using Curriculum-Based Multi-Task Knowledge Distillation
by
Dhieddine Barhoumi
Dhieddine Barhoumi 1,2,
Stefan Hensel
Stefan Hensel 1
and
Marin B. Marinov
Marin B. Marinov 3,*
1
Institute for Machine Learning and Analytics (IMLA), University of Applied Sciences Offenburg, 77652 Offenburg, Germany
2
National Institute of Applied Sciences and Technology, University of Carthage, Tunis Cedex 1080, Tunisia
3
Department of Electronics, Technical University of Sofia, 1756 Sofia, Bulgaria
*
Author to whom correspondence should be addressed.
Appl. Sci. 2026, 16(18), 9359; https://doi.org/10.3390/app16189359 (registering DOI)
Submission received: 14 July 2026
/
Revised: 9 September 2026
/
Accepted: 15 September 2026
/
Published: 20 September 2026
Abstract
Multimodal sensor fusion architectures such as TransFuser achieve strong waypoint-based driving performance by fusing RGB panoramas and LiDAR bird’s-eye-view (BEV) representations through transformer attention, but their dual backbones and quadratic-cost fusion blocks preclude deployment on resource-constrained platforms. This work proposes LightFuser, a hybrid Performer–Transformer fusion architecture that retains multimodal reasoning at substantially reduced cost. LightFuser halves backbone parameters and FLOPs, applies Performer attention with FAVOR+ at the early, high-resolution fusion stage to obtain linear time and memory complexity, and reserves exact Transformer attention for the mid and late stages, where token counts are small and semantics are richer. A multi-task knowledge distillation framework transfers the decision quality of a high-capacity TransFuser teacher to the lightweight student, organized as a task-ordered curriculum: supervision starts with waypoint control and successively adds depth, semantic segmentation, and BEV objectives. On the CARLA Leaderboard, LightFuser attains a Driving Score of 42.4 (teacher: 47.3) while reducing network latency from 128.4 ms to 67.6 ms per frame. Latency is measured over network modules only and excludes simulator stepping, sensor acquisition, preprocessing, and I/O overhead; the resulting 14.8 FPS sustains a 10 Hz control cycle, meaning near-real-time operation throughout.
Share and Cite
MDPI and ACS Style
Barhoumi, D.; Hensel, S.; Marinov, M.B.
Compressing Transformer-Based Sensor Fusion Models for Autonomous Driving Using Curriculum-Based Multi-Task Knowledge Distillation. Appl. Sci. 2026, 16, 9359.
https://doi.org/10.3390/app16189359
AMA Style
Barhoumi D, Hensel S, Marinov MB.
Compressing Transformer-Based Sensor Fusion Models for Autonomous Driving Using Curriculum-Based Multi-Task Knowledge Distillation. Applied Sciences. 2026; 16(18):9359.
https://doi.org/10.3390/app16189359
Chicago/Turabian Style
Barhoumi, Dhieddine, Stefan Hensel, and Marin B. Marinov.
2026. "Compressing Transformer-Based Sensor Fusion Models for Autonomous Driving Using Curriculum-Based Multi-Task Knowledge Distillation" Applied Sciences 16, no. 18: 9359.
https://doi.org/10.3390/app16189359
APA Style
Barhoumi, D., Hensel, S., & Marinov, M. B.
(2026). Compressing Transformer-Based Sensor Fusion Models for Autonomous Driving Using Curriculum-Based Multi-Task Knowledge Distillation. Applied Sciences, 16(18), 9359.
https://doi.org/10.3390/app16189359
Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details
here.
Article Metrics
Article metric data becomes available approximately 24 hours after publication online.