D3SSTrack: Center-Focused State-Space Modeling for Monocular 3D Multi-Object Tracking
Abstract
1. Introduction
- We propose SST (Solid State Track), a dual-branch temporal modeling block that separates state-space dynamics from spatial gating, enabling robust tracking and efficient spatial–temporal modeling from monocular video input.
- We demonstrate that Mamba sequence modeling improves temporal consistency and identity preservation in 3D tracking tasks, especially in cases of occlusion or fast motion, reducing the total number of identity switches by 9.75% and track fragmentation by 6.6% for the Car category on the KITTI dataset.
- We introduce a contrastive embedding loss that explicitly enforces temporal consistency by pulling together embeddings of the same object across frames while pushing apart distractors and background clutter.
- We evaluate our model on the KITTI 3D tracking benchmark, where it achieves state-of-the-art performance in key metrics, outperforming the best model in sAMOTA by 0.16% and AMOTA by 0.22%, validating the effectiveness of our design.
2. Related Work
2.1. 3D Tracking Across Sensing Modalities
2.2. Spatial Object Perception and Tracking
2.3. Monocular 3D Multi-Object Tracking
3. Proposed Approach
3.1. Preliminaries
3.1.1. Centers of Objects
3.1.2. Solid State Blocks
3.2. Solid State Track
3.3. Contrastive Embedding Loss
4. Experiments
4.1. Setups
4.1.1. Dataset
4.1.2. Metrics
4.1.3. Implementation Details
4.2. Main Results
4.3. Qualitative Results
5. Ablation Study
5.1. Ablation SST DropPath Probability
- Lowering DropPath’s keep probability by begins to hurt performance across all metrics. While regularization is stronger, it appears to overregularize the temporal filtering paths, reducing both tracking accuracy and matching quality.
- At , we obtain the best overall tracking performance, with the highest sAMOTA () and AMOTA (), indicating an optimal balance between information flow and regularization.
- Losing only paths yields the highest localization accuracy (AMOTP = ) due to more consistent feature propagation. However, it leads to slightly lower tracking accuracy, suggesting that regularization is not strong enough to prevent overfitting.
5.2. Ablation SST Block Depth
- Shallow [1, 2, 4, 2]: Achieves the highest inference speed (45 FPS) but lowest tracking accuracy, indicating insufficient capacity for temporal modeling when processing concatenated frame features.
- Normal [1, 3, 8, 4]: Delivers the best tracking metrics (sAMOTA = 97.06%, AMOTA = 49.5%) at a practical inference speed (38 FPS), with eight Mamba-Attention blocks providing optimal capacity for KITTI tracking.
- Deep [1, 3, 12, 4]: Shows marginal localization improvement (AMOTP = 77.19%) but significant speed degradation (25 FPS), suggesting diminishing returns due to overfitting on the relatively small KITTI dataset.
5.3. Ablation Contrastive Embedding Loss
6. Limitations and Future Direction
6.1. Dataset and Class Scope
6.1.1. Evaluation Scope
6.1.2. Single-Class Focus
6.2. Challenging Scenarios
7. Conclusions
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- Sharma, S.; Ansari, J.A.; Murthy, J.K.; Krishna, K.M. Beyond Pixels: Leveraging Geometry and Shape Cues for Online Multi-Object Tracking. In Proceedings of the 2018 IEEE International Conference on Robotics and Automation (ICRA), Brisbane, QLD, Australia, 21–25 May 2018; pp. 3508–3515. [Google Scholar] [CrossRef]
- Hu, H.N.; Yang, Y.H.; Fischer, T.; Darrell, T.; Yu, F.; Sun, M. Monocular quasi-dense 3D object tracking. IEEE Trans. Pattern Anal. Mach. 2023, 45, 1992–2008. [Google Scholar] [CrossRef] [PubMed]
- Hu, H.; Zhu, M.; Li, M.; Chan, K.-L. Deep Learning-Based Monocular 3D Object Detection with Refinement of Depth Information. Sensors 2022, 22, 2576. [Google Scholar] [CrossRef] [PubMed]
- Masoumian, A.; Rashwan, H.A.; Cristiano, J.; Asif, M.S.; Puig, D. Monocular Depth Estimation Using Deep Learning: A Review. Sensors 2022, 22, 5353. [Google Scholar] [CrossRef] [PubMed]
- Shi, Y.; Shen, J.; Sun, Y.; Wang, Y.; Li, J.; Sun, S.; Jiang, K.; Yang, D. SRCN3D: Sparse R-CNN 3D for Compact Convolutional Multi-View 3D Object Detection and Tracking. arXiv 2022, arXiv:2206.14451. [Google Scholar] [CrossRef]
- Hu, H.N.; Cai, Q.Z.; Wang, D.; Lin, J.; Sun, M.; Krahenbuhl, P.; Darrell, T.; Yu, F. Joint monocular 3D vehicle detection and tracking. In Proceedings of the 2019 IEEE/CVF International Conference on Computer Vision (ICCV), Seoul, Republic of Korea, 27–28 October 2019. [Google Scholar] [CrossRef]
- Bradler, H.; Kretz, A.; Mester, R. Urban Traffic Surveillance (UTS): A fully probabilistic 3D tracking approach based on 2D detections. In Proceedings of the 2021 IEEE Intelligent Vehicles Symposium (IV), Nagoya, Japan, 11–17 July 2021. [Google Scholar] [CrossRef]
- Meinhardt, T.; Kirillov, A.; Leal-Taixe, L.; Feichtenhofer, C. TrackFormer: Multi-Object Tracking with Transformers. arXiv 2021. [Google Scholar] [CrossRef]
- Chaabane, M.; Zhang, P.; Beveridge, J.R.; O’Hara, S. DEFT: Detection Embeddings for Tracking. In Proceedings of the 2021 CVPR Conference on Computer Vision and Pattern Recognition, Virtual, 19–25 June 2021. [Google Scholar]
- Li, P.; Jin, J. Time3D: End-to-end joint monocular 3D object detection and tracking for autonomous driving. In Proceedings of the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA, 18–24 June 2022. [Google Scholar] [CrossRef]
- Huang, K.C.; Yang, M.H.; Tsai, Y.H. Delving into motion-aware matching for monocular 3D object tracking. In Proceedings of the 2023 IEEE/CVF International Conference on Computer Vision (ICCV), Paris, France, 1–6 October 2023. [Google Scholar] [CrossRef]
- Zhou, X.; Koltun, V.; Krähenbühl, P. Tracking objects as points. In Proceedings of the 2020 European Conference on Computer Vision, Glasgow, UK, 23–28 August 2020; pp. 474–490. [Google Scholar]
- Yin, T.; Zhou, X.; Krahenbuhl, P. Center-based 3D object detection and tracking. In Proceedings of the 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA, 20–25 June 2021. [Google Scholar] [CrossRef]
- Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, Ł.; Polosukhin, I. Attention is All you Need. In Proceedings of the 2017 Neural Information Processing Systems, Long Beach, CA, USA, 4–9 December 2017. [Google Scholar]
- Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. An image is worth 16 × 16 words: Transformers for image recognition at scale. In Proceedings of the 2021 International Conference on Learning Representations, Virtual, 3–7 May 2021. [Google Scholar]
- Liu, Z.; Lin, Y.; Cao, Y.; Hu, H.; Wei, Y.; Zhang, Z.; Lin, S.; Guo, B. Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the 2021 International Conference on Computer Vision, Virtual, 11–17 October 2021; pp. 10012–10022. [Google Scholar]
- Gu, A.; Goel, K.; Ré, C. Efficiently modeling long sequences with structured state spaces. arXiv 2021, arXiv:2111.00396. [Google Scholar]
- Gu, A.; Dao, T. Mamba: Linear-time sequence modeling with selective state spaces. arXiv 2023, arXiv:2312.00752. [Google Scholar]
- Nguyen, E.; Goel, K.; Gu, A.; Downs, G.; Shah, P.; Dao, T.; Baccus, S.; Re, C. S4nd: Modeling images and videos as multidimensional signals with state spaces. Adv. Neural Inf. Process. Syst. 2022, 35, 2846–2861. [Google Scholar]
- Zhu, L.; Liao, B.; Zhang, Q.; Wang, X.; Liu, W.; Wang, X. Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model. In Proceedings of the 2024 Forty-First International Conference on Machine Learning, Vienna, Austria, 21–27 July 2024. [Google Scholar]
- Liu, Y.; Tian, Y.; Zhao, Y.; Yu, H.; Xie, L.; Wang, Y.; Ye, Q.; Jiao, J.; Liu, Y. VMamba: Visual State Space Model. In 2024 Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2024; pp. 103031–103063. [Google Scholar]
- Hatamizadeh, A.; Kautz, J. Mambavision: A hybrid mamba-transformer vision backbone. In Proceedings of the 2025 Computer Vision and Pattern Recognition Conference, Nashville, TN, USA, 11–15 June 2025; pp. 25261–25270. [Google Scholar]
- Huang, H.-W.; Yang, C.-Y.; Chai, W.; Jiang, Z.; Hwang, J.-N. Exploring learning-based motion models in multi-object tracking. arXiv 2024, arXiv:2403.10826. [Google Scholar]
- Hu, B.; Luo, R.; Liu, Z.; Wang, C.; Liu, W. Trackssm: A general motion predictor by state-space model. arXiv 2024, arXiv:2409.00487. [Google Scholar] [CrossRef]
- Duan, K.; Bai, S.; Xie, L.; Qi, H.; Huang, Q.; Tian, Q. CenterNet: Keypoint Triplets for Object Detection. In Proceedings of the 2019 IEEE/CVF International Conference on Computer Vision (ICCV), Seoul, Republic of Korea, 27 October–2 November 2019; pp. 6568–6577. [Google Scholar] [CrossRef]
- Lin, X.; Pei, Z.; Lin, T.; Huang, L.; Su, Z. Sparse4D v3: Advancing End-to-End 3D Detection and Tracking. arXiv 2023, arXiv:2311.11722. [Google Scholar]
- Giancola, S.; Zarzar, J.; Ghanem, B. Leveraging shape completion for 3D siamese tracking. In Proceedings of the 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA, 15–20 June 2019. [Google Scholar] [CrossRef]
- Fan, Z.; Xu, C.; Chu, M.; Huang, Y.; Ma, Y.; Wang, J.; Xu, Y.; Wu, D. MonoSG: Monocular 3D Object Detection with Stereo Guidance. IEEE Robot. Autom. Lett. 2025, 10, 3604–3611. [Google Scholar] [CrossRef]
- He, C.; Zeng, H.; Huang, J.; Hua, X.S.; Zhang, L. Structure aware single-stage 3D object detection from Point Cloud. In Proceedings of the 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 13–19 June 2020. [Google Scholar] [CrossRef]
- Kim, A.; Ošep, A.; Leal-Taixé, L. agerMOT: 3D Multi-Object Tracking via Sensor Fusion. In Proceedings of the 2021 IEEE International Conference on Robotics and Automation (ICRA), Xi’an, China, 30 May–5 June 2021; pp. 11315–11321. [Google Scholar] [CrossRef]
- Li, P.; Chen, X.; Shen, S. Stereo R-CNN Based 3D Object Detection for Autonomous Driving. In Proceedings of the 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA, 15–20 June 2019; pp. 7636–7644. [Google Scholar] [CrossRef]
- Ding, M.; Huo, Y.; Yi, H.; Wang, Z.; Shi, J.; Lu, Z.; Luo, P. Learning depth-guided convolutions for monocular 3d object detection. In Proceedings of the 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Virtual, 16–18 June 2020; pp. 11672–11681. [Google Scholar]
- Brazil, G.; Liu, X. M3D-RPN: Monocular 3D region proposal network for object detection. In Proceedings of the 2019 IEEE International Conference on Computer Vision (ICCV), Seoul, Republic of Korea, 27 October–2 November 2019. [Google Scholar]
- Wang, Y.; Chao, W.; Garg, D.; Hariharan, B.; Campbell, M.E.; Weinberger, K.Q. Pseudo-lidar from visual depth estimation: Bridging the gap in 3d object detection for autonomous driving. In Proceedings of the 2019 IEEE Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA, 16–20 June 2019; pp. 8445–8453. [Google Scholar]
- Yan, Z.; Feng, S.; Li, X.; Zhou, Y.; Xia, C.; Li, S. S3MOT: Monocular 3D Object Tracking with Selective State Space Model. arXiv 2025, arXiv:2504.18068. [Google Scholar] [CrossRef]
- Liu, Z.; Wu, Z.; Tóth, R. SMOKE: Single-Stage Monocular 3D Object Detection via Keypoint Estimation. In Proceedings of the 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Seattle, WA, USA, 14–19 June 2020; pp. 4289–4298. [Google Scholar] [CrossRef]
- Bergmann, P.; Meinhardt, T.; Leal-Taixe, L. Tracking without bells and whistles. In Proceedings of the 2019 IEEE International Conference on Computer Vision (ICCV), Seoul, Republic of Korea, 27 October–2 November 2019. [Google Scholar]
- Pang, J.; Qiu, L.; Li, X.; Chen, H.; Li, Q.; Darrell, T.; Yu, F. Quasi-dense similarity learning for multiple object tracking. In Proceedings of the 2021 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Virtual, 19–25 June 2021. [Google Scholar]
- He, K.; Fan, H.; Wu, Y.; Xie, S.; Girshick, R. Momentum Contrast for Unsupervised Visual Representation Learning. In Proceedings of the 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 13–19 June 2020; pp. 9726–9735. [Google Scholar] [CrossRef]
- Zhu, X.; Wang, Y.; Dai, J.; Yuan, L.; Wei, Y. Flow-guided feature aggregation for video object detection. In Proceedings of the 2017 International Conference on Computer Vision (ICCV), Venice, Italy, 22–29 October 2017. [Google Scholar]
- Bewley, A.; Ge, Z.; Ott, L.; Ramos, F.; Upcroft, B. Simple online and realtime tracking. In Proceedings of the 2016 IEEE International Conference on Image Processing, Phoenix, AZ, USA, 25–28 September 2016. [Google Scholar]
- Law, H.; Deng, J. Cornernet, Detecting objects as paired keypoints. In Proceedings of the 2018 European Conference on Computer Vision, Munich, Germany, 8–14 September 2018. [Google Scholar]
- Wang, X.; Fu, C.; Li, Z.; Lai, Y.; He, J. DeepFusionMOT: A 3d multi-object tracking framework based on camera-lidar fusion with deep association. IEEE Robot. Autom. Lett. 2022, 7, 8260–8267. [Google Scholar] [CrossRef]
- Shi, S.; Wang, X.; Li, H. PointRCNN: 3D Object Proposal Generation and Detection from Point Cloud. In Proceedings of the 2019 Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA, 16–20 June 2019. [Google Scholar]
- Weng, X.; Wang, J.; Held, D.; Kitani, K. AB3DMOT: A Baseline for 3D Multi-Object Tracking and New Evaluation Metrics. arXiv 2020, arXiv:2008.08063. [Google Scholar]
- Park, D.; Ambrus, R.; Guizilini, V.; Gaidon, J.L.A. Is pseudo-lidar needed for monocular 3d object detection? In Proceedings of the 2021 IEEE/CVF International Conference on Computer Vision, Montreal, BC, Canada, 11–17 October 2021; pp. 3142–3152. [Google Scholar]
- Yu, F.; Wang, D.; Shelhamer, E.; Darrell, T. Deep layer aggregation. In Proceedings of the 2018 Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 18–22 June 2018. [Google Scholar]
- Lin, T.Y.; Goyal, P.; Girshick, R.; He, K.; Doll’ar, P. Focal loss for dense object detection. In Proceedings of the 2017 International Conference on Computer Vision, Venice, Italy, 22–29 October 2017. [Google Scholar]
- Kingma, D.P.; Ba, J. Adam: A method for stochastic optimization. In Proceedings of the 2015 International Conference on Learning Representations, San Diego, CA, USA, 7–9 May 2015. [Google Scholar]
- Rusak, E.; Reizinger, P.; Juhos, A.; Bringmann, O.; Zimmermann, R.; Brendel, W. InfoNCE: Identifying the Gap Between Theory and Practice. arXiv 2024. [Google Scholar] [CrossRef]
- Chen, T.; Kornblith, S.; Norouzi, M.; Hinton, G. A Simple Framework for Contrastive Learning of Visual Representations. In Proceedings of the 2020 International Conference on Machine Learning; PMLR: Cambridge, MA, USA, 2020; Volume 119, pp. 1597–1607. [Google Scholar]
- Geiger, A.; Lenz, P.; Urtasun, R. Are we ready for autonomous driving? The kitti vision benchmark suite. In Proceedings of the 2012 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Providence, RI, USA, 16–21 June 2012. [Google Scholar]
- Wu, H.; Han, W.; Wen, C.; Li, X.; Wang, C. 3d multi-object tracking in point clouds based on prediction confidence-guided data association. IEEE Trans. Intell. Transp. Syst. 2021, 23, 5668–5677. [Google Scholar] [CrossRef]



| Method | sAMOTA (%) ↑ | AMOTA (%) ↑ | AMOTP (%) ↑ | MOTP (%) ↑ | MOTA (%) ↑ | FPS ↑ |
|---|---|---|---|---|---|---|
| CenterTrack [12] (ECCV 2020) | - | - | - | - | 88.7 | 22 |
| DEFT [44] (CVPRw 2021) | - | - | - | - | 88.1 | 13 |
| ine QD-3DT † [2] (PAMI 2022) | 75.94 | 33.69 | 61.87 | 66.61 | - | 6 |
| AB3DMOT † [45] (ECCVw 2020) | 93.28 | 45.43 | 77.41 | 78.43 | 86.24 | 207 * |
| PC3T † [53] | 87.08 | 40.50 | 75.18 | 80.84 | - | 700 * |
| DeepFusionMOT † [43] | 90.51 | 44.87 | 79.47 | 79.83 | - | 104 |
| AB3DMOT ‖ [45] (ECCVw 2020) | 95.78 | 48.74 | 76.22 | 74.56 | - | 207 * |
| PC3T ‖ [53] | 96.52 | 49.22 | 75.23 | 73.39 | - | 700 * |
| DeepFusionMOT ‖ [43] | 96.68 | 49.49 | 76.39 | 74.66 | 88.2 | 104 |
| S3MOT ‖ [35] | 96.96 | 49.73 | 77.25 | 73.25 | 86.93 | 31 |
| D3SSTrack (Ours) | 97.12 | 49.95 | 77.63 | 77.9 | 89.06 | 38 |
| Probability | sAMOTA (%) ↑ | AMOTA (%) ↑ | AMOTP (%) ↑ |
|---|---|---|---|
| 0.75 | 93.43 | 43.18 | 72.84 |
| 0.85 | 97.12 | 49.95 | 77.63 |
| 0.95 | 96.88 | 49.81 | 77.98 |
| Config | Depths | sAMOTA (%) ↑ | AMOTA (%) ↑ | AMOTP (%) ↑ | FPS ↑ |
|---|---|---|---|---|---|
| Shallow | [1, 2, 4, 2] | 96.22 | 48.63 | 74.41 | 45 |
| Normal | [1, 3, 8, 4] | 97.06 | 49.5 | 77.02 | 38 |
| Deep | [1, 3, 12, 4] | 96.89 | 48.73 | 77.19 | 25 |
| Model | sAMOTA (%) ↑ | AMOTA (%) ↑ | AMOTP (%) ↑ | IDSW ↓ | Frag ↓ |
|---|---|---|---|---|---|
| w/o | 95.94 | 48.47 | 76.23 | 46 | 212 |
| Full (w/ ) | 97.12 | 49.95 | 77.63 | 37 | 198 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Firan, D.-O.; Popa, C.-A. D3SSTrack: Center-Focused State-Space Modeling for Monocular 3D Multi-Object Tracking. Mathematics 2026, 14, 1737. https://doi.org/10.3390/math14101737
Firan D-O, Popa C-A. D3SSTrack: Center-Focused State-Space Modeling for Monocular 3D Multi-Object Tracking. Mathematics. 2026; 14(10):1737. https://doi.org/10.3390/math14101737
Chicago/Turabian StyleFiran, Darius-Ovidiu, and Călin-Adrian Popa. 2026. "D3SSTrack: Center-Focused State-Space Modeling for Monocular 3D Multi-Object Tracking" Mathematics 14, no. 10: 1737. https://doi.org/10.3390/math14101737
APA StyleFiran, D.-O., & Popa, C.-A. (2026). D3SSTrack: Center-Focused State-Space Modeling for Monocular 3D Multi-Object Tracking. Mathematics, 14(10), 1737. https://doi.org/10.3390/math14101737

