AttriMOT: Semantic-Aware Multimodal 3D Multi-Object Tracking with Attribute-Level Alignment
Abstract
1. Introdution
- To address the limitation of existing 3D MOT methods that rely on coarse category-level semantics and struggle with fine-grained instance discrimination and semantic-aware target retrieval, we propose AttriMOT, a novel semantic-aware multimodal 3D MOT framework. The proposed framework explicitly aligns object-level attributes to enable controllable instance discrimination and text-guided tracking.
- To address the issue that fine-grained attribute information is easily dominated by category-level semantics in existing multimodal representation learning methods, we design a category semantic anchoring and competition suppression mechanism, which preserves discriminative attribute information by treating category embeddings as stable semantic anchors while suppressing their dominant components in the shared embedding space.
- To address the lack of structured attribute-level correspondence modeling in existing cross-modal association methods, which leads to unstable association and weak interpretability, we introduce an attribute-based multimodal alignment paradigm that establishes fine-grained structured correspondences between visual features and textual embeddings for robust and interpretable cross-modal association.
- To address the inability of existing multimodal fusion methods to adapt to dynamically varying modality reliability, we develop a parameter-free adaptive confidence fusion strategy that dynamically balances LiDAR- and camera-derived information through a Softmax mechanism, thereby improving the stability of trajectory confidence estimation.
2. Relation Work
2.1. 3D Multiobject Tracking
2.2. Pre-Trained Vision-Language Models
3. Proposed Method
3.1. Stage I: Semantic Anchoring with Category Competition Suppression
3.2. Stage II: Attribute-Level Fine-Grained Multimodal Alignment
3.3. Stage III: Adaptive Softmax-Based Confidence Fusion
3.4. Stage IV: Semantic-Aware Trajectory Selector
4. Experiment
4.1. Datasets
4.2. Evaluation Metrics
4.3. Experimental Setup
4.4. Experimental Analysis
4.4.1. Comparison with State-of-the-Art Methods
4.4.2. Ablation Experiment
4.4.3. Qualitative Results and Visualization Analysis
5. Discussion
5.1. Analysis of Tracking Performance with the Same 3D Detector
5.2. Discussion on the Semantic-Aware Trajectory Selector
5.3. Discussion on the Different Object Categories
5.4. Sensitivity Analysis of Hyperparameters
6. Conclusions
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- Chaabane, M.; Zhang, P.; Beveridge, J.R.; O’Hara, S. Deft: Detection embeddings for tracking. arXiv 2021, arXiv:2102.02267. [Google Scholar] [CrossRef] [Scilit]
- Wu, H.; Li, Q.; Wen, C.; Li, X.; Fan, X.; Wang, C. Tracklet proposal network for multi-object tracking on point clouds. In Proceedings of the Thirtieth International Joint Conference on Artificial Intelligence (IJCAI-21), Montreal, QC, Canada, 19–27 August 2021; pp. 1165–1171. [Google Scholar]
- Karim, T.; Mahayuddin, Z.R.; Hasan, M.K. Singular and Multimodal Techniques of 3D Object Detection: Constraints, Advancements and Research Direction. Appl. Sci. 2023, 13, 13267. [Google Scholar] [CrossRef] [Scilit]
- Saif, F.M.S.; Mahayuddin, Z.R. Vision based 3D object detection using deep learning: Methods with challenges and applications towards future directions. Int. J. Adv. Comput. Sci. Appl. 2022, 13, 203–214. [Google Scholar] [CrossRef] [Scilit]
- Li, S.; Chen, Z.; Li, H.; Tao, Y.; Gao, Y.; Yan, J. Three-dimensional multiobject tracking based on voxel masking encoder and deep hashing paradigm. IEEE Trans. Neural Netw. Learn. Syst. 2025, 37, 864–877. [Google Scholar] [CrossRef] [Scilit]
- Wang, Y.; Abd Rahman, A.H.; Nor Rashid, F.’A.; Razali, M.K.M. Tackling Heterogeneous Light Detection and Ranging-Camera Alignment Challenges in Dynamic Environments: A Review for Object Detection. Sensors 2024, 24, 7855. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Zhu, Z.; Nie, J.; Wu, H.; He, Z.; Gao, M. MSA-MOT: Multi-stage association for 3D multimodality multi-object tracking. Sensors 2022, 22, 8650. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Wang, X.; Fu, C.; Li, Z.; Lai, Y.; He, J. DeepFusionMOT: A 3D multi-object tracking framework based on camera-LiDAR fusion with deep association. IEEE Robot. Autom. Lett. 2022, 7, 8260–8267. [Google Scholar] [CrossRef] [Scilit]
- Kim, A.; Brasó, G.; Osep, A. PolarMOT: How far can geometric relations take us in 3D multi-object tracking? In Proceedings of the 17th European Conference on Computer Vision, ECCV 2022, Tel Aviv, Israel, 23–27 October 2022; pp. 41–58. [Google Scholar]
- Sangaiah, A.K.; Anandakrishnan, J.; Kumar, S.; Bian, G.-B.; AlQahtani, S.A.; Draheim, D. Point-KAN: Leveraging trustworthy AI for reliable 3-D point cloud completion with Kolmogorov-Arnold networks for 6G-IoT applications. IEEE Internet Things J. 2026, 13, 7801–7814. [Google Scholar] [CrossRef] [Scilit]
- Zhan, Y.; Yuan, Y.; Xiong, Z. Mono3DVG: 3D visual grounding in monocular images. In Proceedings of the Thirty-Eighth AAAI Conference on Artificial Intelligence, Vancouver, BC, Canada, 20–27 February 2024; Volume 38, pp. 6988–6996. [Google Scholar]
- Su, Z.; Adam, A.; Nasrudin, M.F.; Prabuwono, A.S. Proposal-Free Fully Convolutional Network: Object Detection Based on a Box Map. Sensors 2024, 24, 3529. [Google Scholar] [CrossRef] [Scilit]
- Li, X.; Xie, T.; Liu, D.; Gao, J.; Dai, K.; Jiang, Z. Poly-MOT: A polyhedral framework for 3D multi-object tracking. In Proceedings of the IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Detroit, MI, USA, 1–5 October 2023; pp. 9391–9398. [Google Scholar]
- Li, X.; Liu, D.; Wu, Y.; Wu, X.; Zhao, L.; Gao, J. Fast-Poly: A fast polyhedral algorithm for 3D multi-object tracking. IEEE Robot. Autom. Lett. 2024, 9, 10519–10526. [Google Scholar] [CrossRef] [Scilit]
- Zulkifley, M.A.; Rawlinson, D.; Moran, B. Robust Observation Detection for Single Object Tracking: Deterministic and Probabilistic Patch-Based Approaches. Sensors 2012, 12, 15638–15670. [Google Scholar] [CrossRef] [Scilit]
- He, J.; Fu, C.; Wang, X. 3D multi-object tracking based on uncertainty-guided data association. arXiv 2023, arXiv:2303.01786. [Google Scholar]
- Ding, S.; Rehder, E.; Schneider, L.; Cordts, M.; Gall, J. 3DMOTFormer: Graph transformer for online 3D multi-object tracking. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Paris, France, 1–6 October 2023; pp. 9784–9794. [Google Scholar]
- Zhang, Z.; Liu, J.; Xia, Y.; Huang, T.; Han, Q.; Liu, H. LEGO: Learning and graph-optimized modular tracker for online multi-object tracking with point clouds. IEEE Trans. Circuits Syst. Video Technol. 2025, 36, 2419–2432. [Google Scholar] [CrossRef] [Scilit]
- Willes, J.; Reading, C.; Waslander, S.L. InterTrack: Interaction transformer for 3D multi-object tracking. In Proceedings of the Conference on Robots and Vision (CRV), Montreal, QC, Canada, 6–8 June 2023; pp. 73–80. [Google Scholar]
- Nguyen, P.; Quach, K.G.; Duong, C.N.; Phung, S.L.; Le, N.; Luu, K. Multi-camera multi-object tracking on the move via single-stage global association approach. Pattern Recognit. 2024, 152, 110457. [Google Scholar] [CrossRef] [Scilit]
- Huang, K.; Hao, Q. Joint multi-object detection and tracking with camera-LiDAR fusion for autonomous driving. In Proceedings of the 2021 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Prague, Czech Republic, 27 September–1 October 2021; pp. 6983–6989. [Google Scholar]
- Kim, A.; Osep, A.; Leal-Taixé, L. EagerMOT: 3D multi-object tracking via sensor fusion. In Proceedings of the 2021 IEEE International Conference on Robotics and Automation (ICRA), Xi’an, China, 30 May–5 June 2021; pp. 11315–11321. [Google Scholar]
- Stadler, D.; Beyerer, J. ByteV2: Associating more detection boxes under occlusion for improved multi-person tracking. In Proceedings of the International Conference on Pattern Recognition, Montreal, QC, Canada, 21–25 August 2022; Springer: Cham, Switzerland, 2022; pp. 79–94. [Google Scholar]
- Li, H.; Liu, H.; Du, Z.; Chen, Z.; Tao, Y. MCCA-MOT: Multimodal collaboration-guided cascade association network for 3D multi-object tracking. IEEE Trans. Intell. Transp. Syst. 2024, 26, 974–989. [Google Scholar] [CrossRef] [Scilit]
- Mohammed, S.A.K.; Razak, M.Z.A.; Rahman, A.H.A. 3D-DIoU: 3D Distance Intersection over Union for Multi-Object Tracking in Point Cloud. Sensors 2023, 23, 3390. [Google Scholar] [CrossRef] [Scilit]
- Pang, Z.; Li, J.; Tokmakov, P.; Chen, D.; Zagoruyko, S.; Wang, Y. Standing between past and future: Spatio-temporal modeling for multi-camera 3D multi-object tracking. In Proceedings of the 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, BC, Canada, 17–24 June 2023; pp. 17928–17938. [Google Scholar]
- Wang, X.; Fu, C.; He, J.; Huang, M.; Meng, T.; Zhang, S.; Zhou, H.; Xu, Z.; Zhang, C. You only need two detectors to achieve multi-modal 3D multi-object tracking. arXiv 2023, arXiv:2304.08709. [Google Scholar]
- Liu, Y.; Mahayuddin, Z.R.; Nasrudin, M.F. Text-guided spatio-temporal 2D and 3D data fusion for multi-object tracking with RegionCLIP. Appl. Sci. 2025, 15, 10112. [Google Scholar] [CrossRef] [Scilit]
- Liang, M.; Yang, B.; Zeng, W.; Chen, Y.; Hu, R.; Casas, S.; Urtasun, R. PnPNet: End-to-end perception and prediction with tracking in the loop. In Proceedings of the 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 13–19 June 2020; pp. 11553–11562. [Google Scholar]
- Cho, Y.J.; Kim, D. Rethinking multi-object tracking based on re-identification and appearance model management. IEEE Access 2023, 11, 54337–54351. [Google Scholar] [CrossRef] [Scilit]
- Nagy, M.; Werghi, N.; Hassan, B.; Dias, J.; Khonji, M. RobMOT: 3D multi-object tracking enhancement through observational noise and state estimation drift mitigation in LiDAR point clouds. IEEE Trans. Intell. Transp. Syst. 2025, 26, 16047–16059. [Google Scholar] [CrossRef] [Scilit]
- Shtedritski, A.; Rupprecht, C.; Vedaldi, A. What does CLIP know about a red circle? Visual prompt engineering for VLMs. In Proceedings of the 2023 IEEE/CVF International Conference on Computer Vision (ICCV), Paris, France, 1–6 October 2023; pp. 11987–11997. [Google Scholar]
- Chen, P.; Li, Q.; Biaz, S.; Bui, T.; Nguyen, A. GScoreCAM: What objects is CLIP looking at? In Proceedings of the 16th Asian Conference on Computer Vision (ACCV 2022), Macao, China, 4–8 December 2022; pp. 1959–1975. [Google Scholar]
- Jia, C.; Yang, Y.; Xia, Y.; Chen, Y.-T.; Parekh, Z.; Pham, H.; Le, Q.; Sung, Y.-H.; Li, Z.; Duerig, T. Scaling up visual and vision-language representation learning with noisy text supervision. In Proceedings of the 38th International Conference on Machine Learning, ICML 2021, Virtual, 18–24 July 2021; pp. 4904–4916. [Google Scholar]
- Li, J.; Peng, J.; Li, H.; Chen, L. UniCL: A universal contrastive learning framework for large time series models. arXiv 2024, arXiv:2405.10597. [Google Scholar] [CrossRef] [Scilit]
- Bandraupalli, S.; Purwar, A. VLMs-in-the-Wild: Bridging the gap between academic benchmarks and enterprise reality. arXiv 2025, arXiv:2509.06994. [Google Scholar]
- Zhong, Y.; Yang, J.; Zhang, P.; Li, C.; Codella, N.; Li, L.H.; Zhou, L.; Dai, X.; Yuan, L.; Li, Y.; et al. RegionCLIP: Region-based language-image pretraining. In Proceedings of the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA, 18–24 June 2022; pp. 16793–16803. [Google Scholar]
- Li, L.H.; Zhang, P.; Zhang, H.; Yang, J.; Li, C.; Zhong, Y.; Wang, L.; Yuan, L.; Zhang, L.; Hwang, J.-N.; et al. GLIP: Grounded language-image pre-training. In Proceedings of the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA, 18–24 June 2022; pp. 10965–10975. [Google Scholar]
- Zhang, H.; Li, F.; Liu, S.; Zhang, L.; Su, H.; Zhu, J.; Ni, L.M.; Shum, H. DINO: DETR with improved denoising anchor boxes for end-to-end object detection. arXiv 2022, arXiv:2203.03605. [Google Scholar]
- Cai, Z.; Kwon, G.; Ravichandran, A.; Bas, E.; Tu, Z.; Bhotika, R.; Soatto, S. X-DETR: A versatile architecture for instance-wise vision-language tasks. In Proceedings of the 17th European Conference on Computer Vision, ECCV 2022, Tel Aviv, Israel, 23–27 October 2022; pp. 290–308. [Google Scholar]
- Minderer, M.; Gritsenko, A.; Stone, A.; Neumann, M.; Weissenborn, D.; Dosovitskiy, A.; Mahendran, A.; Arnab, A.; Dehghani, M.; Shen, Z.; et al. Simple open-vocabulary object detection. In Proceedings of the 17th European Conference on Computer Vision, ECCV 2022, Tel Aviv, Israel, 23–27 October 2022; pp. 728–755. [Google Scholar]
- Li, B.; Weinberger, K.Q.; Belongie, S.; Koltun, V.; Ranftl, R. Language-driven semantic segmentation. arXiv 2022, arXiv:2201.03546. [Google Scholar] [CrossRef] [Scilit]
- Rao, Y.; Zhao, W.; Chen, G.; Tang, Y.; Zhu, Z.; Huang, G.; Zhou, J.; Lu, J. DenseCLIP: Language-guided dense prediction with context-aware prompting. In Proceedings of the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA, 18–24 June 2022; pp. 18082–18091. [Google Scholar]
- Ghiasi, G.; Gu, X.; Cui, Y.; Lin, T. Scaling open-vocabulary image segmentation with image-level labels. In Proceedings of the 17th European Conference on Computer Vision, ECCV 2022, Tel Aviv, Israel, 23–27 October 2022; pp. 540–557. [Google Scholar]
- Dong, X.; Bao, J.; Zheng, Y.; Zhang, T.; Chen, D.; Yang, H.; Zeng, M.; Zhang, W.; Yuan, L.; Chen, D.; et al. MaskCLIP: Masked self-distillation advances contrastive language-image pretraining. In Proceedings of the 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, BC, Canada, 17–24 June 2023; pp. 10995–11005. [Google Scholar]
- Peng, C.; Zeng, Z.; Gao, J.; Zhou, J.; Tomizuka, M.; Wang, X. PNAS-MOT: Multi-modal object tracking with pareto neural architecture search. IEEE Robot. Autom. Lett. 2024, 9, 4377–4384. [Google Scholar] [CrossRef] [Scilit]
- Zhou, T.; Ye, Q.; Luo, W.; Ran, H.; Shi, Z.; Chen, J. AppTracker+: Displacement uncertainty for occlusion handling in low-frame-rate multiple object tracking. Int. J. Comput. Vis. 2025, 133, 2044–2069. [Google Scholar] [CrossRef] [Scilit]
- Yan, Z.; Feng, S.; Li, X.; Zhou, Y.; Xia, C.; Li, S. S3MOT: Monocular 3D object tracking with selective state space model. arXiv 2025, arXiv:2504.18068. [Google Scholar] [CrossRef] [Scilit]
- Miah, M.; Bilodeau, G.A.; Saunier, N. Learning data association for multi-object tracking using only coordinates. Pattern Recognit. 2025, 160, 111169. [Google Scholar] [CrossRef] [Scilit]
- Gong, Y.; Chen, M.; Liu, H.; Gao, Y.; Yang, L.; Wang, N.; Song, Z.; Ma, H. Stable at any speed: Speed-driven multi-object tracking with learnable Kalman filtering. arXiv 2025, arXiv:2508.00358. [Google Scholar]
- Ninh, P.P.; Kim, H. CollabMOT stereo camera collaborative multi object tracking. IEEE Access 2024, 12, 21304–21319. [Google Scholar] [CrossRef] [Scilit]
- Zhang, W.; Zhou, H.; Sun, S.; Wang, Z.; Shi, J.; Loy, C.C. Robust Multi-Modality Multi-Object Tracking. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Seoul, Republic of Korea, 27 October–2 November 2019. [Google Scholar]
- Wu, H.; Han, W.; Wen, C.; Li, X.; Wang, C. 3d multi-object tracking in point clouds based on prediction confidence-guided data association. IEEE Trans. Intell. Transp. Syst. 2021, 23, 5668–5677. [Google Scholar] [CrossRef] [Scilit]
- Cho, M.; Kim, E. 3D LiDAR multi-object tracking with short-term and long-term multi-level associations. Remote Sens. 2023, 15, 5486. [Google Scholar] [CrossRef] [Scilit]










| Method | Published | HOTA (%) ↑ | DetA (%) ↑ | AssA (%) ↑ | LocA (%) ↑ | MOTA (%) ↑ | IDSW ↓ |
|---|---|---|---|---|---|---|---|
| MSA-MOT [7] | Sensors 2022 | 78.52 | 75.19 | 82.56 | 87.00 | 88.01 | 91 |
| DeepFusion-MOT [8] | RA-L 2022 | 75.46 | 71.54 | 80.05 | 86.70 | 84.63 | 84 |
| PolarMOT [9] | ECCV 2022 | 75.16 | 73.94 | 76.95 | 87.12 | 85.08 | 462 |
| EagerMOT [22] | ICRA 2021 | 74.39 | 75.27 | 74.16 | 87.17 | 87.82 | 239 |
| YONTD-MOT [27] | RA-L 2024 | 78.08 | 74.16 | 82.86 | 88.23 | 85.09 | 42 |
| TG3MOT [28] | Applied Science 2025 | 78.72 | 74.59 | 83.69 | 87.64 | 86.15 | 35 |
| Mono-3D-KF | FUSION 2021 | 75.47 | 74.10 | 77.63 | 85.48 | 88.48 | 162 |
| PNAS-MOT [46] | RA-L 2024 | 67.32 | 77.69 | 58.99 | 86.94 | 89.59 | 751 |
| APPTracker+ [47] | IJCV 2024 | 75.19 | 75.55 | 75.36 | 86.59 | 89.09 | 176 |
| S3MOT [48] | Arxiv 2025 | 76.86 | 76.95 | 77.41 | 87.87 | 86.93 | 543 |
| C-TWIX [49] | Pattern Recognition 2025 | 77.58 | 76.97 | 78.84 | 86.95 | 89.68 | 381 |
| SG-LKF [50] | Arxiv 2025 | 79.59 | 77.27 | 82.53 | 87.09 | 90.55 | 160 |
| CollabMOT [51] | Access 2024 | 75.26 | 75.46 | 75.74 | 86.44 | 89.08 | 227 |
| mmMOT [52] | ICCV 2019 | 62.05 | 72.29 | 54.02 | 86.58 | 83.23 | 733 |
| PC3T [53] | TITS 2021 | 77.80 | 74.57 | 81.59 | 86.07 | 88.81 | 225 |
| 3DMLA [54] | Remote Sensing 2023 | 75.65 | 71.92 | 80.02 | 86.62 | 85.03 | 39 |
| Ours | - | 80.05 | 77.85 | 83.00 | 87.53 | 90.32 | 174 |
| AFMA | SACS | ASCF | HOTA (%) ↑ | DetA (%) ↑ | AssA (%) ↑ | MOTA (%) ↑ | MOTP (%) ↑ | IDSW ↓ | IDF1 (%) ↑ | FPS ↓ |
|---|---|---|---|---|---|---|---|---|---|---|
| - | - | - | 77.785 | 73.846 | 82.083 | 82.354 | 88.514 | 8 | 90.914 | 0.0 |
| ✓ | - | - | 78.626 | 74.293 | 83.340 | 82.057 | 89.652 | 5 | 90.347 | 0.4 |
| - | ✓ | - | 78.214 | 74.018 | 82.947 | 82.481 | 89.103 | 6 | 90.562 | 0.3 |
| - | - | ✓ | 78.037 | 73.925 | 82.811 | 82.436 | 88.947 | 7 | 90.481 | 0.2 |
| ✓ | ✓ | - | 79.068 | 74.748 | 83.772 | 82.700 | 89.485 | 4 | 90.827 | 0.8 |
| - | ✓ | ✓ | 78.843 | 74.516 | 83.514 | 82.618 | 89.337 | 5 | 90.744 | 0.7 |
| ✓ | - | ✓ | 78.957 | 74.603 | 83.635 | 82.674 | 89.421 | 5 | 90.793 | 0.8 |
| ✓ | ✓ | ✓ | 79.257 | 75.067 | 83.817 | 83.138 | 89.486 | 5 | 90.966 | 1.1 |
| Base | AFMA | SACS | ASCF | AMOTA ↑ | MOTA ↑ | MOTP ↑ | IDSW ↓ | RECALL ↑ |
|---|---|---|---|---|---|---|---|---|
| ✓ | - | - | - | 0.742 | 0.610 | 0.274 | 386 | 0.767 |
| ✓ | ✓ | - | - | 0.746 | 0.607 | 0.282 | 350 | 0.769 |
| ✓ | ✓ | ✓ | - | 0.751 | 0.616 | 0.287 | 351 | 0.772 |
| ✓ | ✓ | ✓ | ✓ | 0.757 | 0.624 | 0.291 | 330 | 0.778 |
| Method | 3D Detector | 2D Detector | HOTA (%) ↑ | AssA (%) ↑ | MOTA (%) ↑ | IDSW ↓ | IDF1 (%) ↑ |
|---|---|---|---|---|---|---|---|
| YONTD-MOT | VoxelRCNN | FasterRCNN | 77.52 | 82.14 | 82.11 | 10 | 90.75 |
| PVRCNN | MaskRCNN | 76.49 | 81.12 | 80.35 | 8 | 88.85 | |
| VoxelRCNN | FasterRCNN | 75.38 | 82.33 | 75.47 | 14 | 87.91 | |
| PVRCNN | MaskRCNN | 75.05 | 81.78 | 75.14 | 13 | 87.68 | |
| TG3MOT | VoxelRCNN | RegionClip | 77.78 | 82.08 | 82.35 | 8 | 90.91 |
| PVRCNN | RegionClip | 75.93 | 81.27 | 78.61 | 4 | 89.01 | |
| Ours | VoxelRCNN | RegionClip | 79.26 | 83.82 | 83.14 | 5 | 90.97 |
| PVRCNN | RegionClip | 76.71 | 81.28 | 79.83 | 7 | 88.87 |
| Category | Method | HOTA ↑ | DetA ↑ | MOTA ↑ |
|---|---|---|---|---|
| Pedestrian | Baseline | 52.12 | 53.09 | 68.33 |
| Ours | 53.77 | 54.23 | 69.85 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Liu, Y.; Nasrudin, M.F.; Mahayuddin, Z.R. AttriMOT: Semantic-Aware Multimodal 3D Multi-Object Tracking with Attribute-Level Alignment. Symmetry 2026, 18, 907. https://doi.org/10.3390/sym18060907
Liu Y, Nasrudin MF, Mahayuddin ZR. AttriMOT: Semantic-Aware Multimodal 3D Multi-Object Tracking with Attribute-Level Alignment. Symmetry. 2026; 18(6):907. https://doi.org/10.3390/sym18060907
Chicago/Turabian StyleLiu, Youlin, Mohammad Faidzul Nasrudin, and Zainal Rasyid Mahayuddin. 2026. "AttriMOT: Semantic-Aware Multimodal 3D Multi-Object Tracking with Attribute-Level Alignment" Symmetry 18, no. 6: 907. https://doi.org/10.3390/sym18060907
APA StyleLiu, Y., Nasrudin, M. F., & Mahayuddin, Z. R. (2026). AttriMOT: Semantic-Aware Multimodal 3D Multi-Object Tracking with Attribute-Level Alignment. Symmetry, 18(6), 907. https://doi.org/10.3390/sym18060907

