Object Detection and Scene Perception for Connected and Autonomous Vehicles Using LM-JEPA
Highlights
- A latent-space LM-JEPA framework enables resource-efficient multi-modal object detection and scene perception for connected and autonomous vehicles, achieving higher perception accuracy with lower inference latency compared to conventional LLM and VLM-based methods.
- Context-aware and adaptive sensor fusion, selective latent transmission, and lightweight edge-assisted reasoning improve cooperative scene understanding, yielding up to 25% better scene understanding, 20% higher intersection success rates, and a 15% reduction in transmitted model parameters.
- Latent representation learning provides a practical alternative to token-based LLM and VLM inference, making real-time multi-modal perception feasible on resource-constrained edge devices in connected and autonomous vehicles.
- The proposed framework demonstrates that adaptive latent communication and collaborative reasoning can enhance the scalability, energy efficiency, and safety of future intelligent transportation and cooperative autonomous driving systems.
Abstract
1. Introduction
1.1. Multi-Modal Perception and Vision–Language Models
1.2. Motivation
1.3. Contributions
- We propose a novel latent model-joint embedding predictive architecture (LM-JEPA) for connected and autonomous vehicles for object detection, scene perception, and multi-modal reasoning in a compact latent representation. The proposed framework minimizes dependence on computationally intensive token-level inference while supporting efficient real-time operation.
- We develop a context-aware multi-modal perception pipeline that integrates camera, LiDAR, radar, and high-definition (HD) map information through adaptive sensor fusion, selective latent feature transmission, and dynamic edge-assisted inference. The framework enables collaborative perception under stringent latency, communication bandwidth, and energy constraints.
- We perform extensive experimental evaluation on the BDD100K and nuScenes-QA benchmark datasets to assess both perception and scene understanding capabilities. The proposed LM-JEPA consistently outperforms LLM and VLM-based baselines in detection accuracy, reasoning performance, inference latency, and computational efficiency across diverse urban and highway driving scenarios.
- We demonstrate that latent-space predictive learning provides an effective and scalable foundation for next-generation cooperative autonomous driving systems by improving perception reliability, reducing active model complexity, and enabling efficient multi-modal decision-making suitable for resource-constrained vehicular edge platforms.
2. System Model
3. LM-JEPA-Based Perception Approach
3.1. Context-Aware Multi-Modal Fusion
3.2. Latent Representation Learning
3.3. Mapping Continuous Latent Embeddings to the Finite-State Space
3.4. Prompts for Scene Perception and 3D Reconstruction in Autonomous Driving
3.4.1. Prompts for BDD100K
- Identify all vehicles in the current frame and estimate their relative distances to the ego vehicle.
- Determine whether a pedestrian is approaching the crosswalk ahead.
- Detect traffic lights and identify their current state.
- Evaluate lane boundary structure and determine whether the ego vehicle is centered in its lane.
- Predict whether any vehicle will perform a lane change in the next few seconds.
- Identify potentially occluded objects that may emerge from behind parked vehicles.
- Estimate the speed and trajectory of the vehicle directly ahead of the ego vehicle.
3.4.2. Prompts for nuScenesQA
- How many pedestrians are present in the scene, and what are their motion directions?
- Is a cyclist approaching from the right side of the ego vehicle?
- Which object is closest to the ego vehicle?
- Which traffic infrastructure elements are visible (e.g., traffic lights, road signs)?
- Is any object likely to cross the road in the next 10 s?
- What is the safest maneuver for the ego vehicle given the current scene context?
3.4.3. Prompts for 3D Reconstruction
- Predict a latent-consistent 3D reconstruction of the scene from the preceding 5 s of video.
- Infer the 3D geometry of surrounding buildings and infrastructure from multi-modal inputs comprising camera and LiDAR, including occluded regions.
- Predict dense depth maps for visible and partially occluded objects inside a 50-m radius.
- Reconstruct the 3D trajectories of surrounding agents from past observations, including temporally unobserved segments.
- Estimate the spatial layout of roads, sidewalks, curbs, and lane structures from incomplete sensory input.
4. Results and Discussion
4.1. Experimental Setup
- Coll: Number of collisions per episode
- SpdVar: Variance of vehicle speed (m/s)2, reflecting driving smoothness
- Jerk: Time derivative of acceleration (m/s3), measuring control stability
- Route Completion: Fraction of the route successfully traversed
4.2. Evaluation Metrics
4.3. Object Detection and Scene Perception Performance
4.4. Latency and Computational Efficiency
4.5. Representation Learning vs. Reasoning Capability
4.6. Driving Task Performance
- Without temporal prediction: Removing temporal modeling in JEPA reduces MPE from 1.09 m to 1.35 m.
- Without multi-modal fusion: Using only camera images reduces RA from 0.93 to 0.87.
- Without LLM: QA accuracy drops from 89.5% to 74.3%, demonstrating the importance of semantic reasoning.
4.7. Vision–Language Alignment and Retrieval Performance
4.8. Model Accuracy and Efficiency Trade-Offs
4.9. Perception Task Benchmarking
4.10. Discussion and Limitations
5. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
Abbreviations
| 3D | Three-dimensional |
| 6G | Sixth generation (communication networks) |
| BDD100K | Berkeley DeepDrive (Dataset) |
| GNN | Graph neural network |
| GPT | Generative pretrained transformer |
| GPT-4V | Generative pretrained transformer 4 with vision |
| HD | High-definition |
| LiDAR | Light detection and ranging |
| LLM | Large language model |
| LM-JEPA | Latent model-joint embedding predictive architecture |
| mAP | Mean average precision |
| MLP | Multi-layer perceptron |
| ViT | Vision transformer |
| VLM | Vision–language model |
References
- Khalil, R.A.; Safelnasr, Z.; Yemane, N.; Kedir, M.; Shafiqurrahman, A.; Saeed, N. Advanced Learning Technologies for Intelligent Transportation Systems: Prospects and Challenges. IEEE Open J. Veh. Technol. 2024, 5, 397–427. [Google Scholar] [CrossRef]
- Cheng, X.; Liu, B.; Liu, X.; Liu, E.; Huang, Z. Foundation Model Empowered Synesthesia of Machines (SoM): AI-native Intelligent Multi-Modal Sensing-Communication Integration. IEEE Trans. Netw. Sci. Eng. 2025, 13, 762–782. [Google Scholar]
- Zhou, H.; Hu, C.; Yuan, Y.; Cui, Y.; Jin, Y.; Chen, C.; Wu, H.; Yuan, D.; Jiang, L.; Wu, D.; et al. Large Language Model (LLM) for Telecommunications: A Comprehensive Survey on Principles, Key Techniques, and Opportunities. IEEE Commun. Surv. Tutor. 2025, 27, 1955–2005. [Google Scholar] [CrossRef]
- Zhou, H.; Hu, C.; Yuan, D.; Yuan, Y.; Wu, D.; Chen, X.; Tabassum, H.; Liu, X. Large Language Models for Wireless Networks: An Overview from the Prompt Engineering Perspective. IEEE Wirel. Commun. 2025, 32, 98–106. [Google Scholar] [CrossRef]
- Ferrag, M.A.; Lakas, A.; Tihanyi, N.; Debbah, M. LLM and AI Agents for Autonomous Systems: A Survey of Applications, Datasets, and Security Challenges. IEEE Open J. Intell. Transp. Syst. 2026, 7, 615–657. [Google Scholar] [CrossRef]
- Tian, H.; Reddy, K.; Feng, Y.; Quddus, M.; Demiris, Y.; Angeloudis, P. Large (Vision) Language Models for Autonomous Vehicles: Current Trends and Future Directions. IEEE Trans. Intell. Transp. Syst. 2026, 27, 187–210. [Google Scholar] [CrossRef]
- Chi, F.; Wang, Y.; Nasiopoulos, P.; Leung, V.C. Multi-Agent Collaborative Decision-Making Using Small Vision-Language Models for Autonomous Driving. IEEE Internet Things J. 2025, 12, 55344–55355. [Google Scholar] [CrossRef]
- Xiong, G.; Liu, S.; Yan, Y.; Li, Q.; Li, H. Efficacy of Autonomous Vehicle’s Adaptive Decision-Making Based on Large Language Models Across Multiple Driving Scenarios. IEEE Access 2025, 13, 108076–108092. [Google Scholar] [CrossRef]
- Sharshar, A.; Khan, L.U.; Ullah, W.; Guizani, M. Vision-Language Models for Edge Networks: A Comprehensive Survey. IEEE Internet Things J. 2025, 12, 32701–32724. [Google Scholar] [CrossRef]
- Wang, J.; Ren, H.; Zhu, X.; Ma, Z. Enhancing Autonomous Vehicle Decision-Making Through Policy Transfer With Large Language Model. IEEE Trans. Intell. Transp. Syst. 2025, 1–10. [Google Scholar] [CrossRef]
- Zhu, Y.; Li, Y.; Li, Z.; Li, Z.; Guo, G. Game-Theoretic Decision-Making for Autonomous Vehicles at Unsignalized Intersections under Communication Interferences: A Novel Risk-Adaptive Approach. IEEE Trans. Veh. Technol. 2025, 75, 5531–5540. [Google Scholar]
- Liu, Q.; Tang, Y.; Li, X.; Du, G.; Li, Z. Enhancing the Collaborative Decision-Making Performance of Connected and Autonomous Vehicles: A Multi-Modal Failure-Aware Graph Representation Approach. IEEE Trans. Intell. Transp. Syst. 2025, 26, 6601–6620. [Google Scholar] [CrossRef]
- Cui, Y.; Huang, S.; Zhong, J.; Liu, Z.; Wang, Y.; Sun, C.; Li, B.; Wang, X.; Khajepour, A. DriveLLM: Charting the Path Toward Full Autonomous Driving With Large Language Models. IEEE Trans. Intell. Veh. 2024, 9, 1450–1464. [Google Scholar] [CrossRef]
- Zheng, Y.; Xing, Z.; Zhang, Q.; Jin, B.; Li, P.; Zheng, Y.; Xia, Z.; Chen, Y.; Zhao, D. PlanAgent: A Multi-modal Large Language Agent for Closed-loop Vehicle Motion Planning. IEEE Trans. Cogn. Dev. Syst. 2026, 1–14. [Google Scholar] [CrossRef]
- Deng, Y.; Tu, Z.; Yao, J.; Zhang, M.; Zhang, T.; Zheng, X. TARGET: Traffic Rule-Based Test Generation for Autonomous Driving via Validated LLM-Guided Knowledge Extraction. IEEE Trans. Softw. Eng. 2025, 51, 1950–1968. [Google Scholar] [CrossRef]
- Noh, H.; Shim, B.; Yang, H.J. Adaptive Resource Allocation Optimization Using Large Language Models in Dynamic Wireless Environments. IEEE Trans. Veh. Technol. 2025, 74, 16630–16635. [Google Scholar] [CrossRef]
- Friha, O.; Amine Ferrag, M.; Kantarci, B.; Cakmak, B.; Ozgun, A.; Ghoualmi-Zine, N. LLM-Based Edge Intelligence: A Comprehensive Survey on Architectures, Applications, Security and Trustworthiness. IEEE Open J. Commun. Soc. 2024, 5, 5799–5856. [Google Scholar] [CrossRef]
- Gong, Y.; Zhang, X.; Lu, J.; Jiang, X.; Wang, Z.; Liu, H.; Li, Z.; Wang, L.; Yang, Q.; Wu, X. Steering Angle-Guided Multimodal Fusion Lane Detection for Autonomous Driving. IEEE Trans. Intell. Transp. Syst. 2025, 26, 1470–1481. [Google Scholar] [CrossRef]
- Mortlock, T.; Chen, L.; Smereka, J.M.; Khargonekar, P.; Abdullah Al Faruque, M. Fuse It or Lose It? Analyzing the Effects of Sensor Diversity on Multimodal Ensembles for Autonomous Vehicle Perception. IEEE Trans. Intell. Transp. Syst. 2025, 26, 19833–19844. [Google Scholar] [CrossRef]
- Rafiq, M.; Sung, M.; Rafiq, G.; Sang Choi, G. Camscribe: Enhanced Dashcam Video Descriptions Through Multimodal Spatiotemporal and Object Detection for Autonomous Vehicles. IEEE Access 2025, 13, 90144–90162. [Google Scholar] [CrossRef]
- Huang, S.; Shi, F.; Sun, C.; Zhong, J.; Ning, M.; Yang, Y.; Lu, Y.; Wang, H.; Khajepour, A. DriveSOTIF: Advancing SOTIF Through Multimodal Large Language Models. IEEE Trans. Veh. Technol. 2025, 75, 3642–3655. [Google Scholar]
- Wei, Z.; Lin, B.; Nie, Y.; Chen, J.; Ma, S.; Xu, H.; Liang, X. Unseen From Seen: Rewriting Observation-Instruction Using Foundation Models for Augmenting Vision-Language Navigation. IEEE Trans. Neural Netw. Learn. Syst. 2025, 1, 1–13. [Google Scholar] [CrossRef]
- Wang, S. Graph neural network–driven text classification for fire-door defect inspection in pre-completion construction. Sci. Rep. 2025, 15, 44382. [Google Scholar] [CrossRef] [PubMed]
- Wang, S. Development of an automated transformer-based text analysis framework for monitoring fire door defects in buildings. Sci. Rep. 2025, 15, 43910. [Google Scholar] [CrossRef] [PubMed]
- Hassan, M.; Kabir, M.E.; Jusoh, M.; Ki An, H.; Negnevitsky, M.; Li, C. Large Language Models in Transportation: A Comprehensive Bibliometric Analysis of Emerging Trends, Challenges, and Future Research. IEEE Access 2025, 13, 132547–132598. [Google Scholar] [CrossRef]
- Mohammed, A.; Kora, R. A Comprehensive Overview and Analysis of Large Language Models: Trends and Challenges. IEEE Access 2025, 13, 95851–95875. [Google Scholar] [CrossRef]
- McIntosh, T.R.; Susnjak, T.; Arachchilage, N.; Liu, T.; Xu, D.; Watters, P.; Halgamuge, M.N. Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence. IEEE Trans. Artif. Intell. 2026, 7, 22–39. [Google Scholar] [CrossRef]
- Luo, H.; Sun, G.; Liu, Y.; Zhao, D.; Niyato, D.; Yu, H.; Dustdar, S. A Weighted Byzantine Fault Tolerance Consensus Driven Trusted Multiple Large Language Models Network. IEEE Trans. Cogn. Commun. Netw. 2026, 12, 3815–3830. [Google Scholar] [CrossRef]
- Lee, H.; Zhou, W.; Debbah, M.; Lee, I. On the Convergence of Large Language Model Optimizer for Black-Box Network Management. IEEE Trans. Commun. 2025, 73, 11385–11402. [Google Scholar] [CrossRef]
- Liu, X.; Gao, S.; Liu, B.; Cheng, X.; Yang, L. LLM4WM: Adapting LLM for Wireless Multi-Tasking. IEEE Trans. Mach. Learn. Commun. Netw. 2025, 3, 835–847. [Google Scholar] [CrossRef]
- Kang, J.; Ko, W.; Lee, Y.; Lee, K.; Yun, I. Large Language Model-Based Functional Scenario Generation for Automated Vehicle Safety Evaluation Using Vehicle and Pedestrian Traffic Accident Data. IEEE Access 2025, 13, 167660–167671. [Google Scholar] [CrossRef]
- Li, J.; Wang, Z.; Gong, D.; Wang, C. SCNet3D: Rethinking the Feature Extraction Process of Pillar-Based 3D Object Detection. IEEE Trans. Intell. Transp. Syst. 2025, 26, 770–784. [Google Scholar] [CrossRef]
- Xue, P.; Wu, L.; Yu, Z.; Jin, Z.; Yang, Z.; Li, X.; Yang, Z.; Tan, Y. Automated Commit Message Generation with Large Language Models: An Empirical Study and Beyond. IEEE Trans. Softw. Eng. 2024, 50, 3208–3224. [Google Scholar] [CrossRef]
- Zhao, J.; Wen, T.; Cheong, K.H. Can Large Language Models Be Trusted as Evolutionary Optimizers for Network-Structured Combinatorial Problems? IEEE Trans. Netw. Sci. Eng. 2026, 13, 1191–1206. [Google Scholar] [CrossRef]
- Liu, C.; Zhao, J. Enhancing Stability and Resource Efficiency in LLM Training for Edge-Assisted Mobile Systems. IEEE Trans. Mob. Comput. 2026, 25, 1–18. [Google Scholar] [CrossRef]
- Wu, M.; Li, J.; Ji, J.; Hao, F.; Sun, X.; Ji, R. Evaluating and Mitigating Relationship Hallucinations in Large Vision-Language Models. IEEE Trans. Pattern Anal. Mach. Intell. 2026, 48, 6332–6346. [Google Scholar] [CrossRef] [PubMed]
- Qiang, P.; Tan, H.; Zhang, H.; Li, X.; Li, R.; Liang, J. Mitigating Hallucinations in Large Vision-Language Models via Visual-Enhanced Contrastive Decoding. IEEE Trans. Multimed. 2026, 28, 3242–3255. [Google Scholar] [CrossRef]
- Li, S.; Xu, X.; Meng, W.; Song, J.; Peng, C.; Shen, H.T. Mitigating Hallucinations in Large Vision-Language Models via Reasoning Uncertainty-Guided Refinement. IEEE Trans. Multimed. 2025, 27, 7380–7391. [Google Scholar] [CrossRef]
- Fan, J.; Wu, J.; Chu, H.; Ge, Q.; Gao, B. Hallucination Elimination and Text Annotation Framework for Large Vision-Language Models in Traffic Scenarios. IEEE Trans. Intell. Transp. Syst. 2026, 27, 358–374. [Google Scholar] [CrossRef]
- Dastagir, M.B.A.; Han, D. Towards Hybrid Quantum-Classical Deep Learning Architecture for Indoor-Outdoor Detection Using QCNN-LSTM and Cluster State Signal Processing. IEEE Signal Process. Lett. 2024, 31, 2945–2949. [Google Scholar] [CrossRef]
- Schafer, M.; Nadi, S.; Eghbali, A.; Tip, F. An Empirical Evaluation of Using Large Language Models for Automated Unit Test Generation. IEEE Trans. Softw. Eng. 2024, 50, 1–21. [Google Scholar] [CrossRef]
- Anne, T.; Syrkis, N.; Elhosni, M.; Turati, F.; Legendre, F.; Jaquier, A.; Risi, S. Harnessing Language for Coordination: A Framework and Benchmark for LLM-Driven Multiagent Control. IEEE Trans. Games 2025, 17, 933–943. [Google Scholar] [CrossRef]
- Gupta, A.; Anpalagan, A.; Guan, L.; Khwaja, A.S. Deep learning for object detection and scene perception in self-driving cars: Survey, challenges, and open issues. Array 2021, 10, 100057–100088. [Google Scholar] [CrossRef]
- Padmasiri, H.; Madurawe, R.; Abeysinghe, C.; Meedeniya, D. Automated Vehicle Parking Occupancy Detection in Real-Time. In Proceedings of the 2020 Moratuwa Engineering Research Conference (MERCon), Moratuwa, Sri Lanka, 28–30 July 2020; IEEE: New York, NY, USA, 2020; pp. 1–6. [Google Scholar]
- Qian, T.; Chen, J.; Zhuo, L.; Jiao, Y.; Jiang, Y.G. NuScenes-QA: A Multi-Modal Visual Question Answering Benchmark for Autonomous Driving Scenario. Proc. Conf. Artif. Intell. 2024, 38, 4542–4550. [Google Scholar] [CrossRef]
- Yu, F.; Chen, H.; Wang, X.; Xian, W.; Chen, Y.; Liu, F.; Madhavan, V.; Darrell, T. BDD100K: A Diverse Driving Dataset for Heterogeneous Multitask Learning. In Proceedings of the (IEEE Computer Society Conference on Computer Vision and Pattern Recognition. Online); IEEE: New York, NY, USA, 2020; pp. 2633–2642. [Google Scholar]
- Mayumu, N.; Deng, X.; Bagula, A.; Khan, S.R.; Mukala, P. V2X-JEPA: Self-Supervised Multi-Agent Joint Embedding Predictive Architecture for Robust Vehicle-to-Everything Perception. IEEE Internet Things J. 2026, 13, 16609–16620. [Google Scholar] [CrossRef]








| Reference | Problem Addressed | Proposed Solution Mechanism and Identified Gaps |
|---|---|---|
| [26] | Limited multi-modal perception and visual grounding in autonomous driving | Text-based reasoning with textual input and generated output; absence of visual understanding limits applicability to scene perception |
| [27] | Joint visual and language understanding | Multi-modal alignment with visual features and language embeddings; high computational demand under constrained resources |
| [28] | Lack of explainability and contextual reasoning | Visual question answering using scene input and query; reliance on large-scale infrastructure and limited efficiency |
| [29] | Handling driving scenarios | Language-guided perception; scalability issues and weak integration with planning modules |
| [30] | Limited generalization and prediction capability | End-to-end modeling with temporal input and predicted actions; insufficient focus on computational tradeoffs |
| [31] | VLMs in autonomous vehicles | Task categorization with task groups; lack of integration analysis and real-world constraints |
| [32] | Need for unified taxonomy of models | Model classification; limited discussion on deployment efficiency |
| [33] | Reasoning and knowledge integration, unified perception | Knowledge-enhanced models with structured knowledge; lack of VLM implementation |
| [34] | Perception decision control and error propagation across modules | Multi-modal perception with sensor fusion inputs, localization, object detection, scene understanding, behavior prediction, tracking; weak coordination and limited robustness |
| [35] | Reasoning and generalization, traffic prediction | Integration of foundation models with VLMs and action generation, forecasting mobility patterns; ignores perception and control tasks, high computational cost and deployment challenges |
| [36,37] | Mitigating hallucinations in VLMs | Hallucinations in inter-object relationships |
| [38,39] | Reasoning uncertainty | Visual contrastive decoding for mitigating hallucinations |
| [40] | Integration of heterogeneous sensor modalities under dynamic conditions | Fusion mapping with adaptive weights; difficulty in accurate weight estimation and real-time stability |
| [41] | High computational and memory cost | Compression for efficient inference; degradation in rare scenario performance and limited scalability |
| [42] | Catastrophic forgetting | Regularization and replay with meta learning storage overhead and insufficient scalability for large models |
| Symbol | Definition |
|---|---|
| Cluster | |
| , , …, | Vehicles |
| Adaptive modality weights | |
| Normal driving scenarios | |
| Rare driving scenarios, | |
| Simulation-based augmentation | |
| Prior knowledge from foundation models | |
| Computational cost | |
| Memory usage | |
| Energy consumption | |
| New knowledge | |
| Forgetting loss | |
| Regularization | |
| Experience replay | |
| Sensor observations | |
| Scene embedding | |
| Semantic token space via a mapping | |
| q | Natural language query |
| y | Generated response |
| Ground-truth answer sequence | |
| Discrete token space used to represent latent embeddings | |
| Latent state at iteration u | |
| Stochastic operator | |
| Latent space approximated by a finite set | |
| Transition probability matrix | |
| Probability of transitioning between latent states | |
| Optimal representations | |
| Suboptimal representations | |
| Raw sensor measurements collected by vehicle v at time t | |
| Feature vector capturing both visual and contextual information | |
| v | Vehicle speed |
| Traffic density | |
| Road geometry and lane configuration | |
| Environmental conditions, illumination and weather | |
| Relevance of the sensing modalities | |
| A lightweight scoring framework | |
| Adaptive importance of sensing modality | |
| M | Total number of sensing modalities |
| Encoding of historical context and previous scene annotations | |
| Semantic annotations and predicted actions for surrounding agents | |
| Set of neighboring vehicles in communication range of vehicle v | |
| Latent representation generated by vehicle u | |
| Adaptive fusion weight | |
| Collection of latent features from vehicle u | |
| Learning rate in backpropagation | |
| Training losses | |
| Trainable network parameters |
| Parameter | Value |
|---|---|
| Number of vehicles | 1–100 |
| Number of latent tokens | 1–8 |
| Edge infrastructure range | 100 m–3 km |
| Urban scene coverage area | 4 km × 2 km |
| Communication frequency | 5.9 GHz |
| Inter-vehicle distance | 50–200 m |
| Perception features per frame | 25 |
| Edge node deployment height | 10 m–100 m |
| Road network length | 1–4 km |
| Vehicle speed range | 0–100 km/h |
| Data packet size | 1 byte–3 MB |
| Datasets used | BDD100K, nuScenes-QA |
| Edge buffer size () | 1 GB |
| Edge node transmit power | 20 dBm (100 mW) |
| Receiver sensitivity | dBm |
| Edge node energy budget | 600 kJ |
| Vehicle transmit power | 25 dBm (316.2 mW) |
| Standard deviation in speed | 10 km/h |
| Development platform | Intel Core i7 Laptop |
| System RAM | 8 GB |
| Operating system | Ubuntu Linux |
| Training platform | Amazon EC2 (GPU Instance) |
| Framework | TensorFlow |
| Optimizer | AdamW |
| Learning rate | to |
| Batch size | 32 |
| Training epochs | 100 |
| Input resolution | |
| Learning rate scheduler | Cosine Annealing |
| Encoder | Lightweight Vision Transformer (ViT) |
| Predictor | Multi-Layer Perceptron (MLP) |
| Decoder | Two-Layer MLP |
| BDD100K (Perception) | nuScenes-QA (Reasoning) | ||||
|---|---|---|---|---|---|
| Method | Accuracy (%) | Latency (ms) | Accuracy (%) | Latency (ms) | Params (M) |
| LLM-Based VLM | 78.4 | 44.7 | 76.9 | 44.7 | 380.2 |
| Vision–Language Model | 80.1 | 42.3 | 78.5 | 42.3 | 350.6 |
| Multi-Modal Transformer | 81.7 | 43.8 | 80.2 | 43.8 | 340.9 |
| LM-JEPA | 86.7 | 41.5 | 85.3 | 41.5 | 325.1 |
| Dataset/Task | LLM (GPT-Style Reasoning) | JEPA (Image Representation) | LM-JEPA (Video Model) | ||||||
|---|---|---|---|---|---|---|---|---|---|
| QA Acc. | Parsing | Logic Cov. | Linear Acc. | Low-Shot | Spatial | Action Acc. | Temporal | Consistency | |
| BDD100K: Object Detection | 61.2 | 55.3 | 58.1 | 79.3 | 73.3 | 74.6 | 77.9 | 72.2 | 75.1 |
| BDD100K: Lane Detection | 58.4 | 52.1 | 54.7 | 79.3 | 73.3 | 90.0 | 77.9 | 72.2 | 81.5 |
| BDD100K: Drivable Area | 60.1 | 53.7 | 56.2 | 81.1 | 77.3 | 74.6 | 77.9 | 72.2 | 78.4 |
| BDD100K: Traffic Sign Recognition | 64.5 | 58.6 | 61.3 | 81.1 | 77.3 | 74.6 | 77.9 | 72.2 | 76.2 |
| nuScenes-QA: Scene Understanding | 65.4 | 59.2 | 61.8 | 81.1 | 77.3 | 74.6 | 81.9 | 72.2 | 77.9 |
| nuScenes-QA: Spatial Reasoning | 62.7 | 55.6 | 58.9 | 79.3 | 73.3 | 90.0 | 81.9 | 72.2 | 80.3 |
| nuScenes-QA: Temporal Reasoning | 59.8 | 52.1 | 55.3 | 79.3 | 73.3 | 74.6 | 81.9 | 72.2 | 82.7 |
| nuScenes-QA: View Consistency | 60.9 | 53.7 | 56.8 | 81.1 | 77.3 | 74.6 | 81.9 | 72.2 | 83.5 |
| Mean | 61.6 | 55.0 | 57.9 | 80.3 | 75.2 | 78.5 | 79.9 | 72.2 | 79.5 |
| Median | 60.5 | 53.7 | 56.5 | 80.2 | 75.3 | 74.6 | 79.9 | 72.2 | 79.3 |
| Scen. | Behav. | Cmd | Coll. | SpdVar | Acc | Jerk | Lat. | Score |
|---|---|---|---|---|---|---|---|---|
| Hwy | Overtake | I | 2.88 | 5.02 | 0.24 | 2.64 | 1.68 | 85.05 |
| II | 1.94 | 4.05 | 0.24 | 2.81 | 1.87 | 86.12 | ||
| III | 3.07 | 1.26 | 0.18 | 2.64 | 1.86 | 91.12 | ||
| Base | 3.26 | 2.91 | 0.35 | 2.83 | – | 80.00 | ||
| Follow | I | 6.52 | 0.94 | 0.15 | 2.35 | 1.61 | 87.10 | |
| II | 7.84 | 1.11 | 0.05 | 2.38 | 1.64 | 86.23 | ||
| III | 6.78 | 1.37 | 0.09 | 2.31 | 1.64 | 86.26 | ||
| Base | 4.02 | 0.78 | 0.22 | 2.50 | – | 86.00 | ||
| Right Lane | I | 8.77 | 1.69 | 0.17 | 2.39 | 1.32 | 90.88 | |
| II | 4.54 | 1.18 | 0.15 | 2.44 | 1.83 | 91.51 | ||
| III | 7.29 | 0.23 | 0.13 | 2.61 | 1.22 | 92.18 | ||
| Base | 4.70 | 7.39 | 0.22 | 2.77 | – | 86.00 | ||
| Int. | No Yield | I | 0.89 | 0.29 | 0.26 | 2.28 | 1.65 | 59.67 |
| II | 0.89 | 0.22 | 0.28 | 2.55 | 1.78 | 59.60 | ||
| III | 1.04 | 0.21 | 0.26 | 2.32 | 1.52 | 60.32 | ||
| Base | 1.14 | 0.46 | 0.46 | 2.34 | – | 56.00 | ||
| Yield | I | – | 0.29 | 0.52 | 2.27 | 1.47 | 91.53 | |
| II | – | 0.22 | 0.82 | 2.54 | 1.43 | 89.50 | ||
| III | – | 0.21 | 0.48 | 2.28 | 1.38 | 91.92 | ||
| Base | – | 1.67 | 0.90 | – | – | – |
| Scene Reasoning | Object Counting | Hallucination (Static) | Hallucination (Dynamic) | |||||
|---|---|---|---|---|---|---|---|---|
| Dataset | Accuracy | Score | Accuracy | Score | Accuracy | Score | Accuracy | Score |
| BDD100K | 58.7 | 61.2 | 62.5 | 65.1 | 81.3 | 84.6 | 79.8 | 83.2 |
| nuScenes-QA | 54.9 | 57.6 | 59.1 | 61.8 | 78.5 | 82.1 | 76.4 | 80.7 |
| Model-wise Highlights (LLM/VLM/JEPA) | ||||||||
| InternVL-Chat (VLM) | 59.5 | 62.0 | 63.8 | 66.4 | 83.4 | 86.2 | 81.7 | 84.9 |
| Qwen-VL (VLM) | 57.3 | 60.1 | 62.9 | 65.0 | 82.7 | 85.6 | 80.9 | 83.8 |
| LLaMA-3.1 + Vision (LLM) | 56.8 | 59.4 | 61.2 | 63.5 | 80.6 | 83.9 | 78.8 | 82.4 |
| GPT-4o (LLM) | 60.9 | 63.7 | 65.4 | 68.2 | 85.1 | 88.3 | 83.5 | 86.7 |
| JEPA (1.2B) | 61.5 | 64.3 | 66.9 | 69.8 | 85.7 | 88.9 | 84.6 | 87.5 |
| LM-JEPA (1.6B) | 64.2 | 67.1 | 69.5 | 72.3 | 88.6 | 91.2 | 87.3 | 90.1 |
| Scenario | Model | Coll. | SpdVar | Acc | Jerk | Lat. | Hum. | Scn. | Score |
|---|---|---|---|---|---|---|---|---|---|
| Acceleration | Base | 2.44 | 28.8 | 0.36 | 0.78 | – | 92.0 | 60.0 | 75.6 |
| GPT-4o | 2.52 | 30.8 | 0.39 | 0.83 | 5.82 | 92.9 | 71.5 | 76.4 | |
| LM-JEPA | 2.46 | 30.8 | 0.39 | 0.81 | 1.98 | 96.3 | 60.9 | 76.5 | |
| Lane Change | Base | 2.44 | 3.91 | 1.65 | 0.37 | – | 88.5 | 60.0 | 74.5 |
| GPT-4o | 2.71 | 3.88 | 2.23 | 0.53 | 4.84 | 90.4 | 88.6 | 78.4 | |
| LM-JEPA | 2.15 | 4.07 | 2.15 | 0.41 | 1.83 | 92.2 | 71.9 | 77.5 | |
| Left Turn | Base | – | 1.12 | 7.52 | 0.22 | – | 88.0 | 60.0 | 70.4 |
| GPT-4o | – | 0.93 | 11.5 | 0.29 | 5.23 | 91.3 | 85.0 | 71.4 | |
| LM-JEPA | – | 0.94 | 6.74 | 0.19 | 1.64 | 90.2 | 67.8 | 74.4 |
| Model | Det | Seg | Trk | QA | Hall | Align | Ovrl | Rank |
|---|---|---|---|---|---|---|---|---|
| LLaMA3.1+V | 0.542 | 0.471 | 0.438 | 0.561 | 0.261 | 0.552 | 0.538 | 6.8 |
| Qwen2.5-VL | 0.557 | 0.482 | 0.449 | 0.574 | 0.248 | 0.569 | 0.551 | 5.9 |
| GPT-4o | 0.589 | 0.503 | 0.472 | 0.612 | 0.221 | 0.601 | 0.583 | 4.1 |
| JEPA | 0.601 | 0.517 | 0.486 | 0.629 | 0.204 | 0.618 | 0.596 | 3.2 |
| LM-JEPA | 0.634 | 0.541 | 0.512 | 0.661 | 0.178 | 0.647 | 0.629 | 1.9 |
| Model | mAP | MPE | RA | QA | Plan |
|---|---|---|---|---|---|
| Standard LLM | 45.2 | 3.82 | 0.58 | 61.7 | 52.1 |
| JEPA-only | 87.5 | 1.15 | 0.92 | 74.3 | 78.6 |
| JEPA + LLM | 88.3 | 1.09 | 0.93 | 89.5 | 91.2 |
| Data | Scene Acc. (%) | R@1 (%) | ||||||
|---|---|---|---|---|---|---|---|---|
| IVL | Qwen | LLaMA | GPT4o | IVL | Qwen | LLaMA | JEPA | |
| BDD100K | 61.8 | 64.5 | 62.7 | 66.4 | 47.8 | 50.2 | 48.9 | 59.8 |
| nuScenes-QA | 56.9 | 59.8 | 57.4 | 60.3 | 43.6 | 45.7 | 44.2 | 54.6 |
| Data | IVL | Qwen-VL | LLaMA | GPT4o | Claude | JEPA | LM-JEPA |
|---|---|---|---|---|---|---|---|
| BDD100K | 54.1 | 49.3 | 53.9 | 52.7 | 54.9 | 60.1 | 63.5 |
| nuScenes-QA | 49.5 | 44.8 | 48.7 | 48.3 | 50.2 | 55.8 | 58.9 |
| Model | Det | Seg | Trk | QA | Hall | Align | Ovrl | Rank |
|---|---|---|---|---|---|---|---|---|
| MLP | 0.527 | 0.463 | 0.412 | 0.498 | 0.221 | 0.531 | 0.502 | 9.2 |
| StratLR | 0.563 | 0.491 | 0.438 | 0.521 | 0.204 | 0.566 | 0.531 | 6.4 |
| SwitchEM | 0.571 | 0.498 | 0.446 | 0.533 | 0.198 | 0.574 | 0.539 | 5.9 |
| MinRec | 0.552 | 0.482 | 0.431 | 0.515 | 0.214 | 0.559 | 0.522 | 7.5 |
| SubTab | 0.545 | 0.479 | 0.429 | 0.512 | 0.218 | 0.553 | 0.519 | 8.1 |
| JEPA | 0.598 | 0.521 | 0.469 | 0.562 | 0.183 | 0.602 | 0.556 | 3.8 |
| ResNet | 0.551 | 0.478 | 0.436 | 0.514 | 0.244 | 0.552 | 0.529 | 9.8 |
| FTARL | 0.586 | 0.503 | 0.459 | 0.539 | 0.221 | 0.584 | 0.552 | 6.3 |
| VIME | 0.574 | 0.496 | 0.452 | 0.531 | 0.228 | 0.571 | 0.544 | 7.2 |
| BinRecon | 0.561 | 0.487 | 0.448 | 0.526 | 0.233 | 0.566 | 0.538 | 7.0 |
| SubTab | 0.558 | 0.485 | 0.446 | 0.524 | 0.236 | 0.563 | 0.536 | 7.5 |
| LM-JEPA | 0.623 | 0.537 | 0.491 | 0.584 | 0.198 | 0.628 | 0.577 | 2.6 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Gupta, A.; Sultana, A. Object Detection and Scene Perception for Connected and Autonomous Vehicles Using LM-JEPA. Sensors 2026, 26, 4894. https://doi.org/10.3390/s26154894
Gupta A, Sultana A. Object Detection and Scene Perception for Connected and Autonomous Vehicles Using LM-JEPA. Sensors. 2026; 26(15):4894. https://doi.org/10.3390/s26154894
Chicago/Turabian StyleGupta, Abhishek, and Ajmery Sultana. 2026. "Object Detection and Scene Perception for Connected and Autonomous Vehicles Using LM-JEPA" Sensors 26, no. 15: 4894. https://doi.org/10.3390/s26154894
APA StyleGupta, A., & Sultana, A. (2026). Object Detection and Scene Perception for Connected and Autonomous Vehicles Using LM-JEPA. Sensors, 26(15), 4894. https://doi.org/10.3390/s26154894

