Next Article in Journal
LLMs and Generative AI for Financial Sentiment Classification: An Explainable Domain-Adaptive Framework
Previous Article in Journal
Digital Resilience in Information Systems: A Systematic Literature Review of Conceptualization, Measurement, and Regulatory Alignment
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Review

Harnessing Multi-Camera Video Fusion: Technologies, Applications, and Future Prospects

1
School of Automotive and Rail Transportation, Luoyang Polytechnic, Luoyang 471099, China
2
Pengcheng Laboratory, Shenzhen 518000, China
*
Author to whom correspondence should be addressed.
Digital 2026, 6(2), 47; https://doi.org/10.3390/digital6020047
Submission received: 24 March 2026 / Revised: 29 May 2026 / Accepted: 10 June 2026 / Published: 12 June 2026

Abstract

The rapid advancement of information technology and multimedia applications has led to an increasing demand for video data processing. In particular, video fusion technology in multi-camera environments, which integrates and optimizes video data from multiple camera viewpoints, plays a crucial role in enhancing visual quality and improving the completeness of information. This technology addresses the challenge of obtaining high-quality video content in complex and dynamic environments. By improving image clarity, expanding perspective information, and enhancing scene understanding, video fusion technology has shown significant potential for a wide range of applications, attracting considerable attention from both academia and industry. Despite the existence of several review articles on video fusion, they tend to focus on isolated aspects of the technology and often lack a comprehensive, systematic overview of the field. To fill this gap, this paper provides an in-depth review of the research on video fusion technology in multi-camera scenarios. The paper covers the definition of video fusion; offers a detailed classification of key technologies, such as geometric correction and alignment, perspective fusion, spatio-temporal fusion, and multi-modal fusion; and explores its applications in diverse fields including surveillance security, virtual reality, film and television production, intelligent transportation, medical imaging, robotics, and unmanned aerial vehicles. Additionally, the paper examines the role of edge caching in video fusion, highlights the current challenges faced by the field, and discusses the potential of video fusion technology for driving innovation across multiple industries.

1. Introduction

Multi-camera video systems are increasingly used in surveillance and security [1,2], virtual and augmented reality (VR/AR) [3,4], intelligent transportation [5,6], and medical imaging [7,8]. Their value comes from complementary viewpoints and modalities, but this also creates specific fusion problems: cameras may have different fields of view, timestamps, exposure conditions, spatial resolutions, and semantic reliability. Video fusion [9,10] addresses these inconsistencies by aligning and integrating multiple streams into task-relevant representations. The main difficulty is therefore not simply collecting more visual data, but determining how geometric alignment, temporal synchronization, modality reliability, and deployment cost should be balanced for a given application.
Definition of video fusion. In this review, video fusion is defined as the process of integrating two or more temporally or spatially related video streams, visual observations, or auxiliary modalities from multiple cameras or sensors into a unified, more informative, and more consistent representation. This process usually involves calibration, geometric alignment, photometric correction, temporal synchronization, and information integration at the pixel, feature, or decision level. Its purpose is not only to improve visual quality, but also to expand the field of view, reduce uncertainty, compensate for occlusion or missing information, and enhance scene understanding in multi-camera or multimodal environments.
Previous studies show that multi-camera and multimodal fusion can improve depth estimation, 3D reconstruction, occlusion handling, target tracking, and autonomous-driving perception [11,12,13,14,15]. However, these benefits depend on camera layout and fusion architecture. Overlapping cameras support geometric reasoning but require accurate calibration and sufficient shared view regions. Non-overlapping cameras are more suitable for wide-area monitoring, but they rely on trajectory continuity, identity association, and camera-topology modeling. Centralized fusion can exploit global information at the cost of bandwidth and latency, whereas distributed fusion improves scalability but makes cross-camera consistency harder to maintain. This trade-off motivates a review that connects fusion algorithms with application and deployment constraints.
Figure 1 illustrates (a) various ways in which fields of view can overlap [16], and (b) an example of a non-overlapping arrangement in multi-camera systems. Its role is to connect camera layout with fusion strategy. Overlapping views support calibration, depth estimation, 3D reconstruction, and cross-view consistency because the same scene points can be observed from multiple angles. Non-overlapping views are more relevant to wide-area monitoring, cross-camera tracking, and trajectory association, where fusion depends less on pixel correspondence and more on temporal continuity, object identity, and camera topology. Thus, the figure is used as a structural guide for distinguishing geometry-oriented fusion from association-oriented fusion.
To clarify the scope of this review and improve the rigor of citation selection, relevant literature was selected through searches in Web of Science, Scopus, IEEE Xplore, ScienceDirect, ACM Digital Library, Google Scholar, and the proceedings or digital libraries of major venues such as CVPR, NeurIPS, ECCV, EMNLP, IEEE INFOCOM, and IEEE TPAMI. The main search period covered studies published from 2018 to 2026, while several earlier representative works were retained only when they provided foundational concepts or widely used methods. The search keywords included “video fusion”, “multi-camera fusion”, “multi-view video”, “multimodal fusion”, “foundation models”, “video understanding”, “geometric correction”, “image stitching”, “spatiotemporal fusion”, “edge video analytics”, “edge caching”, “real-time video processing”, and “privacy-preserving video analysis”. In particular, recent studies from 2024 to 2026 on multimodal foundation models, long-video understanding, camera–LiDAR fusion, edge-AI deployment, 6G-enabled perception, green AI, and privacy-preserving fusion were given special attention to improve the timeliness of the survey. We prioritized peer-reviewed journal and conference papers, highly cited review articles, and recent representative studies that are directly related to video fusion classification, application scenarios, technical challenges, and future research directions. Studies focusing only on single-camera enhancement, lacking an explicit fusion mechanism, or weakly related to multi-camera or multimodal video fusion were excluded.
To avoid loose or redundant citation practice, the related review papers are cited here according to their specific coverage rather than listed as a single undifferentiated group. Yeong et al. [17] are cited for autonomous-vehicle sensor fusion because their review compares cameras, LiDAR, and radar under different driving conditions. Gandhi et al. [18] are used to represent multimodal sentiment analysis, where the focus is datasets, modality interaction, and fusion methods. Kashinath et al. [19] are cited for real-time multi-sensor traffic-flow analysis, while Qi et al. [20] and Majumder et al. [21] are cited only for IoT-enabled activity recognition and vision–inertial human action recognition, respectively. Yao et al. [13] are used for radar–camera perception in autonomous driving. These references are therefore retained because each supports a distinct comparison dimension in Table 1; they are not used as interchangeable evidence. Although these reviews provide valuable insights, they usually focus on a specific sensor combination, task, or application domain and rarely connect the full video fusion pipeline from pre-processing and correction to multi-dimensional fusion, application deployment, edge caching, and future research challenges. In contrast, this paper provides a holistic review of video fusion technology, covering its definition, classification, and application across various domains.
The novelty of this review lies in four aspects. First, it defines video fusion in the context of multi-camera and multimodal environments and organizes the field through a multi-dimensional taxonomy that links geometric correction, perspective fusion, spatio-temporal fusion, and multimodal fusion. Second, it connects low-level pre-processing methods with high-level learning-based fusion strategies, thereby showing how calibration, registration, color correction, deep fusion, adaptive fusion, and foundation-model-based fusion jointly form the video fusion pipeline. Third, it maps fusion technologies to practical application scenarios, including surveillance, VR/AR, film and television production, intelligent transportation, medical imaging, robotics/UAVs, and smart homes. Fourth, it extends the discussion beyond algorithmic accuracy by incorporating edge caching, real-time edge-AI deployment, privacy-preserving fusion, 6G-enabled collaboration, and green AI as deployment-oriented and future-facing dimensions.
Video fusion is treated in this review as a multi-stage system rather than as a single enhancement operation. The analysis follows the technical chain from geometric correction and visual consistency enhancement to perspective fusion, spatio-temporal fusion, multimodal fusion, and edge-assisted deployment. This organization is intended to address two gaps in task-specific reviews: first, many studies discuss individual methods without comparing their assumptions and limitations; second, deployment constraints such as latency, privacy, bandwidth, and energy consumption are often separated from the discussion of fusion algorithms. Therefore, this review compares both algorithmic mechanisms and application bottlenecks across surveillance, VR/AR, film production, intelligent transportation, medical imaging, robotics, UAVs, and smart homes.
In summary, this review is organized to reduce descriptive fragmentation in the literature. Section 1 defines video fusion and clarifies the review scope. Section 2 examines pre-processing and correction methods that determine whether multi-camera data can be reliably aligned. Section 3 compares perspective, spatio-temporal, multimodal, and foundation-model-based fusion according to their assumptions and trade-offs. Section 4 maps these methods to application bottlenecks rather than listing application fields alone. Section 5 analyzes edge caching as a deployment mechanism that couples storage, inference, transmission, and fusion granularity. Section 6 summarizes the resulting roadmap from the perspective of accuracy, latency, privacy, scalability, and energy cost.

2. Video Fusion Technology for Pre-Processing and Correction

This section primarily focuses on the foundational aspects of video fusion technology, specifically addressing the pre-processing and correction methods. These methods, including geometric correction and alignment, as well as color and brightness correction, are essential steps that ensure seamless integration of images captured by multiple cameras. By eliminating geometric discrepancies, adjusting hue and brightness, and ensuring visual consistency across the images, these technologies lay a solid foundation for more advanced video fusion processes.

2.1. Geometric Correction and Alignment

Geometric correction and alignment are critical steps in video fusion, ensuring that images from multiple cameras are accurately integrated. Geometric correction involves transforming and adjusting images captured from various cameras to align them within a common coordinate system. Due to differences in shooting angles, positions, and lens distortions, images from multiple cameras typically exhibit geometric discrepancies. Geometric correction resolves these issues, aligning the images spatially and allowing them to be fused seamlessly.

2.1.1. Camera Calibration

Before performing geometric correction, camera calibration is a prerequisite. Camera calibration determines the internal parameters, such as focal length, principal point coordinates, and lens distortion coefficients, as well as external parameters, such as the camera’s position and orientation. Calibration methods typically involve capturing multiple images using a calibration board, followed by the estimation of the camera parameters through computer vision algorithms. Zhang et al. [22] addressed the visual alignment challenge in video fusion through precise camera calibration. Their approach optimizes both the internal and external camera parameters (e.g., focal length, optical axis offset, position, and direction), ensuring that videos captured from different perspectives can be accurately aligned and integrated. This improves both the accuracy and visual quality of video fusion in multi-camera systems. Sha et al. [23] proposed an end-to-end method for calibrating a single moving camera in dynamic environments, such as sports videos (e.g., basketball). This method includes three key modules: (i) region-based field segmentation, (ii) camera pose estimation using embedded templates, and (iii) homography prediction via a spatial transformation network (STN). These modules are interconnected, enabling end-to-end training and allowing robust and rapid calibration in dynamic settings. Bhardwaj et al. [24] tackled the challenge of traffic camera calibration through the AutoCalib system. This system automatically calibrates a large number of traffic cameras deployed in urban areas, enabling the measurement of real-world distances from video footage. This is crucial for smart city applications such as speeding vehicle detection and urban planning. By using calibrated camera parameters, video streams from different cameras can be integrated to construct a three-dimensional representation or synchronize videos temporally, which is essential for urban traffic monitoring and management. Furthermore, AutoCalib’s automated calibration significantly reduces the need for manual intervention, enhancing both efficiency and accuracy.
Figure 2 illustrates a labeled vehicle-geometry model and an automatic calibration system based on video-frame analysis. The figure is used here to explain why calibration is not only a pre-processing step, but also a measurement mechanism for video fusion. Vehicle dimensions, pitch, roll, yaw, and camera parameters determine whether image-space observations can be converted into real-world distances and aligned across views. Zhang et al. [25] addressed geometric distortion by estimating camera parameters for sports-field calibration, while traffic-camera calibration studies show how geometric extraction and pose estimation support speed measurement and multi-camera traffic analysis. The unresolved limitation is that real-time calibration must remain stable under camera motion, occlusion, weak texture, and nonlinear lens distortion.
Quantitative results further show why calibration is critical for video fusion. In traffic scenarios, the value of Figure 2 lies not only in estimating vehicle pose, but also in converting image measurements into real-world distances that can be used for speed estimation, traffic analysis, and multi-camera alignment. AutoCalib is evaluated on 10 real-world traffic camera feeds and estimated real-world distances with an error of 12% or lower; its mean root mean square (RMS) distance error is 8.98%, compared with 21.59% for a vanishing-point-based baseline, and its average tilt error is 2.04, compared with 4.94 for the baseline [24]. In sport-video calibration, Zhang et al. [25] generated a homography database with more than 91 thousand matrices and evaluated the method on the 2014 World Cup dataset with 209 training images and 186 testing images; the reported mean/median IoU values reached 91.4/94.2 for the whole field and 95.9/97.5 for the visible field region. These results indicate that accurate calibration directly determines the reliability of downstream fusion tasks, especially when geometric measurements, cross-view alignment, or augmented overlays are required.

2.1.2. Image Registration

Image registration is a fundamental step in geometric correction, where images captured by different cameras are aligned spatially to achieve proper overlay. Common methods for image registration include feature point matching and image transformation techniques. Zhi et al. [26] proposed an improved oriented fast and rotated brief (ORB)-based real-time image registration and target localization algorithm leveraging compute unified device architecture (CUDA), addressing the challenge of real-time processing on central processing units (CPUs), especially for high-resolution video images. The algorithm parallelizes the most computationally intensive parts—ORB feature extraction, feature matching using Hamming distance, and the random sample consensus (RANSAC) algorithm for precise matching—thereby accelerating image registration and target localization under the CUDA architecture. This approach efficiently meets the real-time processing needs of high-resolution video images. Ye et al. [27] developed a fast and robust matching framework to tackle multimodal remote sensing image registration, primarily addressing the significant nonlinear intensity differences between images captured by different sensors. The framework effectively handles registration challenges caused by nonlinear intensity variations and geometric deformations between images taken from different sensors or at different times. The paper also introduced an automatic registration system based on this framework, designed for processing large multi-modal remote sensing images, thus addressing real-world engineering applications.
For multi-camera fusion, the unresolved issue in image registration is the balance between cross-view adaptability and computational efficiency. Registration must remain accurate under viewpoint changes, nonlinear intensity differences, and large-scale video streams, but real-time systems cannot rely on expensive global matching for every frame.

2.1.3. Distortion Correction

Distortion, typically caused by the nonlinear characteristics of camera lenses, can manifest as barrel distortion or pincushion distortion. Distortion correction is achieved through specialized algorithms designed to restore the true geometric shape of images, often employing reverse mapping methods. These methods build a distortion model and calculate the correct position of each pixel for proper correction. Gao et al. [28] explored correcting geometric distortions caused by mismatches between image capture, display, and viewing parameters in stereoscopic 3D imaging. They developed a geometric model to account for mismatch parameters, such as camera separation, field of view (FOV), and camera convergence distance. The model adjusts screen distance and camera convergence to eliminate geometric distortions in stereoscopic 3D viewing and compensates for distortions caused by interpupillary distance mismatches. This approach offers a refined method for minimizing distortions in stereoscopic 3D systems.
The main limitation of distortion correction is that geometric rectification may remove lens-induced deformation while still affecting high-frequency details and downstream feature matching. For high-resolution multi-camera fusion, correction models need to preserve local structure while handling complex lens parameters and display-viewing mismatches.

2.2. Color and Brightness Correction

Color and brightness correction in video fusion is critical for adjusting the color and brightness of multiple video sources, ensuring a visually natural and consistent output after fusion. To reduce redundant citation, this subsection uses representative studies for three different photometric problems. Dziembowski et al. [29] are cited for inter-view color correction in immersive video, where global color offsets and color refinement are used to improve visual consistency across views. Li et al. [30] are cited for low-illumination video enhancement because their method targets nighttime surveillance videos through HSV conversion and wavelet-based detail enhancement. Cao et al. [31] are cited for adaptive gamma correction of brightness-distorted images, which is relevant to overexposed and underexposed frames. In this way, each citation is linked to a specific correction problem rather than used as a general reference for all photometric enhancement methods.
While significant progress has been made in color and brightness correction technologies, challenges remain in handling extreme lighting conditions, dynamic scenes, and high-resolution video content. The demand for real-time processing, especially in surveillance and real-time communication applications, also requires algorithms to balance both accuracy and efficiency. Future research will likely focus on leveraging artificial intelligence and machine learning to enhance the adaptability and intelligence of correction algorithms. Additionally, the integration of multi-modal data fusion presents new challenges in how to effectively handle video data from different sensors and platforms, which will be an important direction for future research.
Table 2 summarizes the main pre-processing and correction methods discussed in this section. The table is designed as a quick index, while the surrounding text explains the trade-offs in more detail. Geometric methods, including calibration, registration, and distortion correction, mainly determine whether multi-camera frames can be mapped into a consistent spatial reference. Their main strength is geometric reliability, but they are sensitive to camera movement, occlusion, weak texture, and nonlinear lens distortion. Photometric methods, including color and brightness correction, improve visual consistency across views, but they may introduce flicker, over-enhancement, or color shift in dynamic scenes. Therefore, pre-processing should be selected according to the dominant error source: geometric mismatch, photometric inconsistency, or real-time computational pressure.

3. Video Fusion Technology for Multi-Dimensional Classification and Advanced Application

This section systematically analyzes the multi-dimensional classification and advanced application of video fusion technology. The classification is primarily based on the technical characteristics of fusion processing and application requirements. The three main technological routes of video fusion are perspective fusion, spatio-temporal fusion, and multi-modal fusion. These focus on image stitching, depth information generation, 3D modeling and reconstruction, and the comprehensive processing of spatio-temporal information, respectively. This classification helps clarify the application scenarios and advantages of different technologies, promoting in-depth research and the widespread application of video fusion in fields such as surveillance, VR, and film and television production.

3.1. Perspective Fusion

Perspective fusion involves various essential techniques: image stitching, multi-view stereo fusion, and 3D modeling and reconstruction. These technologies are used to generate panoramic images, create 3D depth information, and reconstruct high-precision 3D models, respectively, thereby enhancing visual effects and broadening application possibilities in VR, AR, and digital heritage preservation.

3.1.1. Image Stitching

Image stitching is a key technique in perspective fusion, enabling the seamless merging of partially overlapping images from multiple viewpoints into a single high-resolution panoramic image. This process requires overcoming challenges such as parallax errors, scene movement, lighting changes, and uneven feature distribution. Abbadi et al. [32] reviewed various panoramic image stitching techniques, focusing on how to merge multiple images into a high-quality panoramic image through feature matching, image alignment, and fusion. The review addresses the key challenges in stitching, including parallax errors, changes in lighting, and scene movement. These techniques provide solutions for each of these issues, optimizing the overall quality of the stitched image. Wang et al. [33] elaborated on image stitching techniques specifically applied to video images. The article introduces the importance of image stitching in various fields, such as remote sensing, aerospace, VR, and medical imaging. Key processes in image stitching, including feature matching using the SIFT algorithm, image registration, and seam removal, are discussed. The feature matching stage involves identifying corresponding points in overlapping areas between two images, while the registration step aligns the images within a common coordinate system. Finally, the seam removal step eliminates visible seams caused by environmental factors such as exposure differences between images. Nie et al. [34] proposed an unsupervised deep image stitching framework that consists of two stages: unsupervised rough image alignment and unsupervised image reconstruction. The framework uses a loss function based on ablation to constrain the affine network for large baseline scenes and introduces a transformer layer to distort the input image. The reconstruction stage uses features to minimize misalignment, thereby improving resolution through a two-branch network designed to refine the image.
The main limitation of image stitching is its dependence on stable overlap and consistent appearance. Dynamic objects, large parallax, weak texture, and exposure variation can turn seam removal into a temporal consistency problem rather than a purely spatial alignment problem.

3.1.2. Multi-View Stereo Fusion

Multi-view stereo fusion combines data from multiple viewpoints to estimate depth and construct 3D representations. Its performance depends on stereo correspondence, view-baseline selection, depth uncertainty, and memory consumption. Voynov et al. [35] explored the incorporation of low-quality sensor depth data into multi-view stereo (MVS) methods and showed that auxiliary depth can improve reconstruction in regions with missing depth values. Cai et al. [36] introduced MFNet, which uses a coarse-to-fine strategy and a multi-level fusion-aware feature pyramid to reduce cost-volume burden while estimating high-resolution depth. Lou et al. [37] proposed the evidence-based local-global fusion (ELF) framework, which estimates uncertainty and fuses local and global evidence for stereo matching.
For multi-view stereo fusion, depth quality is constrained by occlusion, weak texture, view-baseline selection, and memory consumption. The practical bottleneck is whether depth estimation can remain reliable while reducing cost-volume size and inference latency for real-time or edge-assisted reconstruction.

3.1.3. 3D Modeling and Reconstruction

Three-dimensional modeling and reconstruction approaches create three-dimensional representations of objects or scenes from multi-view images, videos, or designed virtual assets. As shown in Figure 3, the visual workflow is relevant to video fusion only when it is connected to measurable reconstruction criteria, such as geometric accuracy, completeness, runtime, and memory cost. Alldieck et al. [38] reconstructed 3D human body models from monocular RGB video by transforming dynamic body postures into a standard reference frame through deformation cancellation, reporting a reconstruction accuracy of 4.5 mm. Sun et al. [39] introduced NeuralRecon, which reconstructs local scene surfaces as sparse truncated signed distance function (TSDF) volumes from monocular video. These methods show that 3D reconstruction should be evaluated as a fusion problem involving view geometry, temporal consistency, and computational efficiency rather than as a visual modeling example alone. González Izard et al. [40] detailed the process of 3D modeling and reconstruction using volume rendering techniques (such as vtkFlyingEdges3D) to generate isosurfaces. The technique then reduces polygon counts using vtkDecimatePro and applies Laplace smoothing filters for a finer 3D model. These reconstructed models can be exported in formats like .obj or .stl for 3D printing and visualization in augmented or virtual reality.
Although Figure 3 is visually illustrative, it corresponds to a measurable 3D reconstruction workflow. For example, Alldieck et al. [38] reconstructed clothed 3D human models from monocular RGB video with a reported reconstruction accuracy of 4.5 mm, and the mean average errors on BUFF, D-FAUST, and KinectCap were 5.37 mm, 4.44 mm, and 3.97 mm, respectively. NeuralRecon further shows the real-time requirement of this pipeline: it reconstructs sparse TSDF volumes from monocular video at 30 ms per key frame, or about 33 key frames per second, which is approximately 10 times faster than Atlas while maintaining competitive 3D geometry quality on ScanNet [39]. These results indicate that 3D reconstruction figures should be read together with accuracy, completeness, F-score, runtime, and memory cost rather than only as visual demonstrations.
For 3D modeling and reconstruction, the key bottleneck is the joint control of surface detail, temporal coherence, and runtime. A method that improves geometric completeness but increases memory use or frame latency may be unsuitable for VR, AR, and online multi-camera applications.

3.2. Spatio-Temporal Fusion

Spatio-temporal fusion involves the comprehensive processing and integration of information from multiple video sources across both spatial and temporal dimensions. This fusion aims to extract richer and more accurate spatio-temporal data to enhance video analysis and understanding. Different levels of fusion methods, including pixel-level fusion, feature-level fusion, and decision-level fusion, can be applied depending on the specific application. Wang et al. [41] proposed a spatio-temporal fusion-based method for mobile target tracking and segmentation, specifically addressing the challenges of target loss during occlusion. The approach combines Kalman filtering with SiamMask technology, establishes a motion model, and employs an elliptical fitting strategy to assess the bounding box angle and size. The use of an attention mechanism further enhances focus on the target’s primary area, minimizing the impact of background distractions. This method demonstrated robust tracking performance in various complex environments. He et al. [42] developed a method for spatio-temporal fusion of multi-source remote sensing images, which integrates linear stretching (Ls), maximum value composition (MVC), and flexible spatio-temporal data fusion (FSDAF). This method aims to map abandoned land distribution in remote sensing images, overcoming challenges related to fragmented terrain and cloud pollution. Liang et al. [43] introduced the spatial–temporal feature fusion enhancement (STFFE) method, which enhances the discriminative ability of video segments by fusing spatial and temporal features. This method leverages the top-k mechanism for feature selection and temporal information fusion to improve anomaly detection accuracy. The approach effectively increases the model’s ability to distinguish between normal and abnormal features, improving the accuracy of anomaly detection in videos. Xu et al. [44] proposed the nonlocal spatial–temporal feature fusion network (NLMF-Net) for real-time infrared small target detection. This model fuses spatio-temporal information in the feature field and improves detection performance while maintaining real-time processing capabilities. By using high-confidence correlation operations between current and past frames, the model achieves a significant performance boost without requiring substantial computational resources. Fu et al. [45] presented Cuboid-Net, a multi-branch convolutional neural network for joint spatio-temporal video super-resolution. The model treats input low-resolution video as a cubic structure, dividing it into slices to feed into different branches for directional processing. The method includes various enhancement modules for feature extraction and quality enhancement, improving the resolution of video frames while maintaining real-time processing capabilities. Jiang et al. [46] introduced the decomposed spatio-temporal fusion graph convolutional network (DSTGCN), a spatio-temporal fusion traffic prediction model based on input traffic signal decomposition. This model uses graph convolutional networks to capture global spatial information and temporal dependencies to predict traffic conditions, offering significant improvements in traffic flow predictions.
Taken together, these studies show that spatio-temporal fusion is not a single algorithmic category, but a set of methods with different assumptions about temporal continuity, spatial correspondence, and computational budget. CNN-based and graph-based methods are efficient when spatial relations are relatively stable, but they may struggle when cross-camera topology changes or when long-range dependencies dominate. Transformer-based and recurrent models can capture longer temporal contexts, but their cost increases with video length, resolution, and camera number. For multi-camera systems, the central trade-off is therefore not only accuracy, but also whether temporal association, spatial alignment, and deployment latency can be optimized jointly.
Multi-camera vehicle tracking provides a representative downstream application scenario for spatio-temporal fusion because it requires not only the spatial association of vehicle appearances across different camera views, but also the temporal integration of trajectories, transition intervals, motion continuity, and cross-camera identity consistency. Therefore, Table 3 is used here as an empirical comparison rather than a simple list of methods. It shows how different designs balance ID accuracy, computational complexity, latency, and scalability on the CityFlowV2 benchmark.
Early high-performing methods mainly adopt multi-stage offline pipelines, including vehicle detection, Re-ID feature extraction, single-camera tracking, and inter-camera trajectory association. Liu et al. [47] introduced crossroad-zone guidance and direction-based temporal masks to constrain cross-camera matching regions and reduce trajectory association ambiguity. Yang et al. [48] further improved inter-camera association by employing box-grained re-ranking matching, achieving an identification F1 (IDF1) score of 84.86%. CityTrack [49] combined location-aware single-camera tracking with box-grained matching, obtaining the highest IDF1 score of 84.91% among the compared methods. These results indicate that explicit spatial–temporal priors and global association can improve identity consistency, but they also create a practical cost: the pipeline depends on strong detectors, heavy Re-ID backbones, and offline trajectory optimization.
Recent studies have shifted toward more deployable, scalable, and low-latency designs for practical intelligent transportation applications. Lin et al. [50] proposed a self-supervised camera-link model to automatically infer camera relationships, thereby reducing dependence on manually designed spatial–temporal constraints and improving scalability in large camera networks, although its tracking accuracy remains lower than heavily optimized offline systems. Huang et al. [51] designed an online multi-camera multi-vehicle tracking framework based on lightweight YOLO11, improved single-camera association, and hierarchical cross-camera clustering, achieving 81.64% IDF1 with low latency and real-time edge deployment capability. Tseng et al. [52] further enhanced YOLOv9 with attention mechanisms to improve vehicle detection robustness and spatial–temporal association performance, achieving 83.44% IDF1 on CityFlowV2. The comparison therefore suggests a deployment-oriented trade-off: offline methods generally obtain stronger identity consistency through richer appearance modeling and global association, whereas online, lightweight, and self-supervised methods sacrifice part of the IDF1 score to reduce latency, manual topology design, and deployment cost.

3.3. Multi-Modal Fusion

Multi-modal fusion integrates various modalities of information from videos, such as visual, audio, and text, to enable more comprehensive and accurate understanding and analysis. This technology enhances recognition, classification, and retrieval capabilities, allowing systems to better understand complex scenes and events in videos. Applications span fields like autonomous driving, surveillance, video search, and sentiment analysis.

3.3.1. Foundation-Model-Based Fusion

Foundation models introduce a new route for multi-modal video fusion by learning reusable representations across text, image, video, audio, sensor, and other data types. Compared with task-specific fusion modules, generalist multi-modal models provide a shared representation space that can potentially align camera streams with language, audio, depth, LiDAR, radar, and contextual signals [53,54,55]. Video-LLaVA, for example, studies unified visual representation learning for image and video inputs before projection into a language model, which is relevant to video fusion because it shifts the fusion objective from simple feature concatenation to cross-modal representation alignment [54]. This ability is particularly useful for multi-camera systems that require both semantic reasoning and cross-modal alignment.
Recent video-oriented large multi-modal models and benchmarks show this trend more clearly. Video-XL reduces the cost of hour-scale video understanding through visual summarization tokens and dynamic compression [56], while Apollo analyzes how video sampling, architecture, data composition, and training schedules affect video large multimodal model (LMM) performance [57]. Video-MME provides a comprehensive benchmark for evaluating multi-modal LLMs in video analysis [58], and LongVideoBench focuses on long-context interleaved video-language understanding [59]. Streaming Long Video Understanding and LongVLM further indicate that long-video reasoning requires memory-efficient temporal modeling and streaming or compressed representations rather than treating all frames equally [60,61]. A recent survey on video temporal grounding with multimodal large language models also shows that temporal localization, language-guided reasoning, and long-range video context have become central issues in current video understanding research [62]. These studies suggest that video fusion is moving from hand-designed concatenation toward foundation-model-based alignment, compression, grounding, and reasoning. The main bottlenecks are computational cost, weak modeling of calibrated cross-view geometry, limited synchronization awareness, and difficult multi-camera data curation. These bottlenecks make lightweight adapters, retrieval-augmented video memory, cross-view pretraining, temporal grounding, and edge-deployable foundation models more relevant than generic model scaling alone.

3.3.2. Deep Fusion

Deep fusion combines modality-specific representations within deep neural networks to improve prediction and scene understanding. Earlier deep-fusion studies are retained only as methodological background, including shared-private DNNs for emotion recognition [63], multi-layer fusion for video classification [64], and split-attention modules for flexible CNN/RNN integration [65]. More recent references are used for current technical directions, such as transformer-based video retrieval [66], adaptive camera–LiDAR fusion [67], and resilient sensor fusion under adverse sensor failures [68]. This citation arrangement separates historical fusion modules from recent robustness-oriented and representation-oriented fusion studies.
To clarify the structural differences among common DNN-based fusion pipelines, Figure 4 separately compares early fusion, intermediate fusion, late fusion, and hybrid fusion. The figure is used as an architectural comparison because the position of modality interaction determines the main trade-off of each pipeline. Early fusion can exploit low-level complementarity but is sensitive to misalignment and noise. Intermediate fusion preserves modality-specific encoders while allowing feature-level interaction. Late fusion is easier to deploy with separate models but may miss fine-grained cross-modal dependencies. Hybrid fusion increases flexibility by combining multiple stages, but it also increases model complexity and training difficulty.
Zhou et al. [69] proposed a feature-level fusion approach that addresses perspective differences between radar and camera features by projecting them into a Bird’s Eye View (BEV) representation. Chen et al. [70] introduced a connected-vehicle-assisted roadside radar and video data fusion framework. As shown in Figure 5, radar and camera streams are not merely combined as parallel inputs; the framework uses connected-vehicle information as calibration data for a backpropagation (BP) network and dynamically updates training samples according to road conditions. The figure therefore illustrates a traffic-specific fusion mechanism in which heterogeneous sensing, communication, calibration, and flow-state estimation are coupled.
As shown in Figure 6, deep fusion can transform heterogeneous observations, including LiDAR points and camera images, into aligned representations for object detection, localization, and scene understanding. Recent camera–LiDAR fusion studies further show that feature fusion should adapt to modality reliability and failure conditions. GAFusion adaptively fuses LiDAR and camera features with multiple guidance cues for 3D object detection [67], while resilient sensor fusion uses multi-modal expert fusion to improve robustness under adverse sensor failures [68]. The figure therefore supports the view that deep fusion is not a simple numerical combination of sensor outputs, but a representation-learning process that must preserve complementary cues while suppressing modality conflict. Its practical value is clear in autonomous driving, robot navigation, and spatial analysis, but real-time deployment still depends on reducing feature dimensionality, controlling cross-modal interference, and simplifying network structures without sacrificing accuracy.
To further provide a task-consistent quantitative comparison, Table 4 summarizes representative camera-, LiDAR-, radar-, and multi-modal 3D object detection methods on the nuScenes benchmark. All entries focus on the same detection task and use the same nuScenes metric family, including mAP, NDS, latency, and detection-error metrics; however, because the reported methods still differ in backbone, implementation details, and training settings, the comparison is intended as a quantitative reference for modality and design trade-offs rather than a strict ranking. Overall, camera–LiDAR fusion achieves the strongest accuracy in this table, with the best entries reaching 73.04% mAP and 72.48% mAP, which reflects the advantage of combining image semantics with LiDAR geometry. In contrast, radar-centered designs are generally faster but less accurate: the radar-only method has the lowest latency of 42 ms but only 14.11% mAP, while the camera–radar method reaches 51.90% mAP and 59.37% NDS at 52 ms, suggesting that radar can provide useful motion and localization cues but remains limited by sparse and noisy measurements. The results also show that using more modalities does not automatically yield better performance, as the camera–LiDAR–radar setting obtains 67.02% mAP and 70.56% NDS with 308 ms latency, which is below the best camera–LiDAR results. Therefore, this comparison supports the view that multi-modal fusion improves 3D detection only when complementary information is effectively aligned, and that practical deployment must balance accuracy, localization error, and latency rather than maximizing the number of input modalities.

3.3.3. Hybrid Fusion

Hybrid fusion combines multiple fusion strategies, integrating information from different modalities at various levels. This approach aims to overcome the limitations of traditional deep and shallow fusion methods, ensuring more comprehensive and robust fusion of multi-modal data. Hybrid fusion can dynamically adjust its fusion strategies based on the specific requirements of the task at hand. Li et al. [80] proposed a Transformer-based video captioning model called MFVC (Multi-modal Fusion for Video Caption), addressing the low performance and high computational complexity of existing multi-modal fusion methods. The introduction of audio modality data and an attention bottleneck module in this model improves performance while reducing operational costs, making the approach more efficient. Joze et al. [81] introduced the multi-modal Transfer module, which achieves slow modality fusion by adding fusion units at various levels of the CNN. This module addresses the issue that traditional deep and shallow fusion methods often fail to fully utilize multi-modal data. It uses squeeze and excitation operations to recalibrate channel features across different CNN streams, facilitating feature fusion across convolutional layers with varying spatial dimensions. Vielzeuf et al. [82] proposed a multi-modal fusion method where each modality is processed by an independent deep convolutional network, with a central network that connects these modality-specific networks. This central network provides common feature embeddings and regularizes modality-specific networks through multi-task learning, improving the accuracy of existing multi-modal fusion methods on multiple computer vision tasks.
Hybrid fusion faces several challenges, including the effective integration of results from fusion at different levels, avoiding conflicts between different fusion methods, and simplifying model structures to reduce the number of parameters. Although hybrid fusion combines different strategies to improve robustness, it may also increase model complexity, making training and optimization more difficult. The large number of adjustable parameters can also complicate the process of finding the optimal combination for the model, requiring careful balancing of efficiency and performance.

3.3.4. Adaptive Fusion

Adaptive fusion refers to methods that adjust the fusion strategy based on the dynamic characteristics of the data being processed. By adapting the fusion strategy in real-time, adaptive fusion models improve the flexibility and generalization of multi-modal fusion systems. Xue et al. [83] proposed DynMM, a dynamic multi-modal fusion method that can adaptively fuse multi-modal data during inference. By introducing gating functions and resource-aware loss functions, DynMM reduces computational costs while maintaining accuracy. This method is particularly effective when processing different types of multi-modal data, enhancing the adaptability of fusion systems. Wu et al. [84] introduced a denoising bottleneck fusion (DBF) model that addresses redundancy and noise issues in multi-modal signals. By utilizing a bottleneck mechanism and a mutual information maximization module, DBF preserves key information while denoising the multi-modal data. This model has shown significant improvements in multi-modal emotion analysis and summarization tasks.
Adaptive fusion is useful because it adjusts modality weighting or fusion paths according to input conditions, but this flexibility introduces a stability–cost trade-off. Dynamic gating can reduce unnecessary computation when modalities are redundant, yet it may also produce unstable decisions under noisy, missing, or conflicting inputs. Therefore, adaptive fusion should be evaluated not only by average accuracy, but also by inference cost, gating stability, and robustness to modality degradation.

3.3.5. Attention-Based Fusion

Attention-based fusion leverages attention mechanisms to selectively focus on the most important information from different modalities. This approach improves the model’s recognition accuracy by highlighting the most relevant features and enhancing the interpretability of the model’s decision-making process. Liu et al. [85] proposed a hierarchical attention-based multi-modal fusion network for video emotion recognition. The network includes a multi-modal feature extraction module and a multi-modal feature fusion module. It solves the problem of emotional cue variation in different video frames by using a local attention network to address emotional differences between frames and a global attention network to handle emotional discrepancies between different modalities. Earlier attention-based fusion studies are cited here only to show the development of modality weighting and selective feature use. Low-rank tensor self-attention fusion provides an example of parameter-efficient multimodal interaction [86], while early video-description work illustrates word-conditioned modality selection [87]. The recent foundation-model-based video understanding studies discussed above are used as the main evidence for current video-language reasoning and long-context multimodal understanding.
Attention-based fusion offers the advantage of identifying and highlighting the most important information from different modalities, improving the model’s accuracy. The attention mechanism also adds interpretability, helping to understand the model’s decision-making process. However, determining an effective attention distribution can be challenging, especially with complex multi-modal data. Designing attention mechanisms that can generalize to different tasks and datasets while improving robustness remains a challenge. Current research focuses on designing attention mechanisms that can adapt to various tasks, improving generalization and robustness.

3.4. Other Types of Fusion

The fusion method proposed in [88] introduces a new approach that does not fit into the traditional categories of fusion discussed earlier. This method presents a neural multi-modal cooperative learning (NMCL) model that explicitly distinguishes and processes consistent and complementary features in multi-modal data. Using a relation-aware attention mechanism, the model separates the consistent and complementary components by learning thresholds. It then integrates the consistent features to enhance the representation and complements the complementary parts to strengthen the information within each modality. This approach specifically addresses the challenges of modal information consistency and complementarity, particularly in the context of micro-video understanding.
A cross-method comparison shows that the three main fusion routes solve different parts of the multi-camera problem and therefore should not be treated as interchangeable techniques. Perspective fusion is geometry-constrained: it is effective for panoramic stitching, 3D reconstruction, and free-viewpoint rendering when calibration is reliable and view overlap is sufficient, but it is vulnerable to parallax, occlusion, exposure inconsistency, and moving objects. Spatio-temporal fusion is association-constrained: it is suitable for tracking, traffic analysis, anomaly detection, and video prediction, but its performance depends on temporal continuity, camera topology, and the ability to prevent error accumulation across long sequences. Multi-modal fusion is representation-constrained: it can combine visual, audio, textual, LiDAR, radar, and semantic cues, but it must handle modality conflict, sensor failure, missing inputs, and unequal reliability across modalities. Foundation-model-based fusion extends multi-modal fusion by improving semantic generalization and language-guided reasoning, yet it introduces new constraints related to cross-view geometry, synchronization awareness, data curation, and edge deployment cost. This comparison indicates that the main research issue is no longer simply which fusion model is more accurate, but under what assumptions a method remains reliable, scalable, and deployable.
Table 5 compares the major fusion strategies discussed in this section. The table uses short keywords to highlight the main differences, while the text provides the interpretation. Perspective fusion is geometry-oriented and is suitable for extending the field of view or reconstructing 3D scenes, but it depends on accurate calibration and sufficient view overlap. Spatio-temporal fusion is temporal-relation-oriented and is useful for tracking, prediction, anomaly detection, and traffic analysis, but the cost increases with video length, resolution, and camera number. Deep, hybrid, adaptive, and attention-based fusion methods are more flexible for multimodal perception, yet they must balance feature richness, modality conflict, model complexity, and deployment cost. Foundation-model-based fusion improves semantic generalization, but current models still need better cross-view geometry, synchronization awareness, and edge efficiency.

4. Application Fields of Video Fusion Technology

4.1. Surveillance and Security

In surveillance and security, video fusion is mainly used to extend spatial coverage, reduce blind spots, associate targets across cameras, and support anomaly detection [89,90]. The key issue is not only whether more cameras improve monitoring accuracy, but whether cross-camera association can remain reliable under occlusion, illumination changes, crowd density, and privacy constraints. Because surveillance data often contain sensitive identities, locations, and trajectories, privacy-aware fusion mechanisms such as federated model fusion and synthetic domain adaptation are becoming increasingly important for multi-camera deployments.
In this application, the core trade-off is between wide-area association accuracy and privacy risk. Stronger cross-camera linkage can improve long-term person or vehicle tracking, but it also requires stricter anonymization, federated learning, or feature-level protection [91]. Thus, surveillance fusion should be evaluated through both recognition performance and privacy exposure rather than through detection accuracy alone.

4.2. Virtual Reality and Augmented Reality

For VR and AR, video fusion is primarily constrained by multi-view consistency, registration accuracy, latency, and visual artifacts [92,93]. As shown in Figure 7, ARSim integrates 3D synthetic objects with real multi-view image data and uses domain adaptation and randomization strategies to address covariate shift between simulated and real scenes. This example shows that VR/AR fusion is not only a rendering problem; it also requires consistent geometry, appearance alignment, and temporal stability across real and virtual views.
The most relevant techniques are panoramic stitching, multi-view consistency modeling, view synthesis, and real-time registration [94]. The main bottleneck is the joint control of parallax artifacts and interaction latency; therefore, robust calibration, neural rendering, and edge-based view synthesis are more important than simply increasing visual resolution.
Figure 8 should also be interpreted as a production workflow rather than a decorative example. Li et al. analyzed the evolution of virtual production from 2009 to 2019 using four representative films, namely Avatar (2009), The Jungle Book (2015), Ready Player One (2018), and The Lion King (2019), based on 16 text sources and 15 video sources [95]. Their analysis shows that virtual production evolved from real-time low-quality CG preview and virtual-camera control to game-engine-based scene scouting, collaborative production design, and LED-wall/live XR workflows. Therefore, the topology in Figure 8 is technically relevant because it reflects the need to jointly manage camera tracking, virtual-scene rendering, real-time compositing, color consistency, and synchronization between physical cameras and virtual content. The figure also clarifies why XR fusion is latency-sensitive: tracking errors or rendering delays can propagate directly into compositing artifacts.

4.3. Film and Television Production

In film and television production, video fusion supports CGI compositing, virtual production, XR stage integration, multi-camera color correction, and AI-assisted visual content generation [95,96]. Compared with general VR/AR applications, production-oriented fusion places stronger emphasis on temporal continuity, camera tracking, lighting consistency, and real-time compositing because visual discontinuities are directly visible in the final content. The practical challenge is therefore visual continuity across cameras, physical lighting, and virtual scenes. Production-oriented fusion should prioritize temporally consistent enhancement, lighting-aware correction, and real-time rendering rather than treating all views as independent video sources.

4.4. Intelligent Transportation

In intelligent transportation systems, video fusion technology is employed to integrate data from cameras, radar, vehicles, and edge nodes. The citations in this paragraph are organized by function. For robust perception, recent radar–camera and RGB–thermal fusion studies are cited because they address adverse weather and multimodal detection reliability [97,98]. For vehicle identification and multi-vehicle sensing, mmWave-radar registration and multi-camera tracking studies are more directly relevant [51,99]. For scalable deployment, multi-edge collaborative data acquisition is cited because it explicitly considers continuous video analytics under edge-resource constraints [100]. This organization avoids mixing traffic-flow estimation, object detection, tracking, and edge deployment as if they were supported by the same type of reference.
For intelligent transportation, current fusion methods combine radar-camera perception, multi-view traffic monitoring, real-time object detection, vehicle re-identification, and collaborative edge perception [97,98,99,100]. The central requirement is low-latency robustness: all-weather detection, cross-intersection tracking, and accident response must be achieved under bandwidth, synchronization, and edge-computing constraints.

4.5. Medical Imaging

In medical imaging, video fusion technology integrates data from various modalities, such as computed tomography (CT), magnetic resonance imaging (MRI), and ultrasound, to provide doctors with a richer set of diagnostic information. Earlier image-fusion studies are retained only as background for traditional multiscale medical image fusion [101,102]. Recent references are used for more specific current issues: medical ultrasound image and video segmentation [8], medical data protection and encrypted retrieval [7,103]. This distinction makes the citation role clearer: older studies support conventional fusion techniques, whereas recent studies support the discussion of video-oriented and privacy-aware medical fusion.
In medical imaging, typical techniques include CT–MRI–ultrasound fusion, multiscale image fusion, cross-modal registration, video segmentation, and privacy-preserving multimodal retrieval. Unlike public-scene video fusion, this field must balance diagnostic accuracy, clinical interpretability, and patient-data protection, especially in cross-institutional retrieval and collaborative diagnosis [8,103].

4.6. Robotics and Unmanned Aerial Vehicles

In the fields of robotics and unmanned aerial vehicles (UAVs), video fusion technology enhances environmental perception by integrating data from multiple cameras and sensors. Earlier work is cited only for representative robotic monitoring and drone-navigation use cases [104,105]. More recent studies are used to support the current deployment discussion: edge-based open-vocabulary perception highlights the accuracy–latency trade-off on mobile robots [106], and on-device multimodal inference indicates how sensing and encoding can be pipelined on edge devices [107]. This separation prevents older application examples from being used as evidence for current edge-AI trends.
For robotics and UAVs, representative techniques include visual–inertial fusion, multi-camera obstacle perception, edge-based detection, and vision-language-model-assisted scene understanding. The dominant constraint is onboard resource limitation: fusion models must support navigation and inspection while controlling latency, memory use, and battery consumption [106,107,108].

4.7. Smart Home

In smart home systems, video fusion combines cameras and household sensors for activity recognition, anomaly detection, elderly care, and context-aware automation [109]. Unlike traffic or public surveillance scenarios, smart-home fusion is constrained by private indoor data, low-power devices, intermittent connectivity, and the need for local decision-making.
Typical techniques therefore include camera-sensor fusion, local inference, activity recognition, and federated learning for privacy-sensitive video analysis [110]. The main design goal is private and low-power perception, so on-device fusion and green AI are more suitable than cloud-heavy pipelines for elderly care, intrusion detection, and context-aware automation.
Table 6 summarizes these application-specific requirements. Overall, surveillance and transportation emphasize coverage, robustness, latency, and scalability; VR/AR and film production emphasize visual consistency and view synthesis; medical imaging emphasizes registration accuracy and privacy compliance; robotics, UAVs, and smart homes emphasize lightweight, energy-efficient, and privacy-preserving fusion. These differences show that fusion strategies should be selected according to application bottlenecks rather than treated as interchangeable modules.

5. Collaboration Between Video Fusion and Edge Caching Technologies

Edge caching is treated in this review not only as a content-delivery technique, but also as a deployment mechanism for multi-camera fusion. Its relevance depends on four linked questions: what information should be cached, where inference should be executed, when data should be transmitted, and what level of representation should be fused. In video fusion systems, edge nodes may cache raw views, key frames, compressed features, object trajectories, depth maps, virtual-view references, or semantic descriptors. They may also execute early detection, feature extraction, view selection, trajectory association, or partial fusion before sending compact information to the cloud. Therefore, the benefit of edge caching is conditional: it reduces latency and bandwidth only when caching, inference, transmission, and fusion granularity are jointly designed.

5.1. Challenges and Potential Solutions of Video Fusion and Edge Caching

5.1.1. Complexity of Data Synchronization and Real-Time Processing

In multi-camera systems, achieving high-precision synchronization of video streams is a central challenge for video fusion. The real-time nature of video data requires edge caching technology to ensure both data consistency and synchronization while minimizing latency [111].
The references in this part are separated according to the deployment issue they support. Mallick et al. [112] are cited for scalable multi-cloud IoT synchronization, which is relevant to timestamp alignment and distributed video streams. Cardoso Nunes et al. [113] are cited for dynamic hybrid synchronization in distributed learning, which is relevant to heterogeneous edge nodes and delayed updates. These works support the synchronization argument, whereas communication-efficient fusion and robust perception are discussed separately below. Timestamp-based synchronization, which uses high-precision clocks to synchronize the time basis across cameras, is a feasible solution. Additionally, content-based synchronization methods, such as image feature matching, can help correct minor temporal differences between video frames, improving synchronization accuracy.
Recent real-time edge fusion studies further indicate that multi-camera systems should avoid transmitting all raw video streams to the cloud. Instead, edge nodes can execute early-stage detection, tracking, feature extraction, and spatial–temporal filtering, while only compact metadata or task-relevant features are forwarded to the central server. Chen et al. [114] are cited here specifically for quantized communication and transferable fusion in multi-agent collaborative perception. Xing et al. [98] are cited for robust real-time multimodal detection under adverse weather. This citation organization distinguishes communication reduction from perception robustness, and both are necessary for practical edge video fusion.

5.1.2. Optimization and Allocation of Computing Resources

Video fusion processes are computationally intensive, involving tasks such as geometric correction, feature extraction, and the application of fusion algorithms. The limited computing resources of edge nodes present a significant challenge in video fusion systems [115].
To improve processing efficiency and reduce costs, it is crucial to distinguish general resource allocation from cooperative inference. Adaptive task scheduling studies are cited for workload allocation across edge resources [116,117]. In contrast, federated inference and collaborative LLM inference are cited for the more recent problem of splitting or ensembling large models across edge devices [108,118]. This distinction reduces redundant references and clarifies why each citation is relevant to edge-assisted video fusion.
The deployment of edge-AI models also requires balancing accuracy, latency, memory footprint, and energy consumption. Le et al. [119] reviewed video anomaly detection for edge-based internet of things (IoT) systems and emphasized that deployment-oriented evaluation should include latency, model size, floating-point operations (FLOPs), and frames per second (FPS) rather than relying only on benchmark accuracy. This perspective is directly relevant to multi-camera video fusion because fusion models must process multiple input streams under strict real-time and resource constraints. Park et al. [106] systematically analyzed real-time open-vocabulary perception on edge devices and quantified the accuracy–latency trade-off of detection and segmentation pipelines on an NVIDIA Jetson platform. These findings indicate that future edge video fusion systems should combine adaptive task offloading, model compression, lightweight feature transmission, and dynamic scheduling so that high-priority streams or safety-critical events receive more computing resources while low-priority data are processed with cheaper approximations.

5.1.3. Real-Time Edge Fusion and Edge-AI Deployment

Real-time edge fusion extends edge caching from passive content placement to active perception and inference near the data source. In multi-camera systems, this shift is important because transmitting all raw streams to the cloud increases bandwidth consumption, end-to-end latency, and privacy exposure. Recent work on 6G-oriented integrated sensing and edge AI suggests that sensing, communication, and inference should be jointly designed so that edge nodes can provide low-latency intelligent perception [120]. Edge perception further emphasizes that wireless sensing and AI processing at the network edge can reduce redundant data transmission and support fast local decisions [121]. For video fusion, these trends imply that calibration, lightweight feature extraction, object-level association, event filtering, and partial fusion can be executed at edge nodes, while only selected features, alerts, or fused representations are transmitted to cloud servers.
On-device multimodal inference also provides a useful reference for future edge video fusion. Huang et al. [107] proposed MMEdge, which accelerates on-device multimodal inference through pipelined sensing and encoding. Although its focus is general multimodal inference, the underlying idea is highly relevant to multi-camera fusion: sensing and encoding should be scheduled as a pipeline rather than treated as isolated stages. In addition, dynamic edge caching for short video services shows that content popularity and crowd distribution can guide cache placement and reduce service latency [122]. These studies indicate that real-time video fusion systems should jointly optimize where to cache, where to infer, when to transmit, and how much information to fuse.
Table 7 summarizes the main deployment constraints that should be considered when multi-camera fusion moves from offline analysis to real-time edge systems. Computational complexity mainly affects whether a fusion model can run on cameras, UAVs, robots, or edge servers. Latency determines whether the fused output can support time-critical applications such as traffic control and public safety. Synchronization affects the spatial and temporal consistency of multi-view fusion. Scalability becomes important when the number of cameras, users, and modalities increases. Energy consumption is especially critical for battery-powered IoT devices and mobile platforms.

5.1.4. Storage and Bandwidth Limitations of Edge Devices

Edge caching reduces dependence on centralized servers, but it also shifts the bottleneck to the storage capacity, bandwidth, and update frequency of edge devices [125,126]. For video fusion, this bottleneck is more complex than ordinary content caching because multi-camera streams are temporally correlated and may be needed for later calibration, tracking, re-identification, or view synthesis.
Storage and bandwidth limitations should therefore be analyzed at the representation level. Caching full-resolution synchronized streams preserves visual fidelity but quickly exhausts edge storage and uplink bandwidth. Caching compressed streams reduces transmission pressure but may degrade downstream calibration, detection, or view synthesis. Caching intermediate features or object trajectories is more efficient for tracking and retrieval, but it makes the system dependent on the upstream detector and may discard information needed for later reprocessing. Advanced video coding methods [127,128], dynamic content placement, and high-speed storage can mitigate these constraints, but they do not remove the basic trade-off among fidelity, latency, reusability, and privacy exposure.

5.1.5. Security and Privacy Protection Issues

With the decentralization of video data processing to edge nodes, security and privacy protection remain necessary deployment requirements for video fusion and edge caching systems. At the current implementation level, the main concerns include data leakage, unauthorized access, insecure transmission, and exposure of sensitive visual content. Therefore, practical systems should adopt end-to-end encryption [129,130], access control, authentication protocols, and video anonymization techniques to protect video data during transmission, storage, and local processing. In addition, compliance with data protection regulations, such as the general data protection regulation (GDPR), is essential for the legal and ethical use of video data [131,132]. More forward-looking privacy-preserving fusion mechanisms are discussed later in the Section 5.3 as emerging research challenges.

5.2. Applications of Edge Caching Technology in Video Fusion

In the literature, Zhang et al. [133] introduced an edge-assisted free viewpoint video (FVV) system called Edge-FVV, which provides a concrete case for analyzing the mechanism rather than merely claiming that edge caching reduces latency. In this system, edge caches are placed between servers and users, and each user request must be mapped to cached reference views, network downloading, and virtual-view synthesis. When a requested viewpoint is not directly recorded by a camera, adjacent reference streams must be retrieved and synthesized, so the total latency depends jointly on cache placement, bandwidth, synthesis time, and request allocation.
Figure 9 illustrates this three-tier architecture comprising servers, edge caches, and user devices. Its mechanism can be interpreted as a joint optimization problem. First, reference views must be selected and cached under limited edge capacity. Second, user requests must be assigned to edge caches according to network conditions and cache availability. Third, virtual-view synthesis must be scheduled so that computation delay does not offset the communication benefit of caching. This architecture is therefore relevant to multi-camera video fusion because it connects view selection, edge storage, transmission delay, and synthesis computation within one deployment pipeline.
The reported evaluation gives empirical support for this mechanism. In the Edge-FVV experiments, processing delay was modeled by considering the number of edge caches, bandwidth, GPU concurrency, total users, and synthesis time. Under representative settings such as G = 10 , B = 10 , T v = 0.1 or 1, and U = 256 or 1024, cache deployment showed different behaviors: when synthesis time was small, adding more caches could create a non-monotonic latency curve, whereas when synthesis time dominated, additional caches produced diminishing returns. For request allocation, the distributed multi-armed-bandit algorithms reduced total processing time by 4.2–7.4% over benchmarks, while the centralized DQN-based allocation reduced it by 4.6–6.8% [133]. These results show that edge caching benefits video fusion only when communication, computation, and request allocation are optimized together; caching alone does not guarantee lower latency.
The limitations of this example are also instructive. Edge-FVV focuses on FVV request latency, so its conclusions cannot be directly generalized to all video fusion tasks such as detection-level fusion, cross-camera tracking, or privacy-preserving analytics. Virtual-view interpolation may reduce reference-view demand, but it introduces additional computation and possible artifacts [134,135]. The distributed design also increases system complexity because cache states, user requests, network conditions, and synthesis resources must be coordinated. Thus, the main academic value of Edge-FVV is not simply that it uses edge caches, but that it exposes the coupling among cache placement, view synthesis, request allocation, and latency. This coupling is the mechanism that future edge-assisted video fusion systems need to model explicitly.

5.3. Future Prospects and Research Roadmap

The future roadmap of multi-camera video fusion is shaped by the tension between algorithmic capability and deployment constraints. As summarized in Table 8, the main directions are not independent topics but coupled design problems. Foundation-model-driven fusion improves semantic generalization, but it must be adapted to calibrated cross-view geometry and real-time inference. Edge fusion reduces cloud dependence, but its benefit depends on where features are cached, where inference is executed, and how much information is transmitted. Privacy-preserving fusion protects sensitive visual data, but stronger protection may affect latency, model updates, and retrieval accuracy. 6G-enabled collaboration can support dense camera networks, yet it also requires joint sensing, communication, and computation control. Green AI further constrains the roadmap by requiring continuous sensing to be evaluated through energy cost as well as accuracy.
Security and privacy protection are therefore treated here as emerging research challenges rather than mature deployment assumptions. Future privacy-preserving video fusion should protect not only raw frames during transmission, but also intermediate embeddings, retrieval indices, model updates, and cross-modal semantic representations. Recent studies provide early examples of this direction, including privacy-aware multi-camera surveillance based on federated learning, heterogeneous model fusion, and synthetic domain adaptation [91], privacy-preserving multimodal image-text retrieval with federated training and embedding obfuscation [136], lattice-based encrypted multimodal retrieval for healthcare data [103], and federated learning for privacy-enhancing video activity recognition [110]. For edge-assisted video fusion, this direction also requires joint privacy–utility–efficiency evaluation because stronger privacy protection may affect latency, energy consumption, and fusion accuracy [124].
Overall, the next generation of video fusion systems will likely evolve from isolated algorithmic modules toward integrated perception infrastructures. In such systems, foundation models provide semantic generalization, edge intelligence provides real-time responsiveness, privacy-preserving mechanisms protect sensitive data, 6G networks support large-scale collaborative perception, and green AI reduces the cost of continuous visual sensing. A key open problem is how to jointly optimize these goals rather than treating accuracy, latency, privacy, and energy consumption as separate objectives.

6. Conclusions

This review shows that multi-camera video fusion has shifted from isolated image enhancement or stitching operations toward a system-level problem involving geometry, temporal association, multimodal representation, edge deployment, privacy, and energy consumption. The central challenge is not only how to fuse more inputs, but how to decide which information should be aligned, cached, transmitted, inferred, and protected under specific application constraints.
Accordingly, the manuscript has organized the field along three connected dimensions: technical mechanisms, application bottlenecks, and deployment constraints. Perspective fusion depends on calibration and view overlap; spatio-temporal fusion depends on cross-camera association and long-range temporal modeling; multimodal and foundation-model-based fusion depend on representation alignment, modality reliability, and computational feasibility. Edge caching further changes the problem by moving part of storage, inference, and fusion from the cloud to edge nodes, where latency, bandwidth, and resource limits become part of the fusion design. These observations suggest that future progress should be evaluated through joint trade-offs among accuracy, latency, scalability, privacy, and energy cost rather than through fusion accuracy alone.

Author Contributions

Conceptualization, C.M. and L.X.; methodology, C.M. and L.X.; validation, C.M.; data curation, L.X.; writing—original draft preparation, C.M. and L.X.; writing—review and editing, C.M. and L.X.; supervision, C.M. and L.X.; project administration, C.M. and L.X. All authors have read and agreed to the published version of the manuscript.

Funding

This research was funded by the China Postdoctoral Science Foundation (grant number: 252102211014), and in part by the Science and Technology Research Project of Henan Province (grant number: 262102210226).

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

The dataset is available on request from the authors.

Acknowledgments

This research was financially supported by the China Postdoctoral Science Foundation (252102211014), and in part by the Science and Technology Research Project of Henan Province (262102210226).

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Ahmad, T.; Morel, A.; Cheng, N.; Palaniappan, K.; Calyam, P.; Sun, K.; Pan, J. Future UAV/drone systems for intelligent active surveillance and monitoring. ACM Comput. Surv. 2025, 58, 1–37. [Google Scholar] [CrossRef] [Scilit]
  2. Phan, D.T.; Doan, V.H.M.; Choi, J.; Lee, B.; Oh, J. AADC-Net: A multimodal deep learning framework for automatic anomaly detection in real-time surveillance. IEEE Trans. Instrum. Meas. 2025, 74, 1–13. [Google Scholar] [CrossRef] [Scilit]
  3. Laviola, E.; Gattullo, M.; Romano, S.; Uva, A.E. Which Side is the Top? A User Study to Compare Visual Assets for Component Orientation in Assembly with Augmented Reality. IEEE Trans. Vis. Comput. Graph. 2025, 31, 3470–3480. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  4. Zhang, T.; Cui, Y.; Fang, W. Integrative human and object aware online progress observation for human-centric augmented reality assembly. Adv. Eng. Inform. 2025, 64, 103081. [Google Scholar] [CrossRef] [Scilit]
  5. Yan, H.; Li, Y. Generative AI for Intelligent Transportation Systems: Road Transportation Perspective. ACM Comput. Surv. 2025, 57, 1–45. [Google Scholar] [CrossRef] [Scilit]
  6. Liu, Y.; Wang, X.; Hu, E.; Wang, A.; Shiri, B.; Lin, W. VNDHR: Variational single nighttime image Dehazing for enhancing visibility in intelligent transportation systems via hybrid regularization. IEEE Trans. Intell. Transp. Syst. 2025, 26, 10189–10203. [Google Scholar] [CrossRef] [Scilit]
  7. Lai, Q.; Ji, L. A bidirectional cross-scrambling medical image encryption scheme incorporates compressed sensing and its application in IoMT. IEEE Trans. Circuits Syst. Video Technol. 2025, 35, 7697–7705. [Google Scholar] [CrossRef] [Scilit]
  8. Xiao, X.; Zhang, J.; Shao, Y.; Liu, J.; Shi, K.; He, C.; Kong, D. Deep learning-based medical ultrasound image and video segmentation methods: Overview, frontiers, and challenges. Sensors 2025, 25, 2361. [Google Scholar] [CrossRef] [Scilit]
  9. Wang, Q.; Tu, Z.; Li, C.; Tang, J. High performance RGB-Thermal Video Object Detection via hybrid fusion with progressive interaction and temporal-modal difference. Inf. Fusion 2025, 114, 102665. [Google Scholar] [CrossRef] [Scilit]
  10. Yin, Z.; Wu, Z.; Shi, W.; Hu, G.; Lin, W. Video compressed sensing via wavelet residual sampling and dual-domain fusion. IEEE Trans. Multimed. 2025, 27, 4240–4255. [Google Scholar] [CrossRef] [Scilit]
  11. Amosa, T.I.; Sebastian, P.; Izhar, L.I.; Ibrahim, O.; Ayinla, L.S.; Bahashwan, A.A.; Bala, A.; Samaila, Y.A. Multi-camera multi-object tracking: A review of current trends and future advances. Neurocomputing 2023, 552, 126558. [Google Scholar] [CrossRef] [Scilit]
  12. Iguernaissi, R.; Merad, D.; Aziz, K.; Drap, P. People tracking in multi-camera systems: A review. Multimed. Tools Appl. 2019, 78, 10773–10793. [Google Scholar] [CrossRef] [Scilit]
  13. Yao, S.; Guan, R.; Huang, X.; Li, Z.; Sha, X.; Yue, Y.; Lim, E.G.; Seo, H.; Man, K.L.; Zhu, X.; et al. Radar-camera fusion for object detection and semantic segmentation in autonomous driving: A comprehensive review. IEEE Trans. Intell. Veh. 2023, 9, 2094–2128. [Google Scholar] [CrossRef] [Scilit]
  14. Guan, B.; Zhao, J.; Mitra, S.; Kneip, L. Six-point method for multi-camera systems with reduced solution space. Int. J. Comput. Vis. 2025, 133, 7270–7292. [Google Scholar] [CrossRef] [Scilit]
  15. Deng, K.; Zhao, D.; Han, Q.; Wang, S.; Zhang, Z.; Zhou, A.; Ma, H. Geryon: Edge assisted real-time and robust object detection on drones via mmWave radar and camera fusion. In Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies; Association for Computing Machinery: New York, NY, USA, 2022; Volume 6, pp. 1–27. [Google Scholar]
  16. Benrazek, A.E.; Farou, B.; Seridi, H.; Kouahla, Z.; Kurulay, M. Ascending hierarchical classification for camera clustering based on FoV overlaps for WMSN. IET Wirel. Sens. Syst. 2019, 9, 382–388. [Google Scholar] [CrossRef] [Scilit]
  17. Yeong, D.J.; Velasco-Hernandez, G.; Barry, J.; Walsh, J. Sensor and sensor fusion technology in autonomous vehicles: A review. Sensors 2021, 21, 2140. [Google Scholar] [CrossRef] [Scilit]
  18. Gandhi, A.; Adhvaryu, K.; Poria, S.; Cambria, E.; Hussain, A. Multimodal sentiment analysis: A systematic review of history, datasets, multimodal fusion methods, applications, challenges and future directions. Inf. Fusion 2023, 91, 424–444. [Google Scholar] [CrossRef] [Scilit]
  19. Kashinath, S.A.; Mostafa, S.A.; Mustapha, A.; Mahdin, H.; Lim, D.; Mahmoud, M.A.; Mohammed, M.A.; Al-Rimy, B.A.S.; Fudzee, M.F.M.; Yang, T.J. Review of data fusion methods for real-time and multi-sensor traffic flow analysis. IEEE Access 2021, 9, 51258–51276. [Google Scholar] [CrossRef] [Scilit]
  20. Qi, J.; Yang, P.; Newcombe, L.; Peng, X.; Yang, Y.; Zhao, Z. An overview of data fusion techniques for Internet of Things enabled physical activity recognition and measure. Inf. Fusion 2020, 55, 269–280. [Google Scholar] [CrossRef] [Scilit]
  21. Majumder, S.; Kehtarnavaz, N. Vision and inertial sensing fusion for human action recognition: A review. IEEE Sens. J. 2020, 21, 2454–2467. [Google Scholar] [CrossRef] [Scilit]
  22. Zhang, Y.J. Camera calibration. In 3-D Computer Vision: Principles, Algorithms and Applications; Springer: Berlin/Heidelberg, Germany, 2023; pp. 37–65. [Google Scholar]
  23. Sha, L.; Hobbs, J.; Felsen, P.; Wei, X.; Lucey, P.; Ganguly, S. End-to-end camera calibration for broadcast videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 13–19 June 2020; pp. 13627–13636. [Google Scholar]
  24. Bhardwaj, R.; Tummala, G.K.; Ramalingam, G.; Ramjee, R.; Sinha, P. Autocalib: Automatic traffic camera calibration at scale. ACM Trans. Sens. Netw. (TOSN) 2018, 14, 1–27. [Google Scholar] [CrossRef] [Scilit]
  25. Zhang, N.; Izquierdo, E. A high accuracy camera calibration method for sport videos. In Proceedings of the 2021 International Conference on Visual Communications and Image Processing (VCIP); IEEE: Piscataway, NJ, USA, 2021; pp. 1–5. [Google Scholar]
  26. Zhi, X.; Yan, J.; Hang, Y.; Wang, S. Realization of CUDA-based real-time registration and target localization for high-resolution video images. J. Real.-Time Image Process. 2019, 16, 1025–1036. [Google Scholar] [CrossRef] [Scilit]
  27. Ye, Y.; Bruzzone, L.; Shan, J.; Bovolo, F.; Zhu, Q. Fast and robust matching for multimodal remote sensing image registration. IEEE Trans. Geosci. Remote Sens. 2019, 57, 9059–9070. [Google Scholar] [CrossRef] [Scilit]
  28. Gao, Z.; Hwang, A.; Zhai, G.; Peli, E. Correcting geometric distortions in stereoscopic 3D imaging. PLoS ONE 2018, 13, e0205032. [Google Scholar] [CrossRef] [Scilit]
  29. Dziembowski, A.; Mieloch, D.; Różek, S.; Domański, M. Color correction for immersive video applications. IEEE Access 2021, 9, 75626–75640. [Google Scholar] [CrossRef] [Scilit]
  30. Li, Z.; Jia, Z.; Yang, J.; Kasabov, N. Low illumination video image enhancement. IEEE Photonics J. 2020, 12, 1–13. [Google Scholar] [CrossRef] [Scilit]
  31. Cao, G.; Huang, L.; Tian, H.; Huang, X.; Wang, Y.; Zhi, R. Contrast enhancement of brightness-distorted images by improved adaptive gamma correction. Comput. Electr. Eng. 2018, 66, 569–582. [Google Scholar] [CrossRef] [Scilit]
  32. Abbadi, N.K.E.; Al Hassani, S.A.; Abdulkhaleq, A.H. A review over panoramic image stitching techniques. In Proceedings of the Journal of Physics: Conference Series; IOP Publishing: Bristol, UK, 2021; Volume 1999, p. 012115. [Google Scholar]
  33. Wang, Z.; Yang, Z. Review on image-stitching techniques. Multimed. Syst. 2020, 26, 413–430. [Google Scholar] [CrossRef] [Scilit]
  34. Nie, L.; Lin, C.; Liao, K.; Liu, S.; Zhao, Y. Unsupervised deep image stitching: Reconstructing stitched features to images. IEEE Trans. Image Process. 2021, 30, 6184–6197. [Google Scholar] [CrossRef] [Scilit]
  35. Voynov, O.; Safin, A.; Ignatiev, S.; Burnaev, E. How good MVSNets are at depth fusion. In Proceedings of the Thirteenth International Conference on Machine Vision; SPIE: Bellingham, WA, USA, 2021; Volume 11605, pp. 119–125. [Google Scholar]
  36. Cai, Y.; Li, L.; Wang, D.; Liu, X. MFNet: Multi-level fusion aware feature pyramid based multi-view stereo network for 3D reconstruction. Appl. Intell. 2023, 53, 4289–4301. [Google Scholar] [CrossRef] [Scilit]
  37. Lou, J.; Liu, W.; Chen, Z.; Liu, F.; Cheng, J. Elfnet: Evidential local-global fusion for stereo matching. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Paris, France, 1–6 October 2023; pp. 17784–17793. [Google Scholar]
  38. Alldieck, T.; Magnor, M.; Xu, W.; Theobalt, C.; Pons-Moll, G. Video based reconstruction of 3d people models. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 18–23 June 2018; pp. 8387–8397. [Google Scholar]
  39. Sun, J.; Xie, Y.; Chen, L.; Zhou, X.; Bao, H. Neuralrecon: Real-time coherent 3d reconstruction from monocular video. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 20–25 June 2021; pp. 15598–15607. [Google Scholar]
  40. González Izard, S.; Sánchez Torres, R.; Alonso Plaza, O.; Juanes Mendez, J.A.; García-Peñalvo, F.J. Nextmed: Automatic imaging segmentation, 3D reconstruction, and 3D model visualization platform using augmented and virtual reality. Sensors 2020, 20, 2962. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  41. Wang, J.; Xuan, S.; Zhang, H.; Qin, X. The moving target tracking and segmentation method based on space-time fusion. Multimed. Tools Appl. 2023, 82, 12245–12262. [Google Scholar] [CrossRef] [Scilit]
  42. He, S.; Shao, H.; Xian, W.; Zhang, S.; Zhong, J.; Qi, J. Extraction of abandoned land in hilly areas based on the spatio-temporal fusion of multi-source remote sensing images. Remote Sens. 2021, 13, 3956. [Google Scholar] [CrossRef] [Scilit]
  43. Liang, W.; Zhang, J.; Zhan, Y. Weakly supervised video anomaly detection based on spatial-temporal feature fusion enhancement. Signal Image Video Process. 2024, 18, 1111–1118. [Google Scholar] [CrossRef] [Scilit]
  44. Xu, H.; Zhong, S.; Zhang, T.; Zou, X. Real-time Infrared Small Target Detection with NonLocal Spatial-Temporal Feature Fusion. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2024, 17, 7888–7902. [Google Scholar] [CrossRef] [Scilit]
  45. Fu, C.; Yuan, H.; Xu, H.; Zhang, H.; Shen, L. Cuboid-Net: A multi-branch convolutional neural network for joint space-time video super resolution. IET Image Process. 2023, 17, 4089–4101. [Google Scholar] [CrossRef] [Scilit]
  46. Jiang, T.; Guo, M.; Yang, L.; Ma, Z.; Liu, H. Traffic Flow Prediction Based on Decomposed Spatio-Temporal Fusion Graph Convolutional Network. In Intelligent Transportation and Smart Cities; IOS Press: Amsterdam, The Netherlands, 2024; pp. 156–167. [Google Scholar]
  47. Liu, C.; Zhang, Y.; Luo, H.; Tang, J.; Chen, W.; Xu, X.; Wang, F.; Li, H.; Shen, Y. City-Scale Multi-Camera Vehicle Tracking Guided by Crossroad Zones. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, Nashville, TN, USA, 20–25 June 2021; pp. 4129–4137. [Google Scholar]
  48. Yang, X.; Ye, J.; Lu, J.; Gong, C.; Jiang, M.; Lin, X.; Zhang, W.; Tan, X.; Li, Y.; Ye, X.; et al. Box-Grained Reranking Matching for Multi-Camera Multi-Target Tracking. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, New Orleans, LA, USA, 19–20 June 2022; pp. 3095–3105. [Google Scholar] [CrossRef] [Scilit]
  49. Lu, J.; Yang, X.; Ye, J.; Zhang, Y.; Zou, Z.; Zhang, W.; Tan, X. CityTrack: Improving City-Scale Multi-Camera Multi-Target Tracking by Location-Aware Tracking and Box-Grained Matching. arXiv 2023, arXiv:2307.02753. [Google Scholar] [CrossRef] [Scilit]
  50. Lin, Y.; Lockyer, S.; Evans, A.; Zarbock, M.; Zhang, N. City-Scale Multi-Camera Vehicle Tracking System with Improved Self-Supervised Camera Link Model. In Pattern Analysis and Machine Intelligence; Communications in Computer and Information Science; Springer Nature: Singapore, 2025; Volume 2323, pp. 67–76. [Google Scholar] [CrossRef] [Scilit]
  51. Huang, F.; Yao, J.; Su, C.; Xu, S. An Online Multi-Camera Multi-Vehicle Tracking Using Lightweight YOLO11 and Improved Association Strategies. J. King Saud. Univ. Comput. Inf. Sci. 2025, 37, 180. [Google Scholar] [CrossRef] [Scilit]
  52. Tseng, Y.S.; Su, Y.F.; Lin, D.T. Enhancing Multi-Target Multi-Camera Vehicle Tracking with YOLOv9 and Attention Mechanisms for Smart City Traffic Monitoring. Multimed. Tools Appl. 2025, 84, 45095–45117. [Google Scholar] [CrossRef] [Scilit]
  53. Munikoti, S.; Stewart, I.; Horawalavithana, S.; Kvinge, H.; Emerson, T.; Thompson, S.; Pazdernik, K. Generalist Multimodal AI: A Review of Architectures, Challenges and Opportunities. Neurocomputing 2026, 676, 132933. [Google Scholar] [CrossRef] [Scilit]
  54. Lin, B.; Ye, Y.; Zhu, B.; Cui, J.; Ning, M.; Jin, P.; Yuan, L. Video-LLaVA: Learning United Visual Representation by Alignment Before Projection. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing; Al-Onaizan, Y., Bansal, M., Chen, Y.N., Eds.; Association for Computational Linguistics: Stroudsburg, PA, USA, 2024; pp. 5971–5984. [Google Scholar] [CrossRef] [Scilit]
  55. Lu, J.; Wang, H.; Xu, Y.; Wang, Y.; Yang, K.; Fu, Y. Representation Potentials of Foundation Models for Multimodal Alignment: A Survey. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, Suzhou, China; Association for Computational Linguistics: Stroudsburg, PA, USA, 2025; pp. 16669–16684. [Google Scholar] [CrossRef] [Scilit]
  56. Shu, Y.; Liu, Z.; Zhang, P.; Qin, M.; Zhou, J.; Liang, Z.; Huang, T.; Zhao, B. Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 10–17 June 2025; pp. 26160–26169. [Google Scholar]
  57. Zohar, O.; Wang, X.; Dubois, Y.; Mehta, N.; Xiao, T.; Hansen-Estruch, P.; Yu, L.; Wang, X.; Juefei-Xu, F.; Zhang, N.; et al. Apollo: An Exploration of Video Understanding in Large Multimodal Models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 10–17 June 2025; pp. 18891–18901. [Google Scholar]
  58. Fu, C.; Dai, Y.; Luo, Y.; Li, L.; Ren, S.; Zhang, R.; Wang, Z.; Zhou, C.; Shen, Y.; Zhang, M.; et al. Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-Modal LLMs in Video Analysis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 10–17 June 2025; pp. 24108–24118. [Google Scholar]
  59. Chen, B.; Li, D.; Li, J.; Wu, H. LongVideoBench: A Benchmark for Long-Context Interleaved Video-Language Understanding. In Proceedings of the Advances in Neural Information Processing Systems, Vancouver, BC, Canada, 10–15 December 2024; Volume 37, pp. 28828–28857. [Google Scholar] [CrossRef] [Scilit]
  60. Ding, S.; Dong, X.; Lin, D.; Qian, R.; Wang, J.; Zang, Y.; Zhang, P. Streaming Long Video Understanding with Large Language Models. In Proceedings of the Advances in Neural Information Processing Systems 37, Vancouver, BC, Canada, 10–15 December 2024; Volume 37, pp. 119336–119360. [Google Scholar] [CrossRef] [Scilit]
  61. Weng, Y.; Han, M.; He, H.; Chang, X.; Zhuang, B. LongVLM: Efficient Long Video Understanding via Large Language Models. In Proceedings of the Computer Vision–ECCV 2024; Leonardis, A., Ricci, E., Roth, S., Russakovsky, O., Sattler, T., Varol, G., Eds.; Springer Nature: Cham, Switzerland, 2025; pp. 453–470. [Google Scholar] [CrossRef] [Scilit]
  62. Wu, J.; Liu, W.; Liu, Y.; Liu, M.; Nie, L.; Lin, Z.; Chen, C.W. A Survey on Video Temporal Grounding With Multimodal Large Language Model. IEEE Trans. Pattern Anal. Mach. Intell. 2026, 48, 1521–1541. [Google Scholar] [CrossRef] [Scilit]
  63. Ortega, J.D.; Senoussaoui, M.; Granger, E.; Pedersoli, M.; Cardinal, P.; Koerich, A.L. Multimodal fusion with deep neural networks for audio-video emotion recognition. arXiv 2019, arXiv:1907.03196. [Google Scholar] [CrossRef] [Scilit]
  64. Yang, X.; Molchanov, P.; Kautz, J. Multilayer and multimodal fusion of deep neural networks for video classification. In Proceedings of the 24th ACM international conference on Multimedia; Association for Computing Machinery: New York, NY, USA, 2016; pp. 978–987. [Google Scholar]
  65. Su, L.; Hu, C.; Li, G.; Cao, D. Msaf: Multimodal split attention fusion. arXiv 2020, arXiv:2012.07175. [Google Scholar]
  66. Shvetsova, N.; Chen, B.; Rouditchenko, A.; Thomas, S.; Kingsbury, B.; Feris, R.S.; Harwath, D.; Glass, J.; Kuehne, H. Everything at Once-Multi-Modal Fusion Transformer for Video Retrieval. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 19–20 June 2022; pp. 20020–20029. [Google Scholar]
  67. Li, X.; Fan, B.; Tian, J.; Fan, H. GAFusion: Adaptive Fusing LiDAR and Camera with Multiple Guidance for 3D Object Detection. In Proceedings of the 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 16–22 June 2024; pp. 21209–21218. [Google Scholar] [CrossRef] [Scilit]
  68. Park, K.; Kim, Y.; Kim, D.; Choi, J.W. Resilient Sensor Fusion Under Adverse Sensor Failures via Multi-Modal Expert Fusion. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 10–17 June 2025; pp. 6720–6729. [Google Scholar]
  69. Zhou, T.; Chen, J.; Shi, Y.; Jiang, K.; Yang, M.; Yang, D. Bridging the view disparity between radar and camera features for multi-modal fusion 3d object detection. IEEE Trans. Intell. Veh. 2023, 8, 1523–1535. [Google Scholar] [CrossRef] [Scilit]
  70. Chen, R.; Ning, J.; Lei, Y.; Hui, Y.; Cheng, N. Mixed traffic flow state detection: A connected vehicles-assisted roadside radar and video data fusion scheme. IEEE Open J. Intell. Transp. Syst. 2023, 4, 360–371. [Google Scholar] [CrossRef] [Scilit]
  71. Xu, S.; Li, F.; Huang, P.; Song, Z.; Yang, Z.X. TiGDistill-BEV: Multi-View BEV 3D Object Detection via Target Inner-Geometry Learning Distillation. IEEE Trans. Circuits Syst. Video Technol. 2026, 36, 846–860. [Google Scholar] [CrossRef] [Scilit]
  72. Chen, Y.; Yu, Z.; Chen, Y.; Lan, S.; Anandkumar, A.; Jia, J.; Alvarez, J.M. FocalFormer3D: Focusing on Hard Instance for 3D Object Detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Paris, France, 1–6 October 2023; pp. 8394–8405. [Google Scholar]
  73. Bang, G.; Choi, K.; Kim, J.; Kum, D.; Choi, J.W. RadarDistill: Boosting Radar-Based Object Detection Performance via Knowledge Distillation from LiDAR Features. In Proceedings of the 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 16–22 June 2024; pp. 15491–15500. [Google Scholar] [CrossRef] [Scilit]
  74. Wang, Z.; Huang, Z.; Gao, Y.; Wang, N.; Liu, S. MV2DFusion: Leveraging Modality-Specific Object Semantics for Multi-Modal 3D Detection. IEEE Trans. Pattern Anal. Mach. Intell. 2026, 48, 609–623. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  75. Li, S.; Ma, L.; Li, X. ModalPatch: A Plug-and-Play Module for Robust Multi-Modal 3D Object Detection under Modality Drop. In Proceedings of the 2026 IEEE International Conference on Robotics and Automation (ICRA), Vienna, Austria, 1–5 June 2026. [Google Scholar]
  76. Zhang, J.; Zhang, Y.; Qi, Y.; Fu, Z.; Liu, Q.; Wang, Y. GeoBEV: Learning Geometric BEV Representation for Multi-View 3D Object Detection. Proc. Aaai Conf. Artif. Intell. 2025, 39, 9960–9968. [Google Scholar] [CrossRef] [Scilit]
  77. Yin, J.; Shen, J.; Chen, R.; Li, W.; Yang, R.; Frossard, P.; Wang, W. IS-Fusion: Instance-Scene Collaborative Fusion for Multimodal 3D Object Detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 16–22 June 2024; pp. 14905–14915. [Google Scholar]
  78. Chen, X.; Zhang, T.; Wang, Y.; Wang, Y.; Zhao, H. Futr3d: A Unified Sensor Fusion Framework for 3d Detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 17–24 June 2023; pp. 172–181. [Google Scholar]
  79. Li, Y.; Yang, Y.; Lei, Z. RCTrans: Radar-Camera Transformer via Radar Densifier and Sequential Decoder for 3D Object Detection. Proc. Aaai Conf. Artif. Intell. 2025, 39, 5048–5056. [Google Scholar] [CrossRef] [Scilit]
  80. Li, M.; Zhang, H.; Xu, C.; Yan, C.; Liu, H.; Li, X. MFVC: Urban Traffic Scene Video Caption Based on Multimodal Fusion. Electronics 2022, 11, 2999. [Google Scholar] [CrossRef] [Scilit]
  81. Joze, H.R.V.; Shaban, A.; Iuzzolino, M.L.; Koishida, K. MMTM: Multimodal transfer module for CNN fusion. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 13–19 June 2020; pp. 13289–13299. [Google Scholar]
  82. Vielzeuf, V.; Lechervy, A.; Pateux, S.; Jurie, F. Centralnet: A multilayer approach for multimodal fusion. In Proceedings of the European Conference on Computer Vision (ECCV) Workshops, Munich, Germany, 8–14 September 2018. [Google Scholar]
  83. Xue, Z.; Marculescu, R. Dynamic multimodal fusion. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 17–24 June 2023; pp. 2575–2584. [Google Scholar]
  84. Wu, S.; Dai, D.; Qin, Z.; Liu, T.; Lin, B.; Cao, Y.; Sui, Z. Denoising Bottleneck with Mutual Information Maximization for Video Multimodal Fusion. arXiv 2023, arXiv:2305.14652. [Google Scholar] [CrossRef] [Scilit]
  85. Liu, X.; Li, S.; Wang, M. Hierarchical Attention-Based Multimodal Fusion Network for Video Emotion Recognition. Comput. Intell. Neurosci. 2021, 2021, 5585041. [Google Scholar] [CrossRef] [Scilit]
  86. Zhu, H.; Wang, Z.; Shi, Y.; Hua, Y.; Xu, G.; Deng, L. Multimodal Fusion Method Based on Self-Attention Mechanism. Wirel. Commun. Mob. Comput. 2020, 2020, 8843186. [Google Scholar] [CrossRef] [Scilit]
  87. Liu, J.; Yuan, Z.; Wang, C. Towards good practices for multi-modal fusion in large-scale video classification. In Proceedings of the European Conference on Computer Vision (ECCV) Workshops, Munich, Germany, 8–14 September 2018. [Google Scholar]
  88. Wei, Y.; Wang, X.; Guan, W.; Nie, L.; Lin, Z.; Chen, B. Neural multimodal cooperative learning toward micro-video understanding. IEEE Trans. Image Process. 2019, 29, 1–14. [Google Scholar] [CrossRef] [Scilit]
  89. Elharrouss, O.; Almaadeed, N.; Al-Maadeed, S. A review of video surveillance systems. J. Vis. Commun. Image Represent. 2021, 77, 103116. [Google Scholar] [CrossRef] [Scilit]
  90. Socha, R.; Kogut, B. Urban video surveillance as a tool to improve security in public spaces. Sustainability 2020, 12, 6210. [Google Scholar] [CrossRef] [Scilit]
  91. Lu, P.J.; Chen, W.Y.; Huang, Y.T.; Tseng, V.S.M. Heterogeneous Model Fusion for Privacy-Aware Multi-Camera Surveillance via Synthetic Domain Adaptation. Inf. Fusion 2026, 135, 104413. [Google Scholar] [CrossRef] [Scilit]
  92. Anwar, A.; Choe, T.E.; Wang, Z.; Fidler, S.; Park, M. Augmented Reality based Simulated Data (ARSim) with multi-view consistency for AV perception networks. arXiv 2024, arXiv:2403.15370. [Google Scholar] [CrossRef] [Scilit]
  93. Huang, B.; Timmons, N.G.; Li, Q. Augmented reality with multi-view merging for robot teleoperation. In Proceedings of the Companion of the 2020 ACM/IEEE International Conference on Human-Robot Interaction; Association for Computing Machinery: New York, NY, USA, 2020; pp. 260–262. [Google Scholar]
  94. Bernardi, G.; Brisebarre, G.; Roman, S.; Ardabilian, M.; Dellandrea, E. A comprehensive survey on image fusion: Which approach fits which need. Inf. Fusion 2025, 126, 103594. [Google Scholar] [CrossRef] [Scilit]
  95. Li, H.; Lo, C.H.; Smith, A.; Yu, Z. The development of virtual production in film industry in the past decade. In Proceedings of the International Conference on Human-Computer Interaction; Springer: Cham, Switzerland, 2022; pp. 221–239. [Google Scholar]
  96. Zhang, Y.; Pang, Z.; Huang, S.; Wang, C.; Zhou, X. Unmasking AI-created visual content: A review of generated images and deepfake detection technologies. J. King Saud. Univ. Comput. Inf. Sci. 2025, 37, 148. [Google Scholar] [CrossRef] [Scilit]
  97. Deng, K.; Xing, L.; Wu, H.; Ma, H.; Ling, Y.; Gao, J. Advances in object detection for autonomous driving using mmwave radar and camera: A comprehensive survey. J. King Saud. Univ. Comput. Inf. Sci. 2025, 37, 328. [Google Scholar] [CrossRef] [Scilit]
  98. Xing, L.; Ye, J.; Deng, K.; Wu, H.; Ma, H.; Gao, J. Cerberus: Accurate Real-Time Object Detection System Under Adverse Weather Conditions via Multimodal Fusion. IEEE Internet Things J. 2025, 12, 52837–52849. [Google Scholar] [CrossRef] [Scilit]
  99. Deng, K.; Xing, L.; Wu, H.; Wang, Y.; Xu, L.; Ling, Y.; Gao, J. mmReg: Centimeter-Level and Real-Time mmWave Radar Point Cloud Registration for Multivehicle Sensing. IEEE Internet Things J. 2025, 13, 5069–5086. [Google Scholar] [CrossRef] [Scilit]
  100. Zhang, L.; Gao, G.; Yin, H.; Zhang, H. Multi-Edge Reinforced Collaborative Data Acquisition for Continuous Video Analytics by Prioritizing Quality over Quantity. Proc. Aaai Conf. Artif. Intell. 2025, 39, 1084–1092. [Google Scholar] [CrossRef] [Scilit]
  101. Yadav, S.P.; Yadav, S. Image fusion using hybrid methods in multimodality medical images. Med. Biol. Eng. Comput. 2020, 58, 669–687. [Google Scholar] [CrossRef] [Scilit]
  102. Bavirisetti, D.P.; Xiao, G.; Zhao, J.; Dhuli, R.; Liu, G. Multi-scale guided image and video fusion: A fast and efficient approach. Circuits Syst. Signal Process. 2019, 38, 5576–5605. [Google Scholar] [CrossRef] [Scilit]
  103. Hou, Y.; Yao, W.; Zhu, X.; Li, Z. Lattice-Based Privacy-Preserving Multimodal Retrieval for Healthcare. J. Biomed. Inform. 2026, 175, 104990. [Google Scholar] [CrossRef] [Scilit]
  104. Flangas, A.T.; Sattarvand, J.; Dascalu, S.M.; Harris, F.C. Merging live video feeds for remote monitoring of a mining machine. In Proceedings of the 2021 European Symposium on Software Engineering; Association for Computing Machinery: New York, NY, USA, 2021; pp. 6–13. [Google Scholar]
  105. Arshad, M.A.; Khan, S.H.; Qamar, S.; Khan, M.W.; Murtza, I.; Gwak, J.; Khan, A. Drone navigation using region and edge exploitation-based deep CNN. IEEE Access 2022, 10, 95441–95450. [Google Scholar] [CrossRef] [Scilit]
  106. Park, J.; Kim, P.; Ko, D. Real-Time Open-Vocabulary Perception for Mobile Robots on Edge Devices: A Systematic Analysis of the Accuracy-Latency Trade-Off. Front. Robot. AI 2025, 12, 1693988. [Google Scholar] [CrossRef] [Scilit]
  107. Huang, R.; Yu, M.; Tsoi, M.; Ouyang, X. MMEdge: Accelerating On-device Multimodal Inference via Pipelined Sensing and Encoding. In Proceedings of the 24th ACM Conference on Embedded Networked Sensor Systems; Association for Computing Machinery: New York, NY, USA, 2026. [Google Scholar] [CrossRef] [Scilit]
  108. Wu, J.; Liu, W.; Liu, Y.; Liu, M.; Nie, L.; Lin, Z.; Chen, C.W. Jupiter: Fast and Resource-Efficient Collaborative Inference of Generative LLMs on Edge Devices. In Proceedings of the IEEE INFOCOM 2025-IEEE Conference on Computer Communications; IEEE: Piscataway, NJ, USA, 2025; pp. 1–10. [Google Scholar] [CrossRef] [Scilit]
  109. Sharma, R.; Potnis, A.; Chaurasia, V. Enhancing Smart Home Security Using Deep Convolutional Neural Networks and Multiple Cameras. Wirel. Pers. Commun. 2024, 136, 2185–2200. [Google Scholar] [CrossRef] [Scilit]
  110. Tonmoy, M.R.; Hossain, M.M.; Safran, M.; Alfarhood, S.; Che, D.; Mridha, M.F. Leveraging Federated Learning for Efficient Privacy-Enhancing Violent Activity Recognition from Videos. Comput. Mater. Contin. 2025, 85, 5747–5763. [Google Scholar] [CrossRef] [Scilit]
  111. Cheriet, A.; Mekhaznia, T. Optimizing network performance and security in video surveillance data storage: A solution with blockchain and mobile edge computing. Computing 2025, 107, 130. [Google Scholar] [CrossRef] [Scilit]
  112. Mallick, A.; Barnwal, R.P. A Scalable Framework for Multi-cloud IoT Data Synchronization. In Proceedings of the 26th International Conference on Distributed Computing and Networking; Association for Computing Machinery: New York, NY, USA, 2025; pp. 364–369. [Google Scholar]
  113. Cardoso Nunes, D.; Loureiro Coelho, B.; Parizotto, R.; Egon Schaeffer-Filho, A. No Worker Left (Too Far) Behind: Dynamic Hybrid Synchronization for In-Network ML Aggregation. Int. J. Netw. Manag. 2025, 35, e2290. [Google Scholar] [CrossRef] [Scilit]
  114. Chen, J.; Shu, Q.; Lu, Y.; Zhang, Y.; Wang, Y. QCTF: A Quantized Communication and Transferable Fusion Framework for Multi-Agent Collaborative Perception. IEEE Trans. Intell. Transp. Syst. 2025, 26, 15013–15027. [Google Scholar] [CrossRef] [Scilit]
  115. Zhang, M.; Cao, J.; Sahni, Y.; Chen, Q.; Jiang, S.; Yang, L. Blockchain-based collaborative edge intelligence for trustworthy and real-time video surveillance. IEEE Trans. Ind. Inform. 2022, 19, 1623–1633. [Google Scholar] [CrossRef] [Scilit]
  116. Anand, J.; Karthikeyan, B. Dynamic priority-based task scheduling and adaptive resource allocation algorithms for efficient edge computing in healthcare systems. Results Eng. 2025, 25, 104342. [Google Scholar] [CrossRef] [Scilit]
  117. Chen, Y.; Ding, Y.; Hu, Z.Z.; Ren, Z. Geometrized task scheduling and adaptive resource allocation for large-scale edge computing in smart cities. IEEE Internet Things J. 2025, 12, 14398–14419. [Google Scholar] [CrossRef] [Scilit]
  118. Zhou, Z.; Xie, J.; Huang, M.; Ouyang, T.; Liu, F.; Chen, X. Towards Federated Inference: An Online Model Ensemble Framework for Cooperative Edge AI. In Proceedings of the IEEE INFOCOM 2025-IEEE Conference on Computer Communications; IEEE: Piscataway, NJ, USA, 2025; pp. 1–10. [Google Scholar] [CrossRef] [Scilit]
  119. Le, H.; Lu, C.K.; Hsu, C.C. Video Anomaly Detection for Edge-Based IoT Systems: A Survey of Input Modalities and Real-Time Applications. Intell. Syst. Appl. 2026, 29, 200635. [Google Scholar] [CrossRef] [Scilit]
  120. Liu, Z.; Chen, X.; Wu, H.; Wang, Z.; Chen, X.; Niyato, D.; Huang, K. Integrated Sensing and Edge AI: Realizing Intelligent Perception in 6G. arXiv 2025, arXiv:2501.06726. [Google Scholar] [CrossRef] [Scilit]
  121. Cui, Y.; Cao, X.; Zhu, G.; Nie, J.; Xu, J. Edge Perception: Intelligent Wireless Sensing at Network Edge. IEEE Commun. Mag. 2025, 63, 166–173. [Google Scholar] [CrossRef] [Scilit]
  122. Niu, S.; Liu, Y.; Liao, K.; Zhang, B.; Zou, G. Dynamic edge-caching through content popularity and crowd prediction for short video services. Sci. Rep. 2025, 15, 42147. [Google Scholar] [CrossRef] [Scilit]
  123. Zhang, H.; Du, D.; Zheng, K.; Cao, Y.; Zhao, L.; Zhao, Y.; Zhang, J. GreenRP: Task-Aware Discharge-Resilient Routing for Sustainable Edge AI in Satellite Optical Networks. Electronics 2025, 14, 3075. [Google Scholar] [CrossRef] [Scilit]
  124. Chen, H.; Peng, Y.; Wei, C.; Meng, T.; Ai, W.; He, Z.; Li, K. DRL-Based Privacy-Preserving Video Streaming Task Offloading Under Energy Constraints in MEC. Comput. Netw. 2026, 280, 112164. [Google Scholar] [CrossRef] [Scilit]
  125. Mohanta, B.K.; Awad, A.I.; Dehury, M.K.; Mohapatra, H.; Khan, M.K. Protecting IoT-enabled healthcare data at the edge: Integrating blockchain, AES, and off-chain decentralized storage. IEEE Internet Things J. 2025, 12, 15333–15347. [Google Scholar] [CrossRef] [Scilit]
  126. Wang, X.; Tang, Z.; Guo, J.; Meng, T.; Wang, C.; Wang, T.; Jia, W. Empowering edge intelligence: A comprehensive survey on on-device ai models. ACM Comput. Surv. 2025, 57, 1–39. [Google Scholar] [CrossRef] [Scilit]
  127. Mercat, A.; Sainio, J.; Le Moan, S.; Herglotz, C. Do We Need 10 bits? Assessing HEVC Encoders for Energy-Efficient HDR Video Streaming. IEEE J. Emerg. Sel. Top. Circuits Syst. 2025, 15, 31–43. [Google Scholar] [CrossRef] [Scilit]
  128. Gomes, J.S.; Grellert, M.; Ramos, F.L.; Bampi, S. End-to-end Neural Video Compression: A Review. IEEE Open J. Circuits Syst. 2025, 6, 120–134. [Google Scholar] [CrossRef] [Scilit]
  129. Maglaras, L.; Kioskli, K. End-to-End Encryption: Technological and Human Factor Perspectives. In Proceedings of the International Conference on Human-Computer Interaction; Springer: Cham, Switzerland, 2025; pp. 129–139. [Google Scholar]
  130. Ao, W.; Boddeti, V.N. CryptoFace: End-to-End Encrypted Face Recognition. In Proceedings of the Computer Vision and Pattern Recognition Conference, Nashville, TN, USA, 11–15 June 2025; pp. 19197–19206. [Google Scholar]
  131. Tawfik, A.M.; Al-Ahwal, A.; Eldien, A.S.T.; Zayed, H.H. Blockchain-based access control and privacy preservation in healthcare: A comprehensive survey. Clust. Comput. 2025, 28, 529. [Google Scholar] [CrossRef] [Scilit]
  132. Dhinakaran, D.; Kumar, N.J.; Ponnuviji, N. Safeguarding confidentiality and privacy in cloud-enabled healthcare systems with spectrasafe encryption and dynamic k-anonymity algorithm. Expert Syst. Appl. 2025, 279, 127584. [Google Scholar] [CrossRef] [Scilit]
  133. Zhang, H.; Zhang, J.; Feng, W.; Bian, K.; Tuo, H. Edge-FVV: Free Viewpoint Video Streaming by Learning at the Edge. In Proceedings of the 2023 IEEE International Conference on Multimedia and Expo (ICME); IEEE: Piscataway, NJ, USA, 2023; pp. 2009–2014. [Google Scholar]
  134. Dong, J.; Ota, K.; Dong, M. Video frame interpolation: A comprehensive survey. ACM Trans. Multimed. Comput. Commun. Appl. 2023, 19, 1–31. [Google Scholar] [CrossRef] [Scilit]
  135. Parihar, A.S.; Varshney, D.; Pandya, K.; Aggarwal, A. A comprehensive survey on video frame interpolation techniques. Vis. Comput. 2022, 38, 295–319. [Google Scholar] [CrossRef] [Scilit]
  136. Gao, Y.; Luo, W.; Wang, C.; Ahmad, N.S.; Wang, X.; Goh, P. A Privacy-Preserving Multi-User Retrieval System for Multimodal Artificial Intelligence. Sci. Rep. 2026, 16, 10348. [Google Scholar] [CrossRef] [Scilit] [PubMed]
Figure 1. Camera setup in multi-camera systems. The layout differences indicate different requirements for calibration, depth estimation, and wide-area tracking.
Figure 1. Camera setup in multi-camera systems. The layout differences indicate different requirements for calibration, depth estimation, and wide-area tracking.
Digital 06 00047 g001
Figure 2. Automatic traffic camera calibration. The figure links vehicle geometry, pose estimation, and camera parameters to metric measurement and cross-view alignment.
Figure 2. Automatic traffic camera calibration. The figure links vehicle geometry, pose estimation, and camera parameters to metric measurement and cross-view alignment.
Digital 06 00047 g002
Figure 3. The creation of 3D images and videos with Blender5.0. The example is used to connect visual 3D asset creation with reconstruction accuracy, runtime, and model completeness.
Figure 3. The creation of 3D images and videos with Blender5.0. The example is used to connect visual 3D asset creation with reconstruction accuracy, runtime, and model completeness.
Digital 06 00047 g003
Figure 4. DNN-based fusion pipeline structures for early, intermediate, late, and hybrid fusion. The comparison highlights where modality interaction occurs in each architecture.
Figure 4. DNN-based fusion pipeline structures for early, intermediate, late, and hybrid fusion. The comparison highlights where modality interaction occurs in each architecture.
Digital 06 00047 g004
Figure 5. Holoscopic mixed traffic flow perception with radar and camera. The framework connects roadside sensing, connected-vehicle calibration, and traffic-state estimation.
Figure 5. Holoscopic mixed traffic flow perception with radar and camera. The framework connects roadside sensing, connected-vehicle calibration, and traffic-state estimation.
Digital 06 00047 g005
Figure 6. Feature fusion of camera and LiDAR. The figure illustrates heterogeneous feature alignment for object detection, localization, and scene understanding.
Figure 6. Feature fusion of camera and LiDAR. The figure illustrates heterogeneous feature alignment for object detection, localization, and scene understanding.
Digital 06 00047 g006
Figure 7. Multi-view ARSim data: Obstacle Consistency. The example illustrates geometric and appearance consistency between synthetic objects and real camera views.
Figure 7. Multi-view ARSim data: Obstacle Consistency. The example illustrates geometric and appearance consistency between synthetic objects and real camera views.
Digital 06 00047 g007
Figure 8. Carlet XR virtual shooting solution topology diagram. The topology links camera tracking, rendering, compositing, and synchronization in real-time production.
Figure 8. Carlet XR virtual shooting solution topology diagram. The topology links camera tracking, rendering, compositing, and synchronization in real-time production.
Digital 06 00047 g008
Figure 9. Architecture of the Edge-FVV system. The figure shows how edge caches, reference-view retrieval, request allocation, and virtual-view synthesis jointly affect latency.
Figure 9. Architecture of the Edge-FVV system. The figure shows how edge caches, reference-view retrieval, request allocation, and virtual-view synthesis jointly affect latency.
Digital 06 00047 g009
Table 1. Comparison of features in the review.
Table 1. Comparison of features in the review.
Review StudyDefinitionClassificationApplicationMulti-Camera SceneEdge Cache
Yao et al. [13]
Yeong et al. [17]
Gandhi et al. [18]
Kashinath et al. [19]
Qi et al. [20]
Majumder et al. [21]
Ours
Table 2. Comparison of pre-processing and correction methods for video fusion.
Table 2. Comparison of pre-processing and correction methods for video fusion.
CategoryTechniqueStrengthLimitationScenarioRef.
Camera calibrationIntrinsic; extrinsic; homography; PnPGeometry; 3D consistencyDynamic drift; setup costTraffic; sports; 3D[22,23,24]
Image registrationORB; SIFT; RANSAC; multimodal matchingCross-view alignmentWeak texture; occlusion; large baselinePanorama; remote sensing; surveillance[26,27]
Distortion correctionLens model; reverse mapping; stereo correctionShape recoveryNonlinear distortion; detail lossWide-angle; stereo; immersive video[28]
Color; brightness correctionColor calibration; HSV; tone mapping; gammaVisual consistencyFlicker; over-enhancement; color shiftMulti-view video; low light; XR[29,30,31]
Table 3. Comparison of multi-camera vehicle tracking methods on the CityFlowV2 benchmark.
Table 3. Comparison of multi-camera vehicle tracking methods on the CityFlowV2 benchmark.
MethodsYearIDF1 (%)IDP (%)IDR (%)Computational ComplexityLatencyScalability
MCVT-GCZ [47]202180.9585.6976.70Digital 06 00047 i001Digital 06 00047 i001Digital 06 00047 i002
BRM-MCMT [48]202284.8691.3779.21Digital 06 00047 i001Digital 06 00047 i001Digital 06 00047 i002
CityTrack [49]202384.9191.1580.42Digital 06 00047 i001Digital 06 00047 i001Digital 06 00047 i002
MCVT-CLM [50]202561.0764.7157.82Digital 06 00047 i002Digital 06 00047 i002Digital 06 00047 i001
MCMVT-YOLO [51]202581.6481.3281.97Digital 06 00047 i002Digital 06 00047 i003Digital 06 00047 i001
EMTMCT [52]202583.4485.9181.11Digital 06 00047 i001Digital 06 00047 i002Digital 06 00047 i002
IDP and IDR denote identification precision and identification recall, respectively. Digital 06 00047 i003, Digital 06 00047 i002, and Digital 06 00047 i001 indicate low, medium, and high levels, respectively. For computational complexity and latency, lower levels indicate lower deployment burden, whereas for scalability, higher levels indicate better scalability. The qualitative ratings are summarized from the reported pipeline design, detector/ReID backbone complexity, graph optimization strategy, camera-link dependency, online/offline association setting, and available runtime descriptions.
Table 4. Comparison of multi-modal 3D object detection methods on the nuScenes benchmark.
Table 4. Comparison of multi-modal 3D object detection methods on the nuScenes benchmark.
MethodsPublicationModalityLatency (ms)mAP (%)NDS (%)mATE (%)mASE (%)mAOE (%)mAVE (%)mAAE (%)
TiGDistill-BEV [71]TCSVT 2026C33743.4353.5657.8626.4642.9434.1520.17
FocalFormer3D [72]ICCV 2023L17866.3570.9227.9025.5829.3220.9818.75
RadarDistill [73]CVPR 2024R4214.1138.7951.8727.8557.9632.3512.63
MV2DFusion [74]TPAMI 2026C + L25273.0474.7728.2324.5025.2820.4719.01
ModalPatch [75]ICRA 2026C + L35664.9768.8432.6125.7328.2832.6517.17
GeoBEV [76]AAAI 2025C + L5742.9654.5854.5626.1538.4229.1620.76
IS-Fusion [77]CVPR 2024C + L29872.4873.6326.5825.1629.3426.4618.55
FUTR3D-CL [78]CVPR 2023C + L35170.2873.1231.1725.0319.7225.8618.39
RCTrans [79]AAAI 2025C + R5251.9059.3752.1827.1348.2420.4717.82
FUTR3D-LR [78]CVPR 2023L + R11639.8349.2350.4929.0957.9349.9819.38
FUTR3D-CLR [78]CVPR 2023C + L + R30867.0270.5632.6826.2726.5324.8919.11
Latency is measured in milliseconds on an RTX 4090 GPU. mAP denotes mean Average Precision; NDS denotes nuScenes Detection Score; mATE, mASE, mAOE, mAVE, and mAAE denote mean translation, scale, orientation, velocity, and attribute errors, respectively. C, L, and R denote camera, LiDAR, and radar, respectively. Higher mAP and NDS indicate better detection performance, whereas lower latency, mATE, mASE, mAOE, mAVE, and mAAE indicate lower computational or detection errors.
Table 5. Comparison of major video fusion strategies.
Table 5. Comparison of major video fusion strategies.
StrategyInputMechanismStrengthLimitationApplicationRef.
Perspective fusionMulti-view videoStitching; MVS; 3D reconstructionWide FoV; depthCalibration; parallax; occlusionVR; AR; mapping[32,36,39]
Spatio-temporal fusionMulti-frame videoTracking; temporal modeling; graph; transformerMotion consistencyLong sequence; high costTracking; traffic[41,43,46]
Deep fusionVisual; audio; text; sensorFeature-level DNN fusionRich representationHigh dimension; sensor failure; interferenceRetrieval; driving[66,67,68,69]
Hybrid; adaptive fusionMulti-level featuresMulti-stage; gating; dynamic pathsAccuracy-cost balanceTraining complexityReal-time fusion[81,82,83]
Attention; foundation fusionVideo; language; sensorsAttention; LMMs; adaptersSemantic reasoning; long-video understandingLatency; weak geometry; temporal groundingOpen-vocabulary; robotics[54,58,61,62]
Table 6. Application-oriented comparison of video fusion technologies.
Table 6. Application-oriented comparison of video fusion technologies.
ApplicationTechnologyRequirementBottleneckDirectionRef.
SurveillanceTracking; anomaly detection; privacy fusionCoverage; alertsOcclusion; privacy; associationFederated surveillance; edge analytics[89,90,91]
VR; ARStitching; view synthesis; consistencyImmersion; low latencyParallax; artifacts; delayNeural rendering; free-viewpoint video (FVV); edge XR[92,93,94]
Film; TVCGI compositing; virtual production; XRFidelity; seamless compositingLighting; costAI-assisted production; real-time rendering[95,96]
TransportationRadar–camera; traffic fusion; edge perceptionRobust detection; scalabilityWeather; occlusion; bandwidth6G sensing; collaborative perception[97,98,99,100]
Medical imagingCT; MRI; ultrasound; encrypted retrievalAccuracy; privacyRegistration; data sharingPrivacy-preserving fusion; foundation models[7,8,103]
Robotics; UAVsVisual–inertial; edge detection; VLMsNavigation; lightweight inferencePower; dynamic scenesEdge VLMs; energy-aware fusion[106,107,108]
Smart homeCamera-sensor; activity recognition; local inferencePrivacy; convenienceIndoor privacy; weak devicesOn-device fusion; federated learning[109,110]
Table 7. Key deployment constraints in real-time multi-camera video fusion systems.
Table 7. Key deployment constraints in real-time multi-camera video fusion systems.
ConstraintMain IssueTypical MitigationRef.
Computational complexityCalibration; feature extraction; deep inferenceCompression; lightweight models; offloading[106,108,119]
LatencyMulti-stream upload; cloud dependenceEdge inference; pipelined encoding; caching[107,118]
SynchronizationClock drift; frame mismatch; jitterTimestamp alignment; frame matching; buffering[111,112]
ScalabilityMore cameras; modalities; usersHierarchical fusion; scheduling; compact features[114,122]
Energy costContinuous sensing; encoding; inferencegreen AI; adaptive frame rate; energy-aware routing[123,124]
Table 8. Future research roadmap for multi-camera video fusion.
Table 8. Future research roadmap for multi-camera video fusion.
DirectionCore ChallengeResearch FocusRef.
Foundation-model-driven fusionGeometry; deployment costCross-view pretraining; adapters; video memory; temporal grounding[56,57,58,59,62]
Real-time edge fusionLatency; bandwidthEdge inference; pipelined sensing; adaptive caching; cooperative inference[107,108,118]
Privacy-preserving fusionSensitive data; model updatesFederated learning; encryption; anonymization[91,103,110,136]
6G-enabled collaborationDense camera networkingiIntegrated sensing and communication (ISAC); edge perception; semantic communication[120,121]
Green AI fusionEnergy-constrained devicesEnergy-aware routing; adaptive sensing; lightweight fusion[123,124]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Ma, C.; Xu, L. Harnessing Multi-Camera Video Fusion: Technologies, Applications, and Future Prospects. Digital 2026, 6, 47. https://doi.org/10.3390/digital6020047

AMA Style

Ma C, Xu L. Harnessing Multi-Camera Video Fusion: Technologies, Applications, and Future Prospects. Digital. 2026; 6(2):47. https://doi.org/10.3390/digital6020047

Chicago/Turabian Style

Ma, Chicheng, and Leiyang Xu. 2026. "Harnessing Multi-Camera Video Fusion: Technologies, Applications, and Future Prospects" Digital 6, no. 2: 47. https://doi.org/10.3390/digital6020047

APA Style

Ma, C., & Xu, L. (2026). Harnessing Multi-Camera Video Fusion: Technologies, Applications, and Future Prospects. Digital, 6(2), 47. https://doi.org/10.3390/digital6020047

Article Metrics

Back to TopTop