1. Introduction
Multi-camera video systems are increasingly used in surveillance and security [
1,
2], virtual and augmented reality (VR/AR) [
3,
4], intelligent transportation [
5,
6], and medical imaging [
7,
8]. Their value comes from complementary viewpoints and modalities, but this also creates specific fusion problems: cameras may have different fields of view, timestamps, exposure conditions, spatial resolutions, and semantic reliability. Video fusion [
9,
10] addresses these inconsistencies by aligning and integrating multiple streams into task-relevant representations. The main difficulty is therefore not simply collecting more visual data, but determining how geometric alignment, temporal synchronization, modality reliability, and deployment cost should be balanced for a given application.
Definition of video fusion. In this review, video fusion is defined as the process of integrating two or more temporally or spatially related video streams, visual observations, or auxiliary modalities from multiple cameras or sensors into a unified, more informative, and more consistent representation. This process usually involves calibration, geometric alignment, photometric correction, temporal synchronization, and information integration at the pixel, feature, or decision level. Its purpose is not only to improve visual quality, but also to expand the field of view, reduce uncertainty, compensate for occlusion or missing information, and enhance scene understanding in multi-camera or multimodal environments.
Previous studies show that multi-camera and multimodal fusion can improve depth estimation, 3D reconstruction, occlusion handling, target tracking, and autonomous-driving perception [
11,
12,
13,
14,
15]. However, these benefits depend on camera layout and fusion architecture. Overlapping cameras support geometric reasoning but require accurate calibration and sufficient shared view regions. Non-overlapping cameras are more suitable for wide-area monitoring, but they rely on trajectory continuity, identity association, and camera-topology modeling. Centralized fusion can exploit global information at the cost of bandwidth and latency, whereas distributed fusion improves scalability but makes cross-camera consistency harder to maintain. This trade-off motivates a review that connects fusion algorithms with application and deployment constraints.
Figure 1 illustrates (a) various ways in which fields of view can overlap [
16], and (b) an example of a non-overlapping arrangement in multi-camera systems. Its role is to connect camera layout with fusion strategy. Overlapping views support calibration, depth estimation, 3D reconstruction, and cross-view consistency because the same scene points can be observed from multiple angles. Non-overlapping views are more relevant to wide-area monitoring, cross-camera tracking, and trajectory association, where fusion depends less on pixel correspondence and more on temporal continuity, object identity, and camera topology. Thus, the figure is used as a structural guide for distinguishing geometry-oriented fusion from association-oriented fusion.
To clarify the scope of this review and improve the rigor of citation selection, relevant literature was selected through searches in Web of Science, Scopus, IEEE Xplore, ScienceDirect, ACM Digital Library, Google Scholar, and the proceedings or digital libraries of major venues such as CVPR, NeurIPS, ECCV, EMNLP, IEEE INFOCOM, and IEEE TPAMI. The main search period covered studies published from 2018 to 2026, while several earlier representative works were retained only when they provided foundational concepts or widely used methods. The search keywords included “video fusion”, “multi-camera fusion”, “multi-view video”, “multimodal fusion”, “foundation models”, “video understanding”, “geometric correction”, “image stitching”, “spatiotemporal fusion”, “edge video analytics”, “edge caching”, “real-time video processing”, and “privacy-preserving video analysis”. In particular, recent studies from 2024 to 2026 on multimodal foundation models, long-video understanding, camera–LiDAR fusion, edge-AI deployment, 6G-enabled perception, green AI, and privacy-preserving fusion were given special attention to improve the timeliness of the survey. We prioritized peer-reviewed journal and conference papers, highly cited review articles, and recent representative studies that are directly related to video fusion classification, application scenarios, technical challenges, and future research directions. Studies focusing only on single-camera enhancement, lacking an explicit fusion mechanism, or weakly related to multi-camera or multimodal video fusion were excluded.
To avoid loose or redundant citation practice, the related review papers are cited here according to their specific coverage rather than listed as a single undifferentiated group. Yeong et al. [
17] are cited for autonomous-vehicle sensor fusion because their review compares cameras, LiDAR, and radar under different driving conditions. Gandhi et al. [
18] are used to represent multimodal sentiment analysis, where the focus is datasets, modality interaction, and fusion methods. Kashinath et al. [
19] are cited for real-time multi-sensor traffic-flow analysis, while Qi et al. [
20] and Majumder et al. [
21] are cited only for IoT-enabled activity recognition and vision–inertial human action recognition, respectively. Yao et al. [
13] are used for radar–camera perception in autonomous driving. These references are therefore retained because each supports a distinct comparison dimension in
Table 1; they are not used as interchangeable evidence. Although these reviews provide valuable insights, they usually focus on a specific sensor combination, task, or application domain and rarely connect the full video fusion pipeline from pre-processing and correction to multi-dimensional fusion, application deployment, edge caching, and future research challenges. In contrast, this paper provides a holistic review of video fusion technology, covering its definition, classification, and application across various domains.
The novelty of this review lies in four aspects. First, it defines video fusion in the context of multi-camera and multimodal environments and organizes the field through a multi-dimensional taxonomy that links geometric correction, perspective fusion, spatio-temporal fusion, and multimodal fusion. Second, it connects low-level pre-processing methods with high-level learning-based fusion strategies, thereby showing how calibration, registration, color correction, deep fusion, adaptive fusion, and foundation-model-based fusion jointly form the video fusion pipeline. Third, it maps fusion technologies to practical application scenarios, including surveillance, VR/AR, film and television production, intelligent transportation, medical imaging, robotics/UAVs, and smart homes. Fourth, it extends the discussion beyond algorithmic accuracy by incorporating edge caching, real-time edge-AI deployment, privacy-preserving fusion, 6G-enabled collaboration, and green AI as deployment-oriented and future-facing dimensions.
Video fusion is treated in this review as a multi-stage system rather than as a single enhancement operation. The analysis follows the technical chain from geometric correction and visual consistency enhancement to perspective fusion, spatio-temporal fusion, multimodal fusion, and edge-assisted deployment. This organization is intended to address two gaps in task-specific reviews: first, many studies discuss individual methods without comparing their assumptions and limitations; second, deployment constraints such as latency, privacy, bandwidth, and energy consumption are often separated from the discussion of fusion algorithms. Therefore, this review compares both algorithmic mechanisms and application bottlenecks across surveillance, VR/AR, film production, intelligent transportation, medical imaging, robotics, UAVs, and smart homes.
In summary, this review is organized to reduce descriptive fragmentation in the literature.
Section 1 defines video fusion and clarifies the review scope.
Section 2 examines pre-processing and correction methods that determine whether multi-camera data can be reliably aligned.
Section 3 compares perspective, spatio-temporal, multimodal, and foundation-model-based fusion according to their assumptions and trade-offs.
Section 4 maps these methods to application bottlenecks rather than listing application fields alone.
Section 5 analyzes edge caching as a deployment mechanism that couples storage, inference, transmission, and fusion granularity.
Section 6 summarizes the resulting roadmap from the perspective of accuracy, latency, privacy, scalability, and energy cost.
3. Video Fusion Technology for Multi-Dimensional Classification and Advanced Application
This section systematically analyzes the multi-dimensional classification and advanced application of video fusion technology. The classification is primarily based on the technical characteristics of fusion processing and application requirements. The three main technological routes of video fusion are perspective fusion, spatio-temporal fusion, and multi-modal fusion. These focus on image stitching, depth information generation, 3D modeling and reconstruction, and the comprehensive processing of spatio-temporal information, respectively. This classification helps clarify the application scenarios and advantages of different technologies, promoting in-depth research and the widespread application of video fusion in fields such as surveillance, VR, and film and television production.
3.1. Perspective Fusion
Perspective fusion involves various essential techniques: image stitching, multi-view stereo fusion, and 3D modeling and reconstruction. These technologies are used to generate panoramic images, create 3D depth information, and reconstruct high-precision 3D models, respectively, thereby enhancing visual effects and broadening application possibilities in VR, AR, and digital heritage preservation.
3.1.1. Image Stitching
Image stitching is a key technique in perspective fusion, enabling the seamless merging of partially overlapping images from multiple viewpoints into a single high-resolution panoramic image. This process requires overcoming challenges such as parallax errors, scene movement, lighting changes, and uneven feature distribution. Abbadi et al. [
32] reviewed various panoramic image stitching techniques, focusing on how to merge multiple images into a high-quality panoramic image through feature matching, image alignment, and fusion. The review addresses the key challenges in stitching, including parallax errors, changes in lighting, and scene movement. These techniques provide solutions for each of these issues, optimizing the overall quality of the stitched image. Wang et al. [
33] elaborated on image stitching techniques specifically applied to video images. The article introduces the importance of image stitching in various fields, such as remote sensing, aerospace, VR, and medical imaging. Key processes in image stitching, including feature matching using the SIFT algorithm, image registration, and seam removal, are discussed. The feature matching stage involves identifying corresponding points in overlapping areas between two images, while the registration step aligns the images within a common coordinate system. Finally, the seam removal step eliminates visible seams caused by environmental factors such as exposure differences between images. Nie et al. [
34] proposed an unsupervised deep image stitching framework that consists of two stages: unsupervised rough image alignment and unsupervised image reconstruction. The framework uses a loss function based on ablation to constrain the affine network for large baseline scenes and introduces a transformer layer to distort the input image. The reconstruction stage uses features to minimize misalignment, thereby improving resolution through a two-branch network designed to refine the image.
The main limitation of image stitching is its dependence on stable overlap and consistent appearance. Dynamic objects, large parallax, weak texture, and exposure variation can turn seam removal into a temporal consistency problem rather than a purely spatial alignment problem.
3.1.2. Multi-View Stereo Fusion
Multi-view stereo fusion combines data from multiple viewpoints to estimate depth and construct 3D representations. Its performance depends on stereo correspondence, view-baseline selection, depth uncertainty, and memory consumption. Voynov et al. [
35] explored the incorporation of low-quality sensor depth data into multi-view stereo (MVS) methods and showed that auxiliary depth can improve reconstruction in regions with missing depth values. Cai et al. [
36] introduced MFNet, which uses a coarse-to-fine strategy and a multi-level fusion-aware feature pyramid to reduce cost-volume burden while estimating high-resolution depth. Lou et al. [
37] proposed the evidence-based local-global fusion (ELF) framework, which estimates uncertainty and fuses local and global evidence for stereo matching.
For multi-view stereo fusion, depth quality is constrained by occlusion, weak texture, view-baseline selection, and memory consumption. The practical bottleneck is whether depth estimation can remain reliable while reducing cost-volume size and inference latency for real-time or edge-assisted reconstruction.
3.1.3. 3D Modeling and Reconstruction
Three-dimensional modeling and reconstruction approaches create three-dimensional representations of objects or scenes from multi-view images, videos, or designed virtual assets. As shown in
Figure 3, the visual workflow is relevant to video fusion only when it is connected to measurable reconstruction criteria, such as geometric accuracy, completeness, runtime, and memory cost. Alldieck et al. [
38] reconstructed 3D human body models from monocular RGB video by transforming dynamic body postures into a standard reference frame through deformation cancellation, reporting a reconstruction accuracy of 4.5 mm. Sun et al. [
39] introduced NeuralRecon, which reconstructs local scene surfaces as sparse truncated signed distance function (TSDF) volumes from monocular video. These methods show that 3D reconstruction should be evaluated as a fusion problem involving view geometry, temporal consistency, and computational efficiency rather than as a visual modeling example alone. González Izard et al. [
40] detailed the process of 3D modeling and reconstruction using volume rendering techniques (such as vtkFlyingEdges3D) to generate isosurfaces. The technique then reduces polygon counts using vtkDecimatePro and applies Laplace smoothing filters for a finer 3D model. These reconstructed models can be exported in formats like .obj or .stl for 3D printing and visualization in augmented or virtual reality.
Although
Figure 3 is visually illustrative, it corresponds to a measurable 3D reconstruction workflow. For example, Alldieck et al. [
38] reconstructed clothed 3D human models from monocular RGB video with a reported reconstruction accuracy of 4.5 mm, and the mean average errors on BUFF, D-FAUST, and KinectCap were 5.37 mm, 4.44 mm, and 3.97 mm, respectively. NeuralRecon further shows the real-time requirement of this pipeline: it reconstructs sparse TSDF volumes from monocular video at 30 ms per key frame, or about 33 key frames per second, which is approximately 10 times faster than Atlas while maintaining competitive 3D geometry quality on ScanNet [
39]. These results indicate that 3D reconstruction figures should be read together with accuracy, completeness, F-score, runtime, and memory cost rather than only as visual demonstrations.
For 3D modeling and reconstruction, the key bottleneck is the joint control of surface detail, temporal coherence, and runtime. A method that improves geometric completeness but increases memory use or frame latency may be unsuitable for VR, AR, and online multi-camera applications.
3.2. Spatio-Temporal Fusion
Spatio-temporal fusion involves the comprehensive processing and integration of information from multiple video sources across both spatial and temporal dimensions. This fusion aims to extract richer and more accurate spatio-temporal data to enhance video analysis and understanding. Different levels of fusion methods, including pixel-level fusion, feature-level fusion, and decision-level fusion, can be applied depending on the specific application. Wang et al. [
41] proposed a spatio-temporal fusion-based method for mobile target tracking and segmentation, specifically addressing the challenges of target loss during occlusion. The approach combines Kalman filtering with SiamMask technology, establishes a motion model, and employs an elliptical fitting strategy to assess the bounding box angle and size. The use of an attention mechanism further enhances focus on the target’s primary area, minimizing the impact of background distractions. This method demonstrated robust tracking performance in various complex environments. He et al. [
42] developed a method for spatio-temporal fusion of multi-source remote sensing images, which integrates linear stretching (Ls), maximum value composition (MVC), and flexible spatio-temporal data fusion (FSDAF). This method aims to map abandoned land distribution in remote sensing images, overcoming challenges related to fragmented terrain and cloud pollution. Liang et al. [
43] introduced the spatial–temporal feature fusion enhancement (STFFE) method, which enhances the discriminative ability of video segments by fusing spatial and temporal features. This method leverages the top-k mechanism for feature selection and temporal information fusion to improve anomaly detection accuracy. The approach effectively increases the model’s ability to distinguish between normal and abnormal features, improving the accuracy of anomaly detection in videos. Xu et al. [
44] proposed the nonlocal spatial–temporal feature fusion network (NLMF-Net) for real-time infrared small target detection. This model fuses spatio-temporal information in the feature field and improves detection performance while maintaining real-time processing capabilities. By using high-confidence correlation operations between current and past frames, the model achieves a significant performance boost without requiring substantial computational resources. Fu et al. [
45] presented Cuboid-Net, a multi-branch convolutional neural network for joint spatio-temporal video super-resolution. The model treats input low-resolution video as a cubic structure, dividing it into slices to feed into different branches for directional processing. The method includes various enhancement modules for feature extraction and quality enhancement, improving the resolution of video frames while maintaining real-time processing capabilities. Jiang et al. [
46] introduced the decomposed spatio-temporal fusion graph convolutional network (DSTGCN), a spatio-temporal fusion traffic prediction model based on input traffic signal decomposition. This model uses graph convolutional networks to capture global spatial information and temporal dependencies to predict traffic conditions, offering significant improvements in traffic flow predictions.
Taken together, these studies show that spatio-temporal fusion is not a single algorithmic category, but a set of methods with different assumptions about temporal continuity, spatial correspondence, and computational budget. CNN-based and graph-based methods are efficient when spatial relations are relatively stable, but they may struggle when cross-camera topology changes or when long-range dependencies dominate. Transformer-based and recurrent models can capture longer temporal contexts, but their cost increases with video length, resolution, and camera number. For multi-camera systems, the central trade-off is therefore not only accuracy, but also whether temporal association, spatial alignment, and deployment latency can be optimized jointly.
Multi-camera vehicle tracking provides a representative downstream application scenario for spatio-temporal fusion because it requires not only the spatial association of vehicle appearances across different camera views, but also the temporal integration of trajectories, transition intervals, motion continuity, and cross-camera identity consistency. Therefore,
Table 3 is used here as an empirical comparison rather than a simple list of methods. It shows how different designs balance ID accuracy, computational complexity, latency, and scalability on the CityFlowV2 benchmark.
Early high-performing methods mainly adopt multi-stage offline pipelines, including vehicle detection, Re-ID feature extraction, single-camera tracking, and inter-camera trajectory association. Liu et al. [
47] introduced crossroad-zone guidance and direction-based temporal masks to constrain cross-camera matching regions and reduce trajectory association ambiguity. Yang et al. [
48] further improved inter-camera association by employing box-grained re-ranking matching, achieving an identification F1 (IDF1) score of 84.86%. CityTrack [
49] combined location-aware single-camera tracking with box-grained matching, obtaining the highest IDF1 score of 84.91% among the compared methods. These results indicate that explicit spatial–temporal priors and global association can improve identity consistency, but they also create a practical cost: the pipeline depends on strong detectors, heavy Re-ID backbones, and offline trajectory optimization.
Recent studies have shifted toward more deployable, scalable, and low-latency designs for practical intelligent transportation applications. Lin et al. [
50] proposed a self-supervised camera-link model to automatically infer camera relationships, thereby reducing dependence on manually designed spatial–temporal constraints and improving scalability in large camera networks, although its tracking accuracy remains lower than heavily optimized offline systems. Huang et al. [
51] designed an online multi-camera multi-vehicle tracking framework based on lightweight YOLO11, improved single-camera association, and hierarchical cross-camera clustering, achieving 81.64% IDF1 with low latency and real-time edge deployment capability. Tseng et al. [
52] further enhanced YOLOv9 with attention mechanisms to improve vehicle detection robustness and spatial–temporal association performance, achieving 83.44% IDF1 on CityFlowV2. The comparison therefore suggests a deployment-oriented trade-off: offline methods generally obtain stronger identity consistency through richer appearance modeling and global association, whereas online, lightweight, and self-supervised methods sacrifice part of the IDF1 score to reduce latency, manual topology design, and deployment cost.
3.3. Multi-Modal Fusion
Multi-modal fusion integrates various modalities of information from videos, such as visual, audio, and text, to enable more comprehensive and accurate understanding and analysis. This technology enhances recognition, classification, and retrieval capabilities, allowing systems to better understand complex scenes and events in videos. Applications span fields like autonomous driving, surveillance, video search, and sentiment analysis.
3.3.1. Foundation-Model-Based Fusion
Foundation models introduce a new route for multi-modal video fusion by learning reusable representations across text, image, video, audio, sensor, and other data types. Compared with task-specific fusion modules, generalist multi-modal models provide a shared representation space that can potentially align camera streams with language, audio, depth, LiDAR, radar, and contextual signals [
53,
54,
55]. Video-LLaVA, for example, studies unified visual representation learning for image and video inputs before projection into a language model, which is relevant to video fusion because it shifts the fusion objective from simple feature concatenation to cross-modal representation alignment [
54]. This ability is particularly useful for multi-camera systems that require both semantic reasoning and cross-modal alignment.
Recent video-oriented large multi-modal models and benchmarks show this trend more clearly. Video-XL reduces the cost of hour-scale video understanding through visual summarization tokens and dynamic compression [
56], while Apollo analyzes how video sampling, architecture, data composition, and training schedules affect video large multimodal model (LMM) performance [
57]. Video-MME provides a comprehensive benchmark for evaluating multi-modal LLMs in video analysis [
58], and LongVideoBench focuses on long-context interleaved video-language understanding [
59]. Streaming Long Video Understanding and LongVLM further indicate that long-video reasoning requires memory-efficient temporal modeling and streaming or compressed representations rather than treating all frames equally [
60,
61]. A recent survey on video temporal grounding with multimodal large language models also shows that temporal localization, language-guided reasoning, and long-range video context have become central issues in current video understanding research [
62]. These studies suggest that video fusion is moving from hand-designed concatenation toward foundation-model-based alignment, compression, grounding, and reasoning. The main bottlenecks are computational cost, weak modeling of calibrated cross-view geometry, limited synchronization awareness, and difficult multi-camera data curation. These bottlenecks make lightweight adapters, retrieval-augmented video memory, cross-view pretraining, temporal grounding, and edge-deployable foundation models more relevant than generic model scaling alone.
3.3.2. Deep Fusion
Deep fusion combines modality-specific representations within deep neural networks to improve prediction and scene understanding. Earlier deep-fusion studies are retained only as methodological background, including shared-private DNNs for emotion recognition [
63], multi-layer fusion for video classification [
64], and split-attention modules for flexible CNN/RNN integration [
65]. More recent references are used for current technical directions, such as transformer-based video retrieval [
66], adaptive camera–LiDAR fusion [
67], and resilient sensor fusion under adverse sensor failures [
68]. This citation arrangement separates historical fusion modules from recent robustness-oriented and representation-oriented fusion studies.
To clarify the structural differences among common DNN-based fusion pipelines,
Figure 4 separately compares early fusion, intermediate fusion, late fusion, and hybrid fusion. The figure is used as an architectural comparison because the position of modality interaction determines the main trade-off of each pipeline. Early fusion can exploit low-level complementarity but is sensitive to misalignment and noise. Intermediate fusion preserves modality-specific encoders while allowing feature-level interaction. Late fusion is easier to deploy with separate models but may miss fine-grained cross-modal dependencies. Hybrid fusion increases flexibility by combining multiple stages, but it also increases model complexity and training difficulty.
Zhou et al. [
69] proposed a feature-level fusion approach that addresses perspective differences between radar and camera features by projecting them into a Bird’s Eye View (BEV) representation. Chen et al. [
70] introduced a connected-vehicle-assisted roadside radar and video data fusion framework. As shown in
Figure 5, radar and camera streams are not merely combined as parallel inputs; the framework uses connected-vehicle information as calibration data for a backpropagation (BP) network and dynamically updates training samples according to road conditions. The figure therefore illustrates a traffic-specific fusion mechanism in which heterogeneous sensing, communication, calibration, and flow-state estimation are coupled.
As shown in
Figure 6, deep fusion can transform heterogeneous observations, including LiDAR points and camera images, into aligned representations for object detection, localization, and scene understanding. Recent camera–LiDAR fusion studies further show that feature fusion should adapt to modality reliability and failure conditions. GAFusion adaptively fuses LiDAR and camera features with multiple guidance cues for 3D object detection [
67], while resilient sensor fusion uses multi-modal expert fusion to improve robustness under adverse sensor failures [
68]. The figure therefore supports the view that deep fusion is not a simple numerical combination of sensor outputs, but a representation-learning process that must preserve complementary cues while suppressing modality conflict. Its practical value is clear in autonomous driving, robot navigation, and spatial analysis, but real-time deployment still depends on reducing feature dimensionality, controlling cross-modal interference, and simplifying network structures without sacrificing accuracy.
To further provide a task-consistent quantitative comparison,
Table 4 summarizes representative camera-, LiDAR-, radar-, and multi-modal 3D object detection methods on the nuScenes benchmark. All entries focus on the same detection task and use the same nuScenes metric family, including mAP, NDS, latency, and detection-error metrics; however, because the reported methods still differ in backbone, implementation details, and training settings, the comparison is intended as a quantitative reference for modality and design trade-offs rather than a strict ranking. Overall, camera–LiDAR fusion achieves the strongest accuracy in this table, with the best entries reaching 73.04% mAP and 72.48% mAP, which reflects the advantage of combining image semantics with LiDAR geometry. In contrast, radar-centered designs are generally faster but less accurate: the radar-only method has the lowest latency of 42 ms but only 14.11% mAP, while the camera–radar method reaches 51.90% mAP and 59.37% NDS at 52 ms, suggesting that radar can provide useful motion and localization cues but remains limited by sparse and noisy measurements. The results also show that using more modalities does not automatically yield better performance, as the camera–LiDAR–radar setting obtains 67.02% mAP and 70.56% NDS with 308 ms latency, which is below the best camera–LiDAR results. Therefore, this comparison supports the view that multi-modal fusion improves 3D detection only when complementary information is effectively aligned, and that practical deployment must balance accuracy, localization error, and latency rather than maximizing the number of input modalities.
3.3.3. Hybrid Fusion
Hybrid fusion combines multiple fusion strategies, integrating information from different modalities at various levels. This approach aims to overcome the limitations of traditional deep and shallow fusion methods, ensuring more comprehensive and robust fusion of multi-modal data. Hybrid fusion can dynamically adjust its fusion strategies based on the specific requirements of the task at hand. Li et al. [
80] proposed a Transformer-based video captioning model called MFVC (Multi-modal Fusion for Video Caption), addressing the low performance and high computational complexity of existing multi-modal fusion methods. The introduction of audio modality data and an attention bottleneck module in this model improves performance while reducing operational costs, making the approach more efficient. Joze et al. [
81] introduced the multi-modal Transfer module, which achieves slow modality fusion by adding fusion units at various levels of the CNN. This module addresses the issue that traditional deep and shallow fusion methods often fail to fully utilize multi-modal data. It uses squeeze and excitation operations to recalibrate channel features across different CNN streams, facilitating feature fusion across convolutional layers with varying spatial dimensions. Vielzeuf et al. [
82] proposed a multi-modal fusion method where each modality is processed by an independent deep convolutional network, with a central network that connects these modality-specific networks. This central network provides common feature embeddings and regularizes modality-specific networks through multi-task learning, improving the accuracy of existing multi-modal fusion methods on multiple computer vision tasks.
Hybrid fusion faces several challenges, including the effective integration of results from fusion at different levels, avoiding conflicts between different fusion methods, and simplifying model structures to reduce the number of parameters. Although hybrid fusion combines different strategies to improve robustness, it may also increase model complexity, making training and optimization more difficult. The large number of adjustable parameters can also complicate the process of finding the optimal combination for the model, requiring careful balancing of efficiency and performance.
3.3.4. Adaptive Fusion
Adaptive fusion refers to methods that adjust the fusion strategy based on the dynamic characteristics of the data being processed. By adapting the fusion strategy in real-time, adaptive fusion models improve the flexibility and generalization of multi-modal fusion systems. Xue et al. [
83] proposed DynMM, a dynamic multi-modal fusion method that can adaptively fuse multi-modal data during inference. By introducing gating functions and resource-aware loss functions, DynMM reduces computational costs while maintaining accuracy. This method is particularly effective when processing different types of multi-modal data, enhancing the adaptability of fusion systems. Wu et al. [
84] introduced a denoising bottleneck fusion (DBF) model that addresses redundancy and noise issues in multi-modal signals. By utilizing a bottleneck mechanism and a mutual information maximization module, DBF preserves key information while denoising the multi-modal data. This model has shown significant improvements in multi-modal emotion analysis and summarization tasks.
Adaptive fusion is useful because it adjusts modality weighting or fusion paths according to input conditions, but this flexibility introduces a stability–cost trade-off. Dynamic gating can reduce unnecessary computation when modalities are redundant, yet it may also produce unstable decisions under noisy, missing, or conflicting inputs. Therefore, adaptive fusion should be evaluated not only by average accuracy, but also by inference cost, gating stability, and robustness to modality degradation.
3.3.5. Attention-Based Fusion
Attention-based fusion leverages attention mechanisms to selectively focus on the most important information from different modalities. This approach improves the model’s recognition accuracy by highlighting the most relevant features and enhancing the interpretability of the model’s decision-making process. Liu et al. [
85] proposed a hierarchical attention-based multi-modal fusion network for video emotion recognition. The network includes a multi-modal feature extraction module and a multi-modal feature fusion module. It solves the problem of emotional cue variation in different video frames by using a local attention network to address emotional differences between frames and a global attention network to handle emotional discrepancies between different modalities. Earlier attention-based fusion studies are cited here only to show the development of modality weighting and selective feature use. Low-rank tensor self-attention fusion provides an example of parameter-efficient multimodal interaction [
86], while early video-description work illustrates word-conditioned modality selection [
87]. The recent foundation-model-based video understanding studies discussed above are used as the main evidence for current video-language reasoning and long-context multimodal understanding.
Attention-based fusion offers the advantage of identifying and highlighting the most important information from different modalities, improving the model’s accuracy. The attention mechanism also adds interpretability, helping to understand the model’s decision-making process. However, determining an effective attention distribution can be challenging, especially with complex multi-modal data. Designing attention mechanisms that can generalize to different tasks and datasets while improving robustness remains a challenge. Current research focuses on designing attention mechanisms that can adapt to various tasks, improving generalization and robustness.
3.4. Other Types of Fusion
The fusion method proposed in [
88] introduces a new approach that does not fit into the traditional categories of fusion discussed earlier. This method presents a neural multi-modal cooperative learning (NMCL) model that explicitly distinguishes and processes consistent and complementary features in multi-modal data. Using a relation-aware attention mechanism, the model separates the consistent and complementary components by learning thresholds. It then integrates the consistent features to enhance the representation and complements the complementary parts to strengthen the information within each modality. This approach specifically addresses the challenges of modal information consistency and complementarity, particularly in the context of micro-video understanding.
A cross-method comparison shows that the three main fusion routes solve different parts of the multi-camera problem and therefore should not be treated as interchangeable techniques. Perspective fusion is geometry-constrained: it is effective for panoramic stitching, 3D reconstruction, and free-viewpoint rendering when calibration is reliable and view overlap is sufficient, but it is vulnerable to parallax, occlusion, exposure inconsistency, and moving objects. Spatio-temporal fusion is association-constrained: it is suitable for tracking, traffic analysis, anomaly detection, and video prediction, but its performance depends on temporal continuity, camera topology, and the ability to prevent error accumulation across long sequences. Multi-modal fusion is representation-constrained: it can combine visual, audio, textual, LiDAR, radar, and semantic cues, but it must handle modality conflict, sensor failure, missing inputs, and unequal reliability across modalities. Foundation-model-based fusion extends multi-modal fusion by improving semantic generalization and language-guided reasoning, yet it introduces new constraints related to cross-view geometry, synchronization awareness, data curation, and edge deployment cost. This comparison indicates that the main research issue is no longer simply which fusion model is more accurate, but under what assumptions a method remains reliable, scalable, and deployable.
Table 5 compares the major fusion strategies discussed in this section. The table uses short keywords to highlight the main differences, while the text provides the interpretation. Perspective fusion is geometry-oriented and is suitable for extending the field of view or reconstructing 3D scenes, but it depends on accurate calibration and sufficient view overlap. Spatio-temporal fusion is temporal-relation-oriented and is useful for tracking, prediction, anomaly detection, and traffic analysis, but the cost increases with video length, resolution, and camera number. Deep, hybrid, adaptive, and attention-based fusion methods are more flexible for multimodal perception, yet they must balance feature richness, modality conflict, model complexity, and deployment cost. Foundation-model-based fusion improves semantic generalization, but current models still need better cross-view geometry, synchronization awareness, and edge efficiency.