A Survey of Multi-Model Collaboration in Video Understanding
Abstract
1. Introduction
- We provide a structured analysis of video understanding from a collaborative systems perspective by characterizing task requirements, collaborative components, and coordination challenges across heterogeneous video understanding scenarios.
- We introduce a unified system-level framework for collaborative video understanding, in which collaborative methods are analyzed through three core components: functional units, inter-unit communication mechanisms, and collaborative state representations.
- We further organize existing collaborative methods from the perspective of coordination mechanisms into two overarching paradigms, namely static collaboration and dynamic collaboration, revealing a progression from predefined execution structures toward adaptive and increasingly autonomous coordination.
- We review representative benchmarks, evaluation metrics, and empirical evidence across major video understanding tasks, and discuss why current evaluation protocols remain insufficient for measuring collaborative organization, memory use, adaptive execution, and system-level collaborative capability.
- We identify key challenges and future directions for collaborative video understanding, including structured task orchestration, semantically aligned inter-unit communication, persistent shared memory, uncertainty-aware error containment, evidence-grounded reasoning, trustworthy collaboration, and collaboration-centric evaluation.
2. Video Understanding Problem Definition
2.1. Problem Definition
- Recognition and Localization Tasks identify entities or events in video content and estimate their spatial or temporal locations [43,44]. The output Y typically takes the form of categorical labels, spatial regions, temporal segments, trajectories, or spatio-temporal event locations. Representative tasks include action recognition, event classification, object tracking, temporal action localization, and spatio-temporal action localization [45]. These tasks primarily rely on video-internal evidence to determine what occurs and where or when it occurs [46].
- Cross-modal Matching and Grounding Tasks establish semantic correspondences between video content and external modalities, such as language or audio [47]. The output Y typically takes the form of retrieval rankings, similarity scores, matched video-text pairs, or spatially/temporally grounded regions or segments associated with an external query. Representative tasks include video-text retrieval, video-language alignment, video grounding, and audio-visual grounding [48,49]. Their defining characteristic is the explicit alignment of video content with information from another modality [50].
- Generative Description Tasks convert video content into natural-language descriptions or condensed semantic summaries [51,52]. The output Y is typically a free-form or structured textual representation, such as a caption, a sequence of event descriptions, or a condensed summary. Representative tasks include video captioning, dense video captioning, video summarization, and event description [48,49,53]. These tasks emphasize semantic abstraction and narrative organization by converting observed entities, actions, and temporal relations into coherent linguistic representations [50].Reasoning and Decision-making Tasks infer implicit information beyond directly observable video content by integrating evidence across temporal, semantic, and multimodal contexts [26,54]. The output Y may take the form of answers, inferred temporal or causal relations, explanations, decisions, or action plans. Representative tasks include video question answering, temporal reasoning, causal reasoning, commonsense reasoning, embodied video reasoning, and video-agent decision-making [55,56]. These tasks are closely related to collaborative video understanding because they often require coordinated perceptual grounding, evidence retrieval, memory maintenance, and multi-step reasoning [57].

2.2. Fundamental Characteristics of Video Understanding
- Event-Centric Semantics. The semantic meaning of a video is primarily determined by events and interactions that evolve over time rather than isolated visual observations [62]. Consequently, effective video understanding requires identifying participating entities, modeling their relationships, and characterizing event evolution [36]. For example, understanding an interaction such as a person picking up an object and handing it to another person requires coordinated entity perception, relation modeling, and event-level temporal reasoning.
- Multi-level and Multimodal Reasoning. Effective understanding requires integrating information across multiple semantic levels, ranging from visual appearance and motion patterns to event structures, semantic relationships, and contextual knowledge [35]. In addition, video content is often associated with complementary modalities such as language and audio, which provide information unavailable from visual observations alone [63]. Consequently, modern video understanding increasingly relies on multi-level and multimodal reasoning to integrate heterogeneous information sources and construct comprehensive semantic interpretations of dynamic scenes [64].
3. Collaborative Video Understanding Framework
3.1. Collaborative Components
3.2. Coordination Mechanisms
3.2.1. Static Collaboration
3.2.2. Dynamic Collaboration
Controller-Based Collaboration
Agent-Based Collaboration
4. Benchmarks, Metrics, and Empirical Analysis
4.1. Representative Benchmarks
- Recognition and Localization Benchmarks evaluate whether systems can identify actions or events and locate their spatial or temporal occurrences. Representative datasets cover action recognition, temporal action localization, and spatio-temporal localization, with annotations ranging from category labels to temporal segments and spatial regions [10,133]. These benchmarks primarily assess visual perception, motion modeling, and temporal aggregation under predefined label spaces and explicit supervision. For collaborative video understanding, they provide evidence for coordinated perceptual and temporal processing, while offering limited evaluation of long-range memory, multi-step reasoning, and adaptive execution [132].
- Cross-modal Matching and Grounding Benchmarks evaluate whether systems can align video content with language or other modalities. Representative datasets cover video-text retrieval, temporal grounding, query-conditioned moment retrieval and highlight detection, as well as visually grounded video question answering, with annotations including paired instances, ranked results, and temporal spans [137,138]. These benchmarks further differ in whether grounding is evaluated through global retrieval, localized temporal evidence, saliency or highlight identification, or explicit answer-evidence alignment. These benchmarks assess multimodal alignment and retrieval capabilities but exhibit substantial heterogeneity in supervision forms, query formulations, annotation granularity, and modality configurations, which limits direct comparison across datasets [156].
- Generative Description Benchmarks evaluate whether systems can generate high-level textual representations from video content. Representative datasets cover video captioning, dense captioning, event description, and summarization, with annotations typically provided as sentence-level or event-level references [157,158]. These benchmarks assess semantic abstraction, temporal organization, and content selection, but their performance interpretation is highly sensitive to reference quality, annotation diversity, and evaluation protocols, complicating fair comparison across datasets.
- Reasoning and Decision-making Benchmarks evaluate whether systems can infer implicit information by integrating temporal, semantic, and multimodal evidence. Representative datasets cover video question answering, causal reasoning, commonsense inference, and decision-oriented tasks, requiring stronger capabilities in evidence integration, memory utilization, and reasoning consistency [131,150,159]. Recent benchmarks further extend this category toward long-form and long-context video understanding, requiring models to integrate evidence over substantially longer temporal horizons and, in some cases, across multiple input modalities such as visual, audio, and subtitle information [132,153,154]. However, variations in task formulations, input configurations, and supervision strategies introduce substantial challenges for fair comparison across benchmarks. In more advanced long-horizon and agentic settings, these tasks further require coherent state maintenance, tool use, planning, and adaptive execution across extended inference processes [160].
4.2. Evaluation Metrics
- Task Outcome Metrics. Task outcome metrics remain the predominant evaluation paradigm in existing video understanding benchmarks. Recognition tasks commonly rely on accuracy-based measures, such as Top-1 and Top-5 accuracy, whereas localization tasks use overlap-based measures, including mAP under different temporal IoU thresholds and IoU-based matching criteria [161,162]. Cross-modal retrieval tasks typically employ ranking metrics, such as Recall@K, Mean Rank, and Median Rank, while temporal grounding tasks additionally use IoU-based criteria to assess localization accuracy [163,164]. Generation tasks use reference-based metrics, such as BLEU, METEOR, CIDEr, and SPICE, to measure textual similarity and semantic consistency, whereas dense video captioning additionally employs the task-specific SODAc metric [161,165]. Reasoning-oriented and composite video understanding tasks typically adopt accuracy-based or task-specific measures, including answer accuracy, grounded answer accuracy, category-level accuracy, event localization accuracy, and summary-based question answering accuracy [151]. Decision-oriented settings may additionally use success rate or task completion rate to assess decision effectiveness. These examples are representative rather than exhaustive, as task-specific benchmarks may adopt additional evaluation measures tailored to their particular objectives. Although these metrics effectively characterize task outcomes, they primarily evaluate final predictions under predefined task settings and provide limited insight into system-level behaviors.
- Beyond Task-level Evaluation. Beyond task-level outcomes, recent evaluations have increasingly considered system-level properties, including output quality, consistency, and efficiency [17,166]. Quality and consistency assessments complement conventional task metrics by examining textual or semantic similarity, cross-modal alignment, and temporal consistency, while human and model-based evaluations provide additional assessment for open-ended outputs [167]. Efficiency-oriented criteria further characterize the computational and interaction costs of collaborative systems, including token cost, frame count, frame efficiency, inference cost, runtime, FLOPs, and latency [168]. These measures are particularly relevant to multi-model and agent-based systems, where adaptive coordination may introduce additional computation, iterative interactions, or redundant evidence processing [169]. However, such system-level evaluations remain less standardized than conventional task metrics, and their adoption varies substantially across benchmarks and studies [170].
4.3. Quantitative Performance Trends and Comparative Analysis
- Recognition and Localization. Table 5 summarizes representative quantitative evidence for both non-collaborative and collaborative methods across recognition and localization tasks. Recognition and localization exhibit the clearest gains across successive technical paradigms. On Kinetics-400, Top-1 accuracy increases from 74.7 for TSM to 86.1 for MViTv2, 90.0 for VideoMAE V2, and 92.1 for InternVideo2. SSv2 shows a similar but increasingly saturated trajectory, improving from 64.3 for TSM to 73.3 for MViTv2 and 77.4 for InternVideo2. For temporal localization, InternVideo2 reaches 41.2 average mAP on ActivityNet v1.3 and 72.0 on THUMOS14, exceeding InternVideo by 2.2 and 0.4 points, respectively. On AVA v2.2, frame-mAP increases from 27.4 for X3D to 42.6 for VideoMAE V2 among the representative non-collaborative methods. The collaborative methods show a more heterogeneous pattern: Feature Hallucination VideoMAE V2 and Feature Hallucination InternVideo2 reach 87.5 and 91.6 on Kinetics-400, respectively, while their SSv2 results are 77.4 and 77.3. Other collaborative methods report task-specific results, including 68.3 for TrAction + V-JEPA 2-L on SSv2, 49.8 for MLLM4WTAL on THUMOS14, and 45.1 for LART-Hiera on AVA v2.2. These results show substantial gains from advances in model architectures and large-scale pre-training, while the relatively small gap among the strongest recent models on several established recognition benchmarks suggests an emerging saturation trend. Collaborative methods, however, do not exhibit a uniform performance advantage, partly because they target different objectives and are evaluated under heterogeneous backbone, training, and computational settings. Therefore, controlled comparisons with matched single-model baselines and computational budgets are needed to isolate and quantify the actual contribution of collaboration.
- Cross-modal Matching and Grounding. Cross-modal matching and grounding show substantial but less uniform gains across the available dataset–metric combinations. On MSR-VTT retrieval, the reported Recall@1 values range from 20.9 for Collaborative Experts and 29.6 for TeachText to 62.8 for InternVideo2 and 78.6 for MAVIS. On Charades-STA, InternVideo2 obtains 70.03 Recall@1, IoU = 0.5, compared with 57.91 for EMTM and 38.4 for Moment-GPT; on QVHighlights, InternVideo2 reaches 49.24 moment-retrieval mAP, compared with 43.63 for UniVTG. As reflected in Table 6, these results demonstrate substantial progress in cross-modal matching and grounding, but the magnitude of improvement remains highly task-dependent. The mixed results of collaborative methods further suggest that the effectiveness of collaboration depends on how the collaboration mechanism aligns with the specific cross-modal objective and evaluation protocol.
| Method Type | Method | Video Recognition | Temporal Localization | Spatio-Temporal Action Localization | ||
|---|---|---|---|---|---|---|
| Kinetics-400 Top-1 | SSv2 Top-1 | ActivityNet v1.3 Avg. mAP | THUMOS14 Avg. mAP | AVA v2.2 Frame-mAP | ||
| Non-collaborative | TSM [171] | 74.7 | 64.3 | – | – | – |
| X3D [172] | 79.1 | – | – | – | 27.4 | |
| UniFormer [173] | 83.0 | 71.4 | – | – | – | |
| Video Swin [174] | 84.9 | 69.6 | – | – | – | |
| MViTv2 [175] | 86.1 | 73.3 | – | – | 34.4 | |
| MaskFeat [176] | 87.0 | 75.0 | – | – | 39.8 | |
| VideoPrism [177] | 87.2 | 68.5 | 37.8 | – | – | |
| VideoMAE [178] | 87.4 | 75.4 | – | – | 39.5 | |
| VideoMAE V2 [179] | 90.0 | 77.0 | – | 69.6 | 42.6 | |
| InternVideo [180] | 91.1 | 77.2 | 39.0 | 71.6 | 41.0 | |
| InternVideo2 [181] | 92.1 | 77.4 | 41.2 | 72.0 | – | |
| Collaborative | TrAction + DINOv2-B [182] | – | 62.3 | – | – | – |
| TrAction + V-JEPA 2-L [182] | – | 68.3 | – | – | – | |
| LART-Hiera [183] | – | – | – | – | 45.1 | |
| GAP + CLIP [184] | – | – | 31.8 | 32.9 | – | |
| MLLM4WTAL [185] | – | – | – | 49.8 | – | |
| JEDI Student [186] | 82.08 | – | – | – | – | |
| Feature Hallucination VideoMAE V2 [187] | 87.5 | 77.4 | – | – | – | |
| Feature Hallucination InternVideo2 [187] | 91.6 | 77.3 | – | – | – | |
| Method Type | Method | Video-Text Retrieval | Temporal Grounding | Moment Retrieval | |
|---|---|---|---|---|---|
| MSR-VTT Recall@1 | Charades-STA Recall@1, IoU = 0.5 | ActivityNet Captions Recall@1, IoU = 0.5 | QVHighlights mAP | ||
| Non-collaborative models | LGI [188] | – | 59.46 | 41.51 | – |
| Moment-DETR [142] | – | 55.65 | – | 36.14 | |
| UniVTG [189] | – | 60.19 | – | 43.63 | |
| QD-DETR [190] | – | 57.31 | – | 40.19 | |
| VideoCLIP [191] | 30.9 | – | – | – | |
| Frozen in Time [192] | 32.5 | – | – | – | |
| CLIP4Clip [193] | 44.5 | – | – | – | |
| InternVideo2 [181] | 62.8 | 70.03 | – | 49.24 | |
| Collaborative models | Moment-GPT [194] | – | 38.4 | 31.1 | 35.0 |
| EMTM [195] | – | 57.91 | 44.73 | – | |
| Collaborative Experts [77] | 20.9 | – | – | – | |
| TeachText [196] | 29.6 | – | – | – | |
| Learning from Image Captioning [197] | 39.2 | – | – | – | |
| MAVIS [198] | 78.6 | – | – | – | |
- Generative Description. Generative description exhibits substantial progress, although the strongest method varies across datasets and metrics. Table 7 summarizes representative results from both non-collaborative and collaborative methods across global and dense video captioning. On MSVD, CIDEr increases from 49.8 for RETTA and 51.7 for Enc–Dec + Local + Global to 146.2 for Vid2Seq, while Vid2Seq reaches 64.6 on MSR-VTT compared with 53.8 for both SwinBERT and Multiple Networks Consensus + OracleNet. Among the listed methods, no single method consistently dominates all ActivityNet Captions metrics: Masked Transformer records the highest METEOR score of 9.56, while Vid2Seq achieves the highest CIDEr and SODAc scores of 30.10 and 5.80, respectively. The collaborative methods span both global and dense captioning, with APML Ensemble reaching 109.5 CIDEr on MSVD and ECG obtaining 7.06 METEOR on ActivityNet Captions. Overall, the metric- and method-dependent variations indicate that captioning progress cannot be represented by a single performance ranking. The mixed performance of collaborative methods further suggests that collaboration is beneficial only when its design effectively supports the specific generation objective and evaluation criterion.
| Method Type | Method | Global Video Captioning | Dense Video Captioning | ||||
|---|---|---|---|---|---|---|---|
| MSVD | MSR-VTT | VATEX | ActivityNet Captions | ||||
| CIDEr | CIDEr | CIDEr | METEOR | CIDEr | SODAc | ||
| Non-collaborative | Dense-Captioning Events [141] | – | – | – | 4.82 | 17.29 | – |
| Masked Transformer [199] | – | – | – | 9.56 | – | – | |
| PDVC [200] | – | – | – | 7.50 | 25.87 | 5.26 | |
| Enc–Dec + Local + Global [201] | 51.7 | – | – | – | – | – | |
| SwinBERT [202] | 120.6 | 53.8 | 73.0 | – | – | – | |
| Vid2Seq [161] | 146.2 | 64.6 | – | 8.50 | 30.10 | 5.80 | |
| Collaborative | Multiple Networks Consensus + OracleNet [203] | – | 53.8 | – | – | – | – |
| Weakly Supervised Dense Event Captioning [204] | – | – | – | 6.30 | 18.77 | – | |
| ECG [205] | – | – | – | 7.06 | 14.25 | – | |
| RETTA [206] | 49.8 | 24.3 | 23.8 | – | – | – | |
| APML Ensemble [207] | 109.5 | 52.7 | – | – | – | – | |
- Reasoning and Decision-making. Reasoning-oriented tasks show the largest apparent gains but also the greatest sensitivity to evaluation protocols. Table 8 summarizes representative results from non-collaborative and collaborative methods across video question answering, situated reasoning, and long-context multimodal understanding. On STAR, the reported accuracies range from 33.3 for HCRN and 36.8 for ClipBERT to 44.9 for TraveLER, although the latter is obtained under zero-shot evaluation. On Video-MME, VILA-1.5-34B reports 59.0 overall accuracy, while the collaborative methods Video-EM, A4VL, LVAgent, and VideoChat-M1 report 62.0, 77.2, 81.7, and 83.2, respectively; however, the subtitle setting is not explicitly specified for VideoChat-M1. MSR-VTT-QA also requires caution because the available results are reported under different methodological settings, and the table does not support a direct ranking across all methods. More broadly, these results do not isolate whether an improvement arises from visual representation, language-model scale, prompting, memory, or collaboration among functional units. Nevertheless, the stronger results of collaborative methods on Video-MME suggest that explicit coordination may be particularly beneficial for long-context reasoning, where complementary perception, memory, and decision processes must be integrated. This apparent advantage should, however, be interpreted cautiously because the non-collaborative evidence in the current comparison remains limited.
- Cross-benchmark Comparability. Reliable comparison is restricted to results obtained on the same dataset with the same metric and evaluation protocol. Even within an individual column, differences in pre-training data, supervision resources, input modalities, model scale, inference configurations, and collaboration mechanisms can confound attribution. This limitation is particularly evident for reasoning-oriented tasks, where the reported results include both conventional evaluations and zero-shot settings, and the subtitle condition is not uniformly specified. The four tables should therefore be read as evidence of task-specific capability trends rather than as a global leaderboard of video understanding systems or a uniform ranking of collaborative versus non-collaborative approaches.
- Implications for Collaborative Video Understanding. Across the four task regimes, performance gains appear most consistent on standardized recognition benchmarks, whereas grounding, generation, and reasoning exhibit stronger dataset, metric, and protocol dependence. The results also document a broad transition from task-specific architectures to large-scale pre-trained and foundation models, alongside an emerging body of collaborative approaches. However, the quantitative evidence does not establish a corresponding improvement attributable to collaboration across tasks. Existing metrics primarily evaluate final task outputs and do not isolate the contributions of functional-unit selection, inter-unit communication, shared-state maintenance, or adaptive execution. Consequently, Table 5, Table 6, Table 7 and Table 8 provide evidence of task-level capability evolution and the emerging quantitative presence of collaborative methods rather than direct evidence of collaboration effectiveness. Collaboration-centric evaluation should complement outcome metrics with matched-cost single-model baselines, unit-level ablations, communication and routing efficiency, memory fidelity, uncertainty calibration, and robustness to component failures.
| Method Type | Method | Paradigm | Video Question Answering | Situated Reasoning | Long-Context Multimodal Understanding |
|---|---|---|---|---|---|
| MSR-VTT-QA Standard Accuracy | STAR Accuracy | Video-MME Overall Accuracy | |||
| Non-collaborative | VILA-1.5-34B [208] | Video-capable VLM | – | – | 59.0 |
| HME [209] | Heterogeneous multimodal attention | 33.0 | – | – | |
| HCRN [210] | Conditional relation network | 35.6 | 33.3 † | – | |
| ClipBERT [211] | Sparse-sampling Transformer | 37.4 | 36.8 † | – | |
| HQGA [212] | Query-conditioned graph hierarchy | 38.6 | – | – | |
| MERLOT [213] | Temporal multimodal Transformer | 43.1 | – | – | |
| VIOLET [214] | Video-language Transformer | 43.9 | – | – | |
| Collaborative | Video-EM [69] | Multi-model tool orchestration with episodic memory | – | – | 62.0 |
| A4VL [215] | Multi-agent perception-action alliance | – | – | 77.2 | |
| LVAgent [169] | Multi-round MLLM collaboration | – | – | 81.7 | |
| VideoChat-M1 [83] | Multi-agent reinforcement learning | – | – | 83.2 * | |
| TraveLER [18] | Modular multi-LMM agents | – | 44.9 ‡ | – | |
| RADI-P [216] | Training-stage teacher–student ranking distillation | 48.2 | – | – |
5. Challenges and Future Directions
5.1. Adaptive Task Decomposition and Functional Units Orchestration
5.2. Efficient and Semantically Aligned Inter-Unit Communication
5.3. Persistent and Temporally Consistent Shared Memory
5.4. Uncertainty-Aware Collaboration and Error Containment
5.5. Evidence-Grounded and Verifiable Collaborative Reasoning
5.6. Collaboration-Centric Evaluation and Scalable Deployment
6. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Ke, S.R.; Hoang, L.; Lee, Y.J.; Hwang, J.N.; Yoo, J.H.; Choi, K.H. A Review on Video-Based Human Activity Recognition. Computers 2013, 2, 88–131. [Google Scholar] [CrossRef] [Scilit]
- AlShami, A.K.; Rabinowitz, R.; Lam, K.N.; Shleibik, Y.; Mersha, M.; Boult, T.E.; Kalita, J. SMART-vision: Survey of modern action recognition techniques in vision. Multimed. Tools Appl. 2025, 84, 32705–32776. [Google Scholar] [CrossRef] [Scilit]
- Jyothi, H.; Komala, M.; Mallikarjunaswamy, S. A Comprehensive Survey on Technologies in Video-based Event Detection and Recognition Using Machine Learning and Deep Learning Techniques. In 2024 Second International Conference on Networks, Multimedia and Information Technology (NMITCON); IEEE: New York, NY, USA, 2024. [Google Scholar]
- Sargano, A.B.; Angelov, P.; Habib, Z. A Comprehensive Review on Handcrafted and Learning-Based Action Representation Approaches for Human Activity Recognition. Appl. Sci. 2017, 7, 110. [Google Scholar] [CrossRef] [Scilit]
- Xiao, J.; Huang, N.; Qin, H.; Li, D.; Li, Y.; Zhu, F.; Tao, Z.; Yu, J.; Lin, L.; Chua, T.S.; et al. VideoQA in the Era of LLMs: An Empirical Study. Int. J. Comput. Vis. 2025, 133, 3970–3993. [Google Scholar] [CrossRef] [Scilit]
- Jenni, S.; Jin, H. Time-Equivariant Contrastive Video Representation Learning. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Montreal, QC, Canada, 10–17 October 2021; pp. 9950–9960. [Google Scholar]
- Wang, H.; Kläser, A.; Schmid, C.; Liu, C.L. Dense trajectories and motion boundary descriptors for action recognition. Int. J. Comput. Vis. 2013, 103, 60–79. [Google Scholar] [CrossRef] [Scilit]
- Wang, H.; Schmid, C. Action Recognition with Improved Trajectories. In Proceedings of the 2013 IEEE International Conference on Computer Vision, Sydney, NSW, Australia, 1–8 December 2013; pp. 3551–3558. [Google Scholar] [CrossRef] [Scilit]
- Sanchez, J.; Perronnin, F.; Mensink, T.; Verbeek, J. Image Classification with the Fisher Vector: Theory and Practice. Int. J. Comput. Vis. 2013, 105, 222–245. [Google Scholar] [CrossRef] [Scilit]
- Carreira, J.; Zisserman, A. Quo vadis, action recognition? A new model and the kinetics dataset. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA, 21–26 July 2017; pp. 6299–6308. [Google Scholar]
- Feichtenhofer, C.; Fan, H.; Malik, J.; He, K. Slowfast networks for video recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Seoul, Republic of Korea, 27 October–2 November 2019; pp. 6202–6211. [Google Scholar]
- Bertasius, G.; Wang, H.; Torresani, L. Is Space-Time Attention All You Need for Video Understanding? arXiv 2021, arXiv:2102.05095. [Google Scholar]
- Arnab, A.; Dehghani, M.; Heigold, G.; Sun, C.; Lučić, M.; Schmid, C. ViViT: A Video Vision Transformer. arXiv 2021, arXiv:2103.15691. [Google Scholar]
- Maaz, M.; Rasheed, H.; Khan, S.; Khan, F. Video-chatgpt: Towards detailed video understanding via large vision and language models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics, Volume 1: Long Papers, Bangkok, Thailand, 11–16 August 2024; pp. 12585–12602. [Google Scholar]
- Lin, B.; Ye, Y.; Zhu, B.; Cui, J.; Ning, M.; Jin, P.; Yuan, L. Video-LLaVA: Learning United Visual Representation by Alignment Before Projection. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Miami, FL, USA; Al-Onaizan, Y., Bansal, M., Chen, Y.N., Eds.; Association for Computational Linguistics: Kerrville, TX, USA, 2024; pp. 5971–5984. [Google Scholar] [CrossRef] [Scilit]
- Fan, Y.; Ma, X.; Wu, R.; Du, Y.; Li, J.; Gao, Z.; Li, Q. Videoagent: A memory-augmented multimodal agent for video understanding. In Proceedings of the European Conference on Computer Vision; Springer: Berlin/Heidelberg, Germany, 2024; pp. 75–92. [Google Scholar]
- Wang, X.; Zhang, Y.; Zohar, O.; Yeung-Levy, S. Videoagent: Long-form video understanding with large language model as agent. In Proceedings of the European Conference on Computer Vision; Springer: Berlin/Heidelberg, Germany, 2024; pp. 58–76. [Google Scholar]
- Shang, C.; You, A.; Subramanian, S.; Darrell, T.; Herzig, R. Traveler: A modular multi-lmm agent framework for video question-answering. arXiv 2024, arXiv:2404.01476. [Google Scholar]
- Zhang, L.; Zhao, T.; Ying, H.; Ma, Y.; Lee, K. Omagent: A multi-modal agent framework for complex video understanding with task divide-and-conquer. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Miami, FL, USA, 12–16 November 2024; pp. 10031–10045. [Google Scholar]
- Sleeman, W.C., IV; Kapoor, R.; Ghosh, P. Multimodal classification: Current landscape, taxonomy and future directions. ACM Comput. Surv. 2022, 55, 1–31. [Google Scholar] [CrossRef] [Scilit]
- Zeng, Y.; Zhang, X.; Li, H.; Wang, J.; Zhang, J.; Zhou, W. X2-VLM: All-in-One Pre-Trained Model for Vision-Language Tasks. IEEE Trans. Pattern Anal. Mach. Intell. 2023, 46, 3156–3168. [Google Scholar] [CrossRef] [Scilit]
- Liu, Z.; He, Y.; Wang, W.; Wang, W.; Wang, Y.; Chen, S.; Zhang, Q.; Lai, Z.; Yang, Y.; Li, Q.; et al. InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language. arXiv 2023, arXiv:2305.05662. [Google Scholar]
- Ding, G.; Sener, F.; Yao, A. Temporal Action Segmentation: An Analysis of Modern Techniques. IEEE Trans. Pattern Anal. Mach. Intell. 2024, 46, 1011–1030. [Google Scholar] [CrossRef] [Scilit]
- Selva, J.; Johansen, A.S.; Escalera, S.; Nasrollahi, K.; Moeslund, T.B.; Clapés, A. Video transformers: A survey. IEEE Trans. Pattern Anal. Mach. Intell. 2023, 45, 12922–12943. [Google Scholar] [CrossRef] [Scilit]
- Schiappa, M.C.; Rawat, Y.S.; Shah, M. Self-supervised learning for videos: A survey. ACM Comput. Surv. 2023, 55, 1–37. [Google Scholar] [CrossRef] [Scilit]
- Tang, Y.; Bi, J.; Xu, S.; Song, L.; Liang, S.; Wang, T.; Zhang, D.; An, J.; Lin, J.; Zhu, R.; et al. Video understanding with large language models: A survey. IEEE Trans. Circuits Syst. Video Technol. 2025, 36, 1355–1376. [Google Scholar] [CrossRef] [Scilit]
- Shabaninia, E.; Nezamabadi-pour, H.; Shafizadegan, F. Multimodal action recognition: A comprehensive survey on temporal modeling. Multimed. Tools Appl. 2024, 83, 59439–59489. [Google Scholar] [CrossRef] [Scilit]
- Chen, F.L.; Zhang, D.Z.; Han, M.L.; Chen, X.Y.; Shi, J.; Xu, S.; Xu, B. Vlp: A survey on vision-language pre-training. Mach. Intell. Res. 2023, 20, 38–56. [Google Scholar] [CrossRef] [Scilit]
- Xie, J.; Chen, Z.; Zhang, R.; Li, G. Large multimodal agents: A survey. Vis. Intell. 2025, 3, 24. [Google Scholar] [CrossRef] [Scilit]
- Durante, Z.; Huang, Q.; Wake, N.; Gong, R.; Park, J.S.; Sarkar, B.; Taori, R.; Noda, Y.; Terzopoulos, D.; Choi, Y.; et al. Agent AI: Surveying the Horizons of Multimodal Interaction. arXiv 2024, arXiv:2401.03568. [Google Scholar]
- Zhao, F.; Zhang, C.; Geng, B. Deep multimodal data fusion. ACM Comput. Surv. 2024, 56, 1–36. [Google Scholar] [CrossRef] [Scilit]
- Wang, M.; Xing, J.; Su, J.; Chen, J.; Liu, Y. Learning spatiotemporal and motion features in a unified 2d network for action recognition. IEEE Trans. Pattern Anal. Mach. Intell. 2022, 45, 3347–3362. [Google Scholar]
- Zhai, Y.; Wang, L.; Tang, W.; Zhang, Q.; Zheng, N.; Doermann, D.; Yuan, J.; Hua, G. Adaptive two-stream consensus network for weakly-supervised temporal action localization. IEEE Trans. Pattern Anal. Mach. Intell. 2022, 45, 4136–4151. [Google Scholar] [CrossRef] [Scilit]
- Li, J.; Tang, S.; Zhu, L.; Zhang, W.; Yang, Y.; Chua, T.S.; Wu, F.; Zhuang, Y. Variational cross-graph reasoning and adaptive structured semantics learning for compositional temporal grounding. IEEE Trans. Pattern Anal. Mach. Intell. 2023, 45, 12601–12617. [Google Scholar] [CrossRef] [Scilit]
- Li, G.; Ye, H.; Qi, Y.; Wang, S.; Qing, L.; Huang, Q.; Yang, M.H. Learning hierarchical modular networks for video captioning. IEEE Trans. Pattern Anal. Mach. Intell. 2023, 46, 1049–1064. [Google Scholar] [CrossRef] [Scilit]
- Bai, Z.; Wang, R.; Gao, D.; Chen, X. Event graph guided compositional spatial–temporal reasoning for video question answering. IEEE Trans. Image Process. 2024, 33, 1109–1121. [Google Scholar] [CrossRef] [Scilit]
- Huang, Q.; Xiong, Y.; Rao, A.; Wang, J.; Lin, D. Movienet: A holistic dataset for movie understanding. In Proceedings of the European Conference on Computer Vision; Springer: Berlin/Heidelberg, Germany, 2020; pp. 709–727. [Google Scholar]
- Geng, T.; Wang, T.; Duan, J.; Zhang, Y.; Guan, W.; Zheng, F.; Shao, L. UniAV: Unified Audio-Visual Perception for Multi-Task Video Event Localization. IEEE Trans. Pattern Anal. Mach. Intell. 2025, 47, 10280–10294. [Google Scholar] [CrossRef] [Scilit]
- Zhang, H.; Sun, A.; Jing, W.; Zhou, J.T. Temporal sentence grounding in videos: A survey and future directions. IEEE Trans. Pattern Anal. Mach. Intell. 2023, 45, 10443–10465. [Google Scholar] [CrossRef] [Scilit]
- Vahdani, E.; Tian, Y. Deep learning-based action detection in untrimmed videos: A survey. IEEE Trans. Pattern Anal. Mach. Intell. 2022, 45, 4302–4320. [Google Scholar] [CrossRef] [Scilit]
- Abdar, M.; Kollati, M.; Kuraparthi, S.; Pourpanah, F.; McDuff, D.; Ghavamzadeh, M.; Yan, S.; Mohamed, A.; Khosravi, A.; Cambria, E.; et al. A review of deep learning for video captioning. IEEE Trans. Pattern Anal. Mach. Intell. 2024, Early Access. [Google Scholar]
- Cong, Y.; Liao, W.; Ackermann, H.; Rosenhahn, B.; Yang, M.Y. Spatial-temporal transformer for dynamic scene graph generation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Montreal, QC, Canada, 10–17 October 2021; pp. 16372–16382. [Google Scholar]
- Stergiou, A.; Poppe, R. About time: Advances, challenges, and outlooks of action understanding. Int. J. Comput. Vis. 2025, 133, 6251–6315. [Google Scholar] [CrossRef] [Scilit]
- Hutchinson, M.S.; Gadepally, V.N. Video action understanding. IEEE Access 2021, 9, 134611–134637. [Google Scholar] [CrossRef] [Scilit]
- Karim, M.; Khalid, S.; Aleryani, A.; Khan, J.; Ullah, I.; Ali, Z. Human action recognition systems: A review of the trends and state-of-the-art. IEEE Access 2024, 12, 36372–36390. [Google Scholar] [CrossRef] [Scilit]
- Huang, T.E.; Liu, Y.; Van Gool, L.; Yu, F. Video task decathlon: Unifying image and video tasks in autonomous driving. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Paris, France, 1–6 October 2023; pp. 8647–8657. [Google Scholar]
- Nguyen, T.; Bin, Y.; Xiao, J.; Qu, L.; Li, Y.; Wu, J.Z.; Nguyen, C.D.; Ng, S.K.; Tuan, L.A. Video-language understanding: A survey from model architecture, model training, and data perspectives. In Proceedings of the Findings of the Association for Computational Linguistics: ACL 2024, Bangkok, Thailand, 11–16 August 2024; pp. 3636–3657. [Google Scholar]
- Zhang, Z.; Xu, D.; Ouyang, W.; Zhou, L. Dense video captioning using graph-based sentence summarization. IEEE Trans. Multimed. 2020, 23, 1799–1810. [Google Scholar] [CrossRef] [Scilit]
- Song, X.; Chen, J.; Wu, Z.; Jiang, Y.G. Spatial-temporal graphs for cross-modal text2video retrieval. IEEE Trans. Multimed. 2021, 24, 2914–2923. [Google Scholar] [CrossRef] [Scilit]
- Wray, M.; Doughty, H.; Damen, D. On semantic similarity in video retrieval. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 19–25 June 2021; pp. 3650–3660. [Google Scholar]
- Aafaq, N.; Mian, A.; Liu, W.; Gilani, S.Z.; Shah, M. Video description: A survey of methods, datasets, and evaluation metrics. ACM Comput. Surv. (CSUR) 2019, 52, 1–37. [Google Scholar]
- Qasim, I.; Horsch, A.; Prasad, D. Dense video captioning: A survey of techniques, datasets and evaluation protocols. ACM Comput. Surv. 2025, 57, 1–36. [Google Scholar] [CrossRef] [Scilit]
- Huang, J.H.; Huck Yang, C.H.; Chen, P.Y.; Chen, M.H.; Worring, M. Conditional Modeling-Based Automatic Video Summarization. ACM Trans. Multimed. Comput. Commun. Appl. 2026, 22, 1–21. [Google Scholar] [CrossRef] [Scilit]
- Zhong, Y.; Ji, W.; Xiao, J.; Li, Y.; Deng, W.; Chua, T.S. Video question answering: Datasets, algorithms and challenges. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, Abu Dhabi, United Arab Emirates, 7–11 December 2022; pp. 6439–6455. [Google Scholar]
- Liang, T.; Li, L.; Hu, J.F.; Yu, X.; Zheng, W.S.; Lai, J. Rethinking Temporal Context in Video-QA: A Comprehensive Study of Single-Frame Static Bias. IEEE Trans. Multimed. 2025, 27, 5077–5091. [Google Scholar] [CrossRef] [Scilit]
- Li, J.; Niu, L.; Zhang, L. From representation to reasoning: Towards both evidence and commonsense reasoning for video question-answering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 18–24 June 2022; pp. 21273–21282. [Google Scholar]
- Sun, Y.Z.; Sun, H.L.; Ma, J.C.; Zhang, P.; Huang, X.Y. Multimodal agent AI: A survey of recent advances and future directions. J. Comput. Sci. Technol. 2025, 40, 1046–1063. [Google Scholar] [CrossRef] [Scilit]
- Lan, X.; Yuan, Y.; Wang, X.; Wang, Z.; Zhu, W. A survey on temporal sentence grounding in videos. ACM Trans. Multimed. Comput. Commun. Appl. 2023, 19, 1–33. [Google Scholar] [CrossRef] [Scilit]
- Hui, T.; Liu, S.; Ding, Z.; Huang, S.; Li, G.; Wang, W.; Liu, L.; Han, J. Language-aware spatial-temporal collaboration for referring video segmentation. IEEE Trans. Pattern Anal. Mach. Intell. 2023, 45, 8646–8659. [Google Scholar] [CrossRef] [Scilit]
- Kwon, H.; Kim, M.; Kwak, S.; Cho, M. Learning self-similarity in space and time as generalized motion for video action recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Montreal, QC, Canada, 10–17 October 2021; pp. 13065–13075. [Google Scholar]
- Liu, X.; Pintea, S.L.; Nejadasl, F.K.; Booij, O.; Van Gemert, J.C. No frame left behind: Full video action recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 19–25 June 2021; pp. 14892–14901. [Google Scholar]
- Cao, Q.; Huang, H. VSRN: Visual-semantic relation network for video visual relation inference. IEEE Trans. Circuits Syst. Video Technol. 2021, 32, 768–777. [Google Scholar] [CrossRef] [Scilit]
- Zhao, B.; Gong, M.; Li, X. Audiovisual video summarization. IEEE Trans. Neural Netw. Learn. Syst. 2021, 34, 5181–5188. [Google Scholar] [CrossRef] [Scilit]
- Yu, T.; Yu, J.; Yu, Z.; Huang, Q.; Tian, Q. Long-term video question answering via multimodal hierarchical memory attentive networks. IEEE Trans. Circuits Syst. Video Technol. 2020, 31, 931–944. [Google Scholar] [CrossRef] [Scilit]
- Li, G.; Wang, X.; Zhu, W. AV-Unified: A Unified Framework for Audio-visual Scene Understanding. IEEE Trans. Multimed. 2026, Early Access. [Google Scholar]
- Zhao, H.; Ji, G.P.; Yan, R.; Xiong, H.; Li, Z. VideoExpert: Augmented LLM for Temporal-Sensitive Video Understanding. IEEE Trans. Circuits Syst. Video Technol. 2026, 36, 5893–5908. [Google Scholar] [CrossRef] [Scilit]
- Madan, N.; Møgelmose, A.; Modi, R.; Rawat, Y.S.; Moeslund, T.B. Foundation models for video understanding: A survey. arXiv 2024, arXiv:2405.03770. [Google Scholar]
- Xu, J.; Lan, C.; Xie, W.; Chen, X.; Lu, Y. Long video understanding with learnable retrieval in video-language models. IEEE Trans. Multimed. 2026, 28, 5254–5261. [Google Scholar] [CrossRef] [Scilit]
- Wang, Q.; Wang, Y.; Chen, H.; Wang, S.; Du, J.; Lee, C.H. Video Segmentation and Tokenization for Model-Based Video Scene Classification. IEEE Trans. Multimed. 2025, 27, 6489–6502. [Google Scholar] [CrossRef] [Scilit]
- Yang, Z.; Chen, G.; Li, X.; Wang, W.; Yang, Y. DoraemonGPT: Toward Understanding Dynamic Scenes with Large Language Models (Exemplified as A Video Agent). In Proceedings of the Forty-first International Conference on Machine Learning, Vienna, Austria, 21–27 July 2024; pp. 55976–55997. [Google Scholar]
- Li, J.; Liao, Z.; Xiao, F.; Li, T.; Zhang, Q.; Zhao, H.; Niu, L.; Chen, G.; Zhang, L.; Jiang, C. Parse, Align and Aggregate: Graph-driven Compositional Reasoning for Video Question Answering. IEEE Trans. Pattern Anal. Mach. Intell. 2026, 48, 5586–5603. [Google Scholar] [CrossRef] [Scilit]
- Li, L.; Jin, T.; Lin, W.; Jiang, H.; Pan, W.; Wang, J.; Xiao, S.; Xia, Y.; Jiang, W.; Zhao, Z. Multi-granularity relational attention network for audio-visual question answering. IEEE Trans. Circuits Syst. Video Technol. 2023, 34, 7080–7094. [Google Scholar] [CrossRef] [Scilit]
- Pei, B.; Huang, Y.; Chen, G.; Xu, J.; Wang, Y.; Wang, L.; Lu, T.; Qiao, Y.; Wu, F. Guiding audio-visual question answering with collective question reasoning. Int. J. Comput. Vis. 2025, 133, 6912–6929. [Google Scholar] [CrossRef] [Scilit]
- Chen, S.; Xu, Q.; Ma, Y.; Qiao, Y.; Wang, Y. Attentive snippet prompting for video retrieval. IEEE Trans. Multimed. 2023, 26, 4348–4359. [Google Scholar] [CrossRef] [Scilit]
- Chang, Z.; Zhang, X.; Wang, S.; Ma, S.; Gao, W. Stam: A spatiotemporal attention based memory for video prediction. IEEE Trans. Multimed. 2022, 25, 2354–2367. [Google Scholar] [CrossRef] [Scilit]
- Miao, B.; Bennamoun, M.; Gao, Y.; Shah, M.; Mian, A. Temporally consistent referring video object segmentation with hybrid memory. IEEE Trans. Circuits Syst. Video Technol. 2024, 34, 11373–11385. [Google Scholar] [CrossRef] [Scilit]
- Liu, Y.; Albanie, S.; Nagrani, A.; Zisserman, A. Use what you have: Video retrieval using representations from collaborative experts. arXiv 2019, arXiv:1907.13487. [Google Scholar]
- Luo, Y.; Zheng, X.; Li, G.; Yin, S.; Lin, H.; Fu, C.; Huang, J.; Ji, J.; Chao, F.; Luo, J.; et al. Video-rag: Visually-aligned retrieval-augmented long video comprehension. Adv. Neural Inf. Process. Syst. 2025, 38, 168008–168033. [Google Scholar]
- Kugo, N.; Li, X.; Li, Z.; Gupta, A.; Khatua, A.; Jain, N.; Patel, C.; Kyuragi, Y.; Ishii, Y.; Tanabiki, M.; et al. VideoMultiAgents: A Multi-Agent Framework for Video Question Answering. arXiv 2025, arXiv:2504.20091. [Google Scholar]
- Chowdhury, S.; Elmoghany, M.; Abeysinghe, Y.; Fei, J.; Nag, S.; Khan, S.; Elhoseiny, M.; Manocha, D. MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks. In Proceedings of the Advances in Neural Information Processing Systems, San Diego, CA, USA, 2–7 December 2025; pp. 49255–49291. [Google Scholar]
- Min, J.; Buch, S.; Nagrani, A.; Cho, M.; Schmid, C. Morevqa: Exploring modular reasoning models for video question answering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 17–21 June 2024; pp. 13235–13245. [Google Scholar]
- Ma, Z.; Gou, C.; Shi, H.; Sun, B.; Li, S.; Rezatofighi, H.; Cai, J. Drvideo: Document retrieval based long video understanding. In Proceedings of the Computer Vision and Pattern Recognition Conference, Nashville, TN, USA, 11–15 June 2025; pp. 18936–18946. [Google Scholar]
- Chen, B.; Wang, Z.; Yue, Z.; Yan, K.; Yu, C.; Huang, Y.; Liu, Z.; Wen, Y.; Chen, X.; Liu, Y.; et al. Videochat-m1: Collaborative policy planning for video understanding via multi-agent reinforcement learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Denver, CO, USA, 3–7 June 2026; pp. 33772–33783. [Google Scholar]
- Zhu, Y.; Zhao, J.; Zhao, J.; Mao, X.; Zhao, B. HiCrew: Hierarchical Reasoning for Long-Form Video Understanding via Question-Aware Multi-Agent Collaboration. arXiv 2026, arXiv:2604.21444. [Google Scholar]
- Yang, A.; Miech, A.; Sivic, J.; Laptev, I.; Schmid, C. Learning to answer visual questions from web videos. IEEE Trans. Pattern Anal. Mach. Intell. 2022, 47, 3202–3218. [Google Scholar] [CrossRef] [Scilit]
- Jiang, Y.; Yin, J. Clip-powered tass: Target-aware single-stream network for audio-visual question answering. Int. J. Comput. Vis. 2025, 133, 2581–2598. [Google Scholar] [CrossRef] [Scilit]
- Wang, R.; Luo, Y.; Zhang, F.; Liu, M.; Luo, X. HSSHG: Heuristic Semantics-Constrained Spatio-Temporal Heterogeneous Graph for VideoQA. IEEE Trans. Multimed. 2024, 26, 11176–11190. [Google Scholar] [CrossRef] [Scilit]
- Cheng, Y.; Fan, H.; Lin, D.; Sun, Y.; Kankanhalli, M.; Lim, J.H. Keyword-aware relative spatio-temporal graph networks for video question answering. IEEE Trans. Multimed. 2023, 26, 6131–6141. [Google Scholar] [CrossRef] [Scilit]
- Xu, W.; Yu, J.; Miao, Z.; Wan, L.; Tian, Y.; Ji, Q. Deep reinforcement polishing network for video captioning. IEEE Trans. Multimed. 2020, 23, 1772–1784. [Google Scholar] [CrossRef] [Scilit]
- Liu, F.; Liu, J.; Wang, W.; Lu, H. Hair: Hierarchical visual-semantic relational reasoning for video question answering. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Montreal, QC, Canada, 10–17 October 2021; pp. 1698–1707. [Google Scholar]
- Lu, J.; You, S.; Bao, B.K. Question Understanding and Temporality Guiding for Video Question Answering. IEEE Trans. Multimed. 2026, 28, 2772–2783. [Google Scholar] [CrossRef] [Scilit]
- Xu, W.; Miao, Z.; Yu, J.; Tian, Y.; Wan, L.; Ji, Q. Bridging video and text: A two-step polishing transformer for video captioning. IEEE Trans. Circuits Syst. Video Technol. 2022, 32, 6293–6307. [Google Scholar] [CrossRef] [Scilit]
- Qian, T.; Cui, R.; Chen, J.; Peng, P.; Guo, X.; Jiang, Y.G. Locate before answering: Answer guided question localization for video question answering. IEEE Trans. Multimed. 2023, 26, 4554–4563. [Google Scholar] [CrossRef] [Scilit]
- Zhou, S.; Xiao, J.; Yang, X.; Song, P.; Guo, D.; Yao, A.; Wang, M.; Chua, T.S. Scene-text grounding for text-based video question answering. IEEE Trans. Multimed. 2025, 28, 1417–1430. [Google Scholar] [CrossRef] [Scilit]
- Xiao, J.; Zhou, P.; Yao, A.; Li, Y.; Hong, R.; Yan, S.; Chua, T.S. Contrastive video question answering via video graph transformer. IEEE Trans. Pattern Anal. Mach. Intell. 2023, 45, 13265–13280. [Google Scholar] [CrossRef] [Scilit]
- Wu, K.; Li, X.; Li, X.; Zuo, K.; Lv, Z. AVQACL++: Toward a robust framework and benchmark for audio-visual question answering continual learning. IEEE Trans. Circuits Syst. Video Technol. 2026, 36, 8232–8245. [Google Scholar] [CrossRef] [Scilit]
- Gao, J.; Sun, X.; Ghanem, B.; Zhou, X.; Ge, S. Efficient video grounding with which-where reading comprehension. IEEE Trans. Circuits Syst. Video Technol. 2022, 32, 6900–6913. [Google Scholar] [CrossRef] [Scilit]
- Su, H.T.; Chang, C.H.; Shen, P.W.; Wang, Y.S.; Chang, Y.L.; Chang, Y.C.; Cheng, P.J.; Hsu, W.H. End-to-end video question-answer generation with generator-pretester network. IEEE Trans. Circuits Syst. Video Technol. 2021, 31, 4497–4507. [Google Scholar] [CrossRef] [Scilit]
- You, Z.; Wen, Z.; Chen, Y.; Li, X.; Zeng, R.; Wang, Y.; Tan, M. Toward long video understanding via fine-detailed video story generation. IEEE Trans. Circuits Syst. Video Technol. 2024, 35, 4592–4607. [Google Scholar] [CrossRef] [Scilit]
- Sun, X.; Dai, Y.; Wang, Y.; Ma, W.; Lin, X. Video question answering via traffic knowledge database and question classification. Multimed. Syst. 2024, 30, 14. [Google Scholar] [CrossRef] [Scilit]
- Lee, J.; Chang, J.; Lee, D.; Choi, J. CA2ST: Cross-Attention in Audio, Space, and Time for Holistic Video Recognition. IEEE Trans. Pattern Anal. Mach. Intell. 2025, 48, 2803–2819. [Google Scholar] [CrossRef] [Scilit]
- Song, P.; Zhang, L.; Lan, L.; Chen, W.; Guo, D.; Yang, X.; Wang, M. Towards efficient partially relevant video retrieval with active moment discovering. IEEE Trans. Multimed. 2025, 27, 6740–6751. [Google Scholar] [CrossRef] [Scilit]
- Li, X.; He, H.; Yang, Y.; Ding, H.; Yang, K.; Cheng, G.; Tong, Y.; Tao, D. Improving video instance segmentation via temporal pyramid routing. IEEE Trans. Pattern Anal. Mach. Intell. 2022, 45, 6594–6601. [Google Scholar] [CrossRef] [Scilit]
- Lai, C.; Ge, W.; Xue, X. Cross-Modal Complementary Learning and Template-Based Reasoning Chains for Future Event Prediction in Videos. IEEE Trans. Multimed. 2025, 27, 7497–7509. [Google Scholar] [CrossRef] [Scilit]
- Chen, L.; Lu, J.; Song, Z.; Zhou, J. Ambiguousness-aware state evolution for action prediction. IEEE Trans. Circuits Syst. Video Technol. 2022, 32, 6058–6072. [Google Scholar] [CrossRef] [Scilit]
- Xie, Z.; Luo, J.; Wu, K.; Kan, Z.; Guo, D. Learning Confidence-aware Prototypes for Weakly-supervised Video Anomaly Detection. IEEE Trans. Circuits Syst. Video Technol. 2025, 36, 5714–5728. [Google Scholar] [CrossRef] [Scilit]
- Huo, S.; Zhou, Y.; Wang, R.; Xiang, W.; Kung, S.Y. Semantic relevance learning for video-query based video moment retrieval. IEEE Trans. Multimed. 2023, 25, 9290–9301. [Google Scholar] [CrossRef] [Scilit]
- Yang, Z.; Chen, D.; Yu, X.; Shen, M.; Gan, C. Vca: Video curious agent for long video understanding. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Honolulu, HI, USA, 19–25 October 2025; pp. 20168–20179. [Google Scholar]
- Ye, J.; Wang, Z.; Sun, H.; Chandrasegaran, K.; Durante, Z.; Eyzaguirre, C.; Bisk, Y.; Niebles, J.C.; Adeli, E.; Li, F.F.; et al. Re-thinking temporal search for long-form video understanding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 11–15 June 2025; pp. 8579–8591. [Google Scholar]
- Kahatapitiya, K.; Ranasinghe, K.; Park, J.; Ryoo, M.S. Language repository for long video understanding. Proc. Find. Assoc. Comput. Linguist. ACL 2025, 2025, 5627–5646. [Google Scholar] [CrossRef] [Scilit]
- Wang, Z.; Yu, S.; Stengel-Eskin, E.; Yoon, J.; Cheng, F.; Bertasius, G.; Bansal, M. Videotree: Adaptive tree-based video representation for llm reasoning on long videos. In Proceedings of the Computer Vision and Pattern Recognition Conference, Nashville, TN, USA, 11–15 June 2025; pp. 3272–3283. [Google Scholar]
- Xue, Z.; Zhang, J.; Xie, X.; Cai, Y.; Liu, Y.; Li, X.; Tao, D. Adavideorag: Omni-contextual adaptive retrieval-augmented efficient long video understanding. arXiv 2025, arXiv:2506.13589. [Google Scholar]
- Yeo, W.; Kim, K.; Yoon, J.; Hwang, S.J. Worldmm: Dynamic multimodal memory agent for long video reasoning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Denver, CO, USA, 3–7 June 2026; pp. 25599–25609. [Google Scholar]
- Zhi, Z.; Wu, Q.; Li, W.; Li, Y.; Shao, K.; Zhou, K. Videoagent2: Enhancing the llm-based agent system for long-form video understanding by uncertainty-aware cot. arXiv 2025, arXiv:2504.04471. [Google Scholar]
- Xiong, H.; Yang, Z.; Yu, J.; Zhuge, Y.; Zhang, L.; Zhu, J.; Lu, H. Streaming video understanding and multi-round interaction with memory-enhanced knowledge. Proc. Int. Conf. Learn. Represent. 2025, 2025, 69332–69351. [Google Scholar]
- Zhang, K.; Yang, Z.; Han, M.; Hao, H.; Zhuge, Y.; Li, C.; Li, Z.; Chang, X. Progressive online video understanding with evidence-aligned timing and transparent decisions. Proc. Int. Conf. Learn. Represent. 2026, 2026, 5092–5121. [Google Scholar]
- Hu, K.; Gao, F.; Nie, X.; Zhou, P.; Tran, S.; Neiman, T.; Wang, L.; Shah, M.; Hamid, R.; Yin, B.; et al. M-llm based video frame selection for efficient video understanding. In Proceedings of the Computer Vision and Pattern Recognition Conference, Nashville, TN, USA, 11–15 June 2025; pp. 13702–13712. [Google Scholar]
- Tang, X.; Qiu, J.; Xie, L.; Tian, Y.; Jiao, J.; Ye, Q. Adaptive keyframe sampling for long video understanding. In Proceedings of the Computer Vision and Pattern Recognition Conference, Nashville, TN, USA, 11–15 June 2025; pp. 29118–29128. [Google Scholar]
- Zhang, S.; Yang, J.; Yin, J.; Luo, Z.; Luan, J. Q-frame: Query-aware frame selection and multi-resolution adaptation for video-llms. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Honolulu, HI, USA, 19–23 October 2025; pp. 22056–22065. [Google Scholar]
- Wang, Z.; Zhou, H.; Wang, S.; Li, J.; Xiong, C.; Savarese, S.; Bansal, M.; Ryoo, M.S.; Niebles, J.C. Active video perception: Iterative evidence seeking for agentic long video understanding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Denver, CO, USA, 3–7 June 2026; pp. 9088–9099. [Google Scholar]
- Diko, A.; Wang, T.; Swaileh, W.; Sun, S.; Patras, I. Rewind: Understanding long videos with instructed learnable memory. In Proceedings of the Computer Vision and Pattern Recognition Conference, Nashville, TN, USA, 11–15 June 2025; pp. 13734–13743. [Google Scholar]
- Yu, S.; Yoon, J.; Bansal, M. Crema: Generalizable and efficient video-language reasoning via multimodal modular fusion. Proc. Int. Conf. Learn. Represent. 2025, 2025, 74382–74406. [Google Scholar]
- Endo, M.; Hsu, J.; Li, J.; Wu, J. Motion question answering via modular motion programs. In Proceedings of the International Conference on Machine Learning, Honolulu, HI, USA, 23–29 July 2023; pp. 9312–9328. [Google Scholar]
- Choi, M.; Goel, H.; Omama, M.; Yang, Y.; Shah, S.; Chinchali, S. Towards neuro-symbolic video understanding. In Proceedings of the European Conference on Computer Vision; Springer: Berlin/Heidelberg, Germany, 2024; pp. 220–236. [Google Scholar]
- Liu, Y.; Qinghong Lin, K.; Chen, C.W.; Shou, M.Z. Videomind: A chain-of-lora agent for long video reasoning. arXiv 2025, arXiv:2503.13444. [Google Scholar]
- Shinn, N.; Cassano, F.; Gopinath, A.; Narasimhan, K.; Yao, S. Reflexion: Language agents with verbal reinforcement learning. Adv. Neural Inf. Process. Syst. 2023, 36, 8634–8652. [Google Scholar] [CrossRef] [Scilit]
- Du, Y.; Li, S.; Torralba, A.; Tenenbaum, J.B.; Mordatch, I. Improving Factuality and Reasoning in Language Models through Multiagent Debate. arXiv 2023, arXiv:2305.14325. [Google Scholar]
- Park, J.S.; O’Brien, J.; Cai, C.J.; Morris, M.R.; Liang, P.; Bernstein, M.S. Generative Agents: Interactive Simulacra of Human Behavior. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology, San Francisco, CA, USA, 29 October–1 November 2023. [Google Scholar] [CrossRef] [Scilit]
- Wang, Z.; Cai, S.; Liu, A.; Jin, Y.; Hou, J.; Zhang, B.; Lin, H.; He, Z.; Zheng, Z.; Yang, Y.; et al. Jarvis-1: Open-world multi-task agents with memory-augmented multimodal language models. IEEE Trans. Pattern Anal. Mach. Intell. 2024, 47, 1894–1907. [Google Scholar] [CrossRef] [Scilit]
- Szot, A.; Mazoure, B.; Attia, O.; Timofeev, A.; Agrawal, H.; Hjelm, D.; Gan, Z.; Kira, Z.; Toshev, A. From multimodal llms to generalist embodied agents: Methods and lessons. In Proceedings of the Computer Vision and Pattern Recognition Conference, Nashville, TN, USA, 11–15 June 2025; pp. 10644–10655. [Google Scholar]
- Li, K.; Wang, Y.; He, Y.; Li, Y.; Wang, Y.; Liu, Y.; Wang, Z.; Xu, J.; Chen, G.; Luo, P.; et al. Mvbench: A comprehensive multi-modal video understanding benchmark. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 17–21 June 2024; pp. 22195–22206. [Google Scholar]
- Fu, C.; Dai, Y.; Luo, Y.; Li, L.; Ren, S.; Zhang, R.; Wang, Z.; Zhou, C.; Shen, Y.; Zhang, M.; et al. Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-Modal LLMs in Video Analysis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 11–15 June 2025; pp. 24108–24118. [Google Scholar]
- Goyal, R.; Ebrahimi Kahou, S.; Michalski, V.; Materzynska, J.; Westphal, S.; Kim, H.; Haenel, V.; Fruend, I.; Yianilos, P.; Mueller-Freitag, M.; et al. The “something something” video database for learning and evaluating visual common sense. In Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy, 22–29 October 2017; pp. 5842–5850. [Google Scholar]
- Idrees, H.; Zamir, A.R.; Jiang, Y.G.; Gorban, A.; Laptev, I.; Sukthankar, R.; Shah, M. The thumos challenge on action recognition for videos “in the wild”. Comput. Vis. Image Underst. 2017, 155, 1–23. [Google Scholar] [CrossRef] [Scilit]
- Caba Heilbron, F.; Escorcia, V.; Ghanem, B.; Carlos Niebles, J. Activitynet: A large-scale video benchmark for human activity understanding. In Proceedings of the Ieee Conference on Computer Vision and Pattern Recognition, Boston, MA, USA, 7–12 June 2015; pp. 961–970. [Google Scholar]
- Gu, C.; Sun, C.; Ross, D.A.; Vondrick, C.; Pantofaru, C.; Li, Y.; Vijayanarasimhan, S.; Toderici, G.; Ricco, S.; Sukthankar, R.; et al. Ava: A video dataset of spatio-temporally localized atomic visual actions. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 18–22 June 2018; pp. 6047–6056. [Google Scholar]
- Xu, J.; Mei, T.; Yao, T.; Rui, Y. Msr-vtt: A large video description dataset for bridging video and language. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA, 26 June–1 July 2016; pp. 5288–5296. [Google Scholar]
- Anne Hendricks, L.; Wang, O.; Shechtman, E.; Sivic, J.; Darrell, T.; Russell, B. Localizing moments in video with natural language. In Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy, 22–29 October 2017; pp. 5803–5812. [Google Scholar]
- Gao, J.; Sun, C.; Yang, Z.; Nevatia, R. Tall: Temporal activity localization via language query. In Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy, 22–29 October 2017; pp. 5267–5275. [Google Scholar]
- Regneri, M.; Rohrbach, M.; Wetzel, D.; Thater, S.; Schiele, B.; Pinkal, M. Grounding action descriptions in videos. Trans. Assoc. Comput. Linguist. 2013, 1, 25–36. [Google Scholar] [CrossRef] [Scilit]
- Krishna, R.; Hata, K.; Ren, F.; Fei-Fei, L.; Carlos Niebles, J. Dense-captioning events in videos. In Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy, 22–29 October 2017; pp. 706–715. [Google Scholar]
- Lei, J.; Berg, T.L.; Bansal, M. Detecting moments and highlights in videos via natural language queries. Adv. Neural Inf. Process. Syst. 2021, 34, 11846–11858. [Google Scholar]
- Xiao, J.; Yao, A.; Li, Y.; Chua, T.S. Can I trust your answer? Visually grounded video question answering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 17–21 June 2024; pp. 13204–13214. [Google Scholar]
- Chen, D.; Dolan, W.B. Collecting highly parallel data for paraphrase evaluation. In Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies, Portland, OR, USA, 19–24 June 2011; pp. 190–200. [Google Scholar]
- Wang, X.; Wu, J.; Chen, J.; Li, L.; Wang, Y.F.; Wang, W.Y. Vatex: A large-scale, high-quality multilingual dataset for video-and-language research. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Seoul, Republic of Korea, 27 October–2 November 2019; pp. 4581–4591. [Google Scholar]
- Zhou, L.; Xu, C.; Corso, J. Towards automatic learning of procedures from web instructional videos. In Proceedings of the AAAI Conference on Artificial Intelligence, New Orleans, LA, USA, 2–7 February 2018; Volume 32. [Google Scholar]
- Song, Y.; Vallmitjana, J.; Stent, A.; Jaimes, A. Tvsum: Summarizing web videos using titles. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA, 7–12 June 2015; pp. 5179–5187. [Google Scholar]
- Jang, Y.; Song, Y.; Yu, Y.; Kim, Y.; Kim, G. Tgif-qa: Toward spatio-temporal reasoning in visual question answering. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA, 21–26 July 2017; pp. 2758–2766. [Google Scholar]
- Wang, W.; Huang, Y.; Wang, L. Long video question answering: A matching-guided attention model. Pattern Recognit. 2020, 102, 107248. [Google Scholar] [CrossRef] [Scilit]
- Xiao, J.; Shang, X.; Yao, A.; Chua, T.S. Next-qa: Next phase of question-answering to explaining temporal actions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 19–25 June 2021; pp. 9777–9786. [Google Scholar]
- Grunde-McLaughlin, M.; Krishna, R.; Agrawala, M. Agqa: A benchmark for compositional spatio-temporal reasoning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 19–25 June 2021; pp. 11287–11297. [Google Scholar]
- Wu, B.; Yu, S.; Chen, Z.; Tenenbaum, J.; Gan, C. STAR: A Benchmark for Situated Reasoning in Real-World Videos. In Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks 1, NeurIPS Datasets and Benchmarks 2021, Virtual, 6–14 December 2021. [Google Scholar]
- Mangalam, K.; Akshulakov, R.; Malik, J. Egoschema: A diagnostic benchmark for very long-form video language understanding. Adv. Neural Inf. Process. Syst. 2023, 36, 46212–46244. [Google Scholar] [CrossRef] [Scilit]
- Wu, H.; Li, D.; Chen, B.; Li, J. Longvideobench: A benchmark for long-context interleaved video-language understanding. Adv. Neural Inf. Process. Syst. 2024, 37, 28828–28857. [Google Scholar] [CrossRef] [Scilit]
- Pehlivan, S.; Laaksonen, J. Temporal teacher with masked transformers for semi-supervised action proposal generation. Mach. Vis. Appl. 2024, 35, 36. [Google Scholar] [CrossRef] [Scilit]
- Buch, S.; Eyzaguirre, C.; Gaidon, A.; Wu, J.; Fei-Fei, L.; Niebles, J.C. Revisiting the “video” in video-language understanding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 19–24 June 2022; pp. 2917–2927. [Google Scholar]
- Rohrbach, A.; Torabi, A.; Rohrbach, M.; Tandon, N.; Pal, C.; Larochelle, H.; Courville, A.; Schiele, B. Movie Description. Int. J. Comput. Vis. 2017, 123, 94–120. [Google Scholar] [CrossRef] [Scilit]
- Apostolidis, E.; Adamantidou, E.; Metsai, A.I.; Mezaris, V.; Patras, I. Video summarization using deep neural networks: A survey. Proc. IEEE 2021, 109, 1838–1863. [Google Scholar] [CrossRef] [Scilit]
- Liu, Y.; Li, S.; Liu, Y.; Wang, Y.; Ren, S.; Li, L.; Chen, S.; Sun, X.; Hou, L. Tempcompass: Do video llms really understand videos? Proc. Find. Assoc. Comput. Linguist. ACL 2024, 2024, 8731–8772. [Google Scholar] [CrossRef] [Scilit]
- Song, E.; Chai, W.; Wang, G.; Zhang, Y.; Zhou, H.; Wu, F.; Chi, H.; Guo, X.; Ye, T.; Zhang, Y.; et al. Moviechat: From dense token to sparse memory for long video understanding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 17–21 June 2024; pp. 18221–18232. [Google Scholar]
- Yang, A.; Nagrani, A.; Seo, P.H.; Miech, A.; Pont-Tuset, J.; Laptev, I.; Sivic, J.; Schmid, C. Vid2seq: Large-scale pretraining of a visual language model for dense video captioning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 18–22 June 2023; pp. 10714–10726. [Google Scholar]
- Wang, B.; Zhao, Y.; Yang, L.; Long, T.; Li, X. Temporal action localization in the deep learning era: A survey. IEEE Trans. Pattern Anal. Mach. Intell. 2023, 46, 2171–2190. [Google Scholar] [CrossRef] [Scilit]
- Wang, J.; Sun, G.; Wang, P.; Liu, D.; Dianat, S.; Rabbani, M.; Rao, R.; Tao, Z. Text is mass: Modeling as stochastic embedding for text-video retrieval. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 17–21 June 2024; pp. 16551–16560. [Google Scholar]
- Xu, H.; He, K.; Plummer, B.A.; Sigal, L.; Sclaroff, S.; Saenko, K. Multilevel language and vision integration for text-to-clip retrieval. Proc. AAAI Conf. Artif. Intell. 2019, 33, 9062–9069. [Google Scholar] [CrossRef] [Scilit]
- Rafiq, M.; Rafiq, G.; Choi, G.S. Video description: Datasets & evaluation metrics. IEEE Access 2021, 9, 121665–121685. [Google Scholar] [CrossRef] [Scilit]
- Jung, W.; Kim, J. QEVA: A Reference-Free Evaluation Metric for Narrative Video Summarization with Multimodal Question Answering. Proc. Find. Assoc. Comput. Linguist. EMNLP 2025, 2025, 24632–24642. [Google Scholar] [CrossRef] [Scilit]
- Shi, Y.; Yang, X.; Xu, H.; Yuan, C.; Li, B.; Hu, W.; Zha, Z.J. Emscore: Evaluating video captioning via coarse-grained and fine-grained embedding matching. In Proceedings of the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: New York, NY, USA, 2022; pp. 17908–17917. [Google Scholar]
- Jeoung, S.; Huybrechts, G.; Ganesh, B.; Galstyan, A.; Bodapati, S. Adaptive video understanding agent: Enhancing efficiency with dynamic frame sampling and feedback-driven reasoning. arXiv 2024, arXiv:2410.20252. [Google Scholar]
- Chen, B.; Yue, Z.; Chen, S.; Wang, Z.; Liu, Y.; Li, P.; Wang, Y. Lvagent: Long video understanding by multi-round dynamical collaboration of mllm agents. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Honolulu, HI, USA, 19–23 October 2025; pp. 20237–20246. [Google Scholar]
- Kumar, Y. VideoLLM Benchmarks and Evaluation: A Survey. arXiv 2025, arXiv:2505.03829. [Google Scholar]
- Lin, J.; Gan, C.; Han, S. Tsm: Temporal shift module for efficient video understanding. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Seoul, Republic of Korea, 27 October–2 November 2019; pp. 7083–7093. [Google Scholar]
- Feichtenhofer, C. X3d: Expanding architectures for efficient video recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Virtual, 14–19 June 2020; pp. 203–213. [Google Scholar]
- Li, K.; Wang, Y.; Gao, P.; Song, G.; Liu, Y.; Li, H.; Qiao, Y. Uniformer: Unified transformer for efficient spatiotemporal representation learning. arXiv 2022, arXiv:2201.04676. [Google Scholar]
- Liu, Z.; Ning, J.; Cao, Y.; Wei, Y.; Zhang, Z.; Lin, S.; Hu, H. Video swin transformer. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 19–24 June 2022; pp. 3202–3211. [Google Scholar]
- Li, Y.; Wu, C.Y.; Fan, H.; Mangalam, K.; Xiong, B.; Malik, J.; Feichtenhofer, C. Mvitv2: Improved multiscale vision transformers for classification and detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 19–24 June 2022; pp. 4804–4814. [Google Scholar]
- Wei, C.; Fan, H.; Xie, S.; Wu, C.Y.; Yuille, A.; Feichtenhofer, C. Masked feature prediction for self-supervised visual pre-training. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 19–24 June 2022; pp. 14668–14678. [Google Scholar]
- Zhao, L.; Gundavarapu, N.B.; Yuan, L.; Zhou, H.; Yan, S.; Sun, J.J.; Friedman, L.; Qian, R.; Weyand, T.; Zhao, Y.; et al. Videoprism: A foundational visual encoder for video understanding. arXiv 2024, arXiv:2402.13217. [Google Scholar]
- Tong, Z.; Song, Y.; Wang, J.; Wang, L. Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training. Adv. Neural Inf. Process. Syst. 2022, 35, 10078–10093. [Google Scholar] [CrossRef] [Scilit]
- Wang, L.; Huang, B.; Zhao, Z.; Tong, Z.; He, Y.; Wang, Y.; Wang, Y.; Qiao, Y. Videomae v2: Scaling video masked autoencoders with dual masking. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 18–22 June 2023; pp. 14549–14560. [Google Scholar]
- Wang, Y.; Li, K.; Li, Y.; He, Y.; Huang, B.; Zhao, Z.; Zhang, H.; Xu, J.; Liu, Y.; Wang, Z.; et al. Internvideo: General video foundation models via generative and discriminative learning. arXiv 2022, arXiv:2212.03191. [Google Scholar]
- Wang, Y.; Li, K.; Li, X.; Yu, J.; He, Y.; Chen, G.; Pei, B.; Zheng, R.; Xu, J.; Wang, Z.; et al. Internvideo2: Scaling video foundation models for multimodal video understanding. arXiv 2024, arXiv:2403.15377. [Google Scholar]
- Meier, J.F.; Mueller, F.B.; Ecker, A.; Lüddecke, T. TrAction: Action Recognition with Sparse Trajectories. arXiv 2026, arXiv:2606.03490. [Google Scholar]
- Rajasegaran, J.; Pavlakos, G.; Kanazawa, A.; Feichtenhofer, C.; Malik, J. On the benefits of 3d pose and tracking for human action recognition. In Proceedings of the 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: New York, NY, USA, 2023; pp. 640–649. [Google Scholar]
- Du, J.R.; Lin, K.Y.; Meng, J.; Zheng, W.S. Towards completeness: A generalizable action proposal generator for zero-shot temporal action localization. In Proceedings of the International Conference on Pattern Recognition; Springer: Berlin/Heidelberg, Germany, 2024; pp. 252–267. [Google Scholar]
- Zhang, Q.; Fang, J.; Yuan, R.; Tang, X.; Qi, Y.; Zhang, K.; Yuan, C. Weakly supervised temporal action localization via dual-prior collaborative learning guided by multimodal large language models. In Proceedings of the 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: New York, NY, USA, 2025; pp. 24139–24148. [Google Scholar]
- Bicsi, L.; Alexe, B.; Ionescu, R.T.; Leordeanu, M. JEDI: Joint expert distillation in a semi-supervised multi-dataset student-teacher scenario for video action recognition. In Proceedings of the 2023 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW); IEEE: New York, NY, USA, 2023; pp. 953–962. [Google Scholar]
- Wang, L.; Koniusz, P. Feature Hallucination for Self-supervised Action Recognition: L. Wang, P. Koniusz. Int. J. Comput. Vis. 2025, 133, 7612–7646. [Google Scholar]
- Mun, J.; Cho, M.; Han, B. Local-global video-text interactions for temporal grounding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Virtual, 14–19 June 2020; pp. 10810–10819. [Google Scholar]
- Lin, K.Q.; Zhang, P.; Chen, J.; Pramanick, S.; Gao, D.; Wang, A.J.; Yan, R.; Shou, M.Z. Univtg: Towards unified video-language temporal grounding. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Paris, France, 2–6 October 2023; pp. 2794–2804. [Google Scholar]
- Moon, W.; Hyun, S.; Park, S.; Park, D.; Heo, J.P. Query-dependent video representation for moment retrieval and highlight detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 18–22 June 2023; pp. 23023–23033. [Google Scholar]
- Xu, H.; Ghosh, G.; Huang, P.Y.; Okhonko, D.; Aghajanyan, A.; Metze, F.; Zettlemoyer, L.; Feichtenhofer, C. Videoclip: Contrastive pre-training for zero-shot video-text understanding. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, Punta Cana, Dominican Republic, 7–11 November 2021; pp. 6787–6800. [Google Scholar]
- Bain, M.; Nagrani, A.; Varol, G.; Zisserman, A. Frozen in time: A joint video and image encoder for end-to-end retrieval. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Montreal, QC, Canada, 10–17 October 2021; pp. 1728–1738. [Google Scholar]
- Luo, H.; Ji, L.; Zhong, M.; Chen, Y.; Lei, W.; Duan, N.; Li, T. Clip4clip: An empirical study of clip for end to end video clip retrieval and captioning. Neurocomputing 2022, 508, 293–304. [Google Scholar] [CrossRef] [Scilit]
- Xu, Y.; Sun, Y.; Zhai, B.; Li, M.; Liang, W.; Li, Y.; Du, S. Zero-shot video moment retrieval via off-the-shelf multimodal large language models. Proc. AAAI Conf. Artif. Intell. 2025, 39, 8978–8986. [Google Scholar] [CrossRef] [Scilit]
- Liang, R.; Yang, Y.; Lu, H.; Li, L. Efficient temporal sentence grounding in videos with multi-teacher knowledge distillation. arXiv 2023, arXiv:2308.03725. [Google Scholar]
- Croitoru, I.; Bogolin, S.V.; Leordeanu, M.; Jin, H.; Zisserman, A.; Albanie, S.; Liu, Y. Teachtext: Crossmodal generalized distillation for text-video retrieval. In Proceedings of the 2021 IEEE/CVF International Conference on Computer Vision (ICCV); IEEE: New York, NY, USA, 2021; pp. 11563–11573. [Google Scholar]
- Ventura, L.; Schmid, C.; Varol, G. Learning text-to-video retrieval from image captioning. Int. J. Comput. Vis. 2025, 133, 1834–1854. [Google Scholar] [CrossRef] [Scilit]
- Zhang, J.; Ye, Q.; Zhou, H.; Liang, H.; Luo, F. MAVIS: Multi-Agent Video Retrieval via Structured Video Understanding. In Proceedings of the Findings of the Association for Computational Linguistics: ACL 2026, San Diego, CA, USA, 2–7 July 2026; pp. 21751–21764. [Google Scholar]
- Zhou, L.; Zhou, Y.; Corso, J.J.; Socher, R.; Xiong, C. End-to-end dense video captioning with masked transformer. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 18–22 June 2018; pp. 8739–8748. [Google Scholar]
- Wang, T.; Zhang, R.; Lu, Z.; Zheng, F.; Cheng, R.; Luo, P. End-to-end dense video captioning with parallel decoding. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Montreal, QC, Canada, 10–17 October 2021; pp. 6847–6857. [Google Scholar]
- Yao, L.; Torabi, A.; Cho, K.; Ballas, N.; Pal, C.; Larochelle, H.; Courville, A. Describing Videos by Exploiting Temporal Structure. In Proceedings of the 2015 IEEE International Conference on Computer Vision (ICCV); IEEE: New York, NY, USA, 2015; pp. 4507–4515. [Google Scholar] [CrossRef] [Scilit]
- Lin, K.; Li, L.; Lin, C.C.; Ahmed, F.; Gan, Z.; Liu, Z.; Lu, Y.; Wang, L. SwinBERT: End-to-End Transformers with Sparse Attention for Video Captioning. In Proceedings of the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: New York, NY, USA, 2022; pp. 17928–17937. [Google Scholar] [CrossRef] [Scilit]
- Duta, I.; Nicolicioiu, A.L.; Bogolin, S.V.; Leordeanu, M. Mining for meaning: From vision to language through multiple networks consensus. arXiv 2018, arXiv:1806.01954. [Google Scholar]
- Duan, X.; Huang, W.; Gan, C.; Wang, J.; Zhu, W.; Huang, J. Weakly supervised dense event captioning in videos. Adv. Neural Inf. Process. Syst. 2018, 31, 3063–3073. [Google Scholar]
- Wu, B.; Niu, G.; Yu, J.; Xiao, X.; Zhang, J.; Wu, H. Weakly supervised dense video captioning via jointly usage of knowledge distillation and cross-modal matching. arXiv 2021, arXiv:2105.08252. [Google Scholar]
- Ma, Y.; Qing, L.; Li, G.; Qi, Y.; Beheshti, A.; Sheng, Q.Z.; Huang, Q. RETTA: Retrieval-enhanced test-time adaptation for zero-shot video captioning. Pattern Recognit. 2025, 171, 112170. [Google Scholar] [CrossRef] [Scilit]
- Lin, K.; Gan, Z.; Wang, L. Augmented partial mutual learning with frame masking for video captioning. Proc. AAAI Conf. Artif. Intell. 2021, 35, 2047–2055. [Google Scholar] [CrossRef] [Scilit]
- Lin, J.; Yin, H.; Ping, W.; Molchanov, P.; Shoeybi, M.; Han, S. Vila: On pre-training for visual language models. In Proceedings of the 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: New York, NY, USA, 2024; pp. 26679–26689. [Google Scholar]
- Fan, C.; Zhang, X.; Zhang, S.; Wang, W.; Zhang, C.; Huang, H. Heterogeneous memory enhanced multimodal attention model for video question answering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA, 15–20 June 2019; pp. 1999–2007. [Google Scholar]
- Le, T.M.; Le, V.; Venkatesh, S.; Tran, T. Hierarchical conditional relation networks for video question answering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Virtual, 14–19 June 2020; pp. 9972–9981. [Google Scholar]
- Lei, J.; Li, L.; Zhou, L.; Gan, Z.; Berg, T.L.; Bansal, M.; Liu, J. Less is more: Clipbert for video-and-language learning via sparse sampling. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 19–25 June 2021; pp. 7331–7341. [Google Scholar]
- Xiao, J.; Yao, A.; Liu, Z.; Li, Y.; Ji, W.; Chua, T.S. Video as conditional graph hierarchy for multi-granular question answering. Proc. AAAI Conf. Artif. Intell. 2022, 36, 2804–2812. [Google Scholar] [CrossRef] [Scilit]
- Zellers, R.; Lu, X.; Hessel, J.; Yu, Y.; Park, J.S.; Cao, J.; Farhadi, A.; Choi, Y. Merlot: Multimodal neural script knowledge models. Adv. Neural Inf. Process. Syst. 2021, 34, 23634–23651. [Google Scholar]
- Fu, T.J.; Li, L.; Gan, Z.; Lin, K.; Wang, W.Y.; Wang, L.; Liu, Z. Violet: End-to-end video-language transformers with masked visual-token modeling. arXiv 2021, arXiv:2111.12681. [Google Scholar]
- Xu, Y.; Liu, G.; Kompella, R.R.; Huang, T.; Hu, S.; Ilhan, F.; Tekin, S.F.; Yahn, Z.; Liu, L. A Multi-Agent Perception-Action Alliance for Efficient Long Video Reasoning. arXiv 2026, arXiv:2603.14052. [Google Scholar]
- Liang, T.; Tan, C.; Xia, B.; Zheng, W.S.; Hu, J.F. Ranking distillation for open-ended video question answering with insufficient labels. In Proceedings of the 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: New York, NY, USA, 2024; pp. 13161–13170. [Google Scholar]
- Liu, R.; Liu, Z.; Tang, J.; Ma, Y.; Pi, R.; Zhang, J.; Chen, Q. LongVideoAgent: Multi-Agent Reasoning with Long Videos. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics, San Diego, CA, USA, 2–7 July 2026; pp. 40404–40416. [Google Scholar] [CrossRef] [Scilit]
- Cheng, D.; Li, M.; Liu, J.; Guo, Y.; Jiang, B.; Liu, Q.; Chen, X.; Zhao, B. Enhancing long video understanding via hierarchical event-based memory. In Proceedings of the 2025 IEEE International Conference on Multimedia and Expo (ICME); IEEE: New York, NY, USA, 2025; pp. 1–6. [Google Scholar]
- Balažević, I.; Shi, Y.; Papalampidi, P.; Chaabouni, R.; Koppula, S.; Hénaff, O.J. Memory consolidation enables long-context video understanding. arXiv 2024, arXiv:2402.05861. [Google Scholar]
- Fan, Y.; Lin, W.; Chen, R.; Jiang, H.; Chai, W.; Wang, J.; Wang, K. GAM-agent: Game-theoretic and uncertainty-aware collaboration for complex visual reasoning. Adv. Neural Inf. Process. Syst. 2025, 38, 143569–143629. [Google Scholar]
- Menon, S.; Iscen, A.; Nagrani, A.; Weyand, T.; Vondrick, C.; Schmid, C. CAViAR: Critic-Augmented Video Agentic Reasoning. arXiv 2025, arXiv:2509.07680. [Google Scholar]
- Bhattacharyya, A.; Panchal, S.; Pourreza, R.; Lee, M.; Madan, P.; Memisevic, R. Look, remember and reason: Grounded reasoning in videos with language models. Proc. Int. Conf. Learn. Represent. 2024, 2024, 7339–7357. [Google Scholar]
- Niu, J.; Li, Y.; Miao, Z.; Ge, C.; Zhou, Y.; He, Q.; Dong, X.; Duan, H.; Ding, S.; Qian, R.; et al. OVO-Bench: How Far Are Your Video-LLMs from Real-World Online Video Understanding? In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 11–15 June 2025; pp. 18902–18913. [Google Scholar]
- Li, C.; Im, E.W.; Fazli, P. VidHalluc: Evaluating Temporal Hallucinations in Multimodal Large Language Models for Video Understanding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 11–15 June 2025; pp. 13723–13733. [Google Scholar]
- Ossowski, T.; Maqbool, D.; Chen, J.; Cai, Z.; Bradshaw, T.; Hu, J. Comma: A communicative multimodal multi-agent benchmark. arXiv 2024, arXiv:2410.07553. [Google Scholar]
- Davidson, T.R.; Fourney, A.; Amershi, S.; West, R.; Horvitz, E.; Kamar, E. The Collaboration Gap. arXiv 2025, arXiv:2511.02687. [Google Scholar]
- Lowe, R.; Foerster, J.; Boureau, Y.L.; Pineau, J.; Dauphin, Y. On the pitfalls of measuring emergent communication. arXiv 2019, arXiv:1903.05168. [Google Scholar]


| Survey | Primary Scope | Analytical Perspective | Collaboration-Centric |
|---|---|---|---|
| [23] | Temporal action segmentation | Core techniques and supervision regimes | × |
| [24] | Video Transformers | Pipeline design and architectural adaptations | × |
| [25] | Self-supervised learning for videos | Learning objectives and training paradigms | × |
| [28] | Vision-language pre-training | VLP design dimensions | × |
| [26] | Video-LLMs | Framework architectures and LLM functional roles | Partial |
| [29] | Large multimodal agents | Agent components, taxonomy, collaboration, and evaluation | Partial |
| [30] | Multimodal Agent AI | Agent integration, learning, categorization, and applications | Partial |
| Ours | Multi-model collaboration | Collaboration paradigms and mechanisms | ✓ |
| Paradigm | Coordination Type | Policy | Execution Behavior | Main Limitation |
|---|---|---|---|---|
| Static Collaboration | – | Predefined | Fixed execution pipeline | Limited adaptability |
| Dynamic Collaboration | Controller-based | Adaptive | Runtime routing, scheduling, or resource allocation | Policy-bound coordination |
| Agent-based | Agent-driven | Planning and iterative execution | High complexity |
| Coordination | Study | Functional Units | Communication/Load | State/Memory | Tasks | Benchmarks | Evaluation |
|---|---|---|---|---|---|---|---|
| Static | Collaborative Experts [77] | Object, scene, motion, face, audio, ASR, and OCR experts; gating | Pairwise feature gating; one-pass | Fused representation; no persistent memory | Video retrieval | MSR-VTT; MSVD; LSMDC; DiDeMo; ActivityNet-Captions | R@K; MdR; MnR |
| Video-RAG [78] | LVLM, Whisper, EasyOCR, APE, CLIP, Contriever, FAISS; scene graph | JSON retrieval requests and evidence; one-pass | OCR, ASR, and scene-graph databases | Long-video QA; retrieval | Video-MME; MLVU; LongVideoBench; VNBench | Accuracy; token cost; ablation | |
| VideoMulti–Agents [79] | Text, video, and scene-graph agents; Organizer | Independent reports and synthesis; parallel inter-agent execution | Agent memory and Organizer context | VideoQA | NExT-QA; IntentQA; EgoSchema | Accuracy; architecture ablation | |
| MAGNET [80] | AV-RAG, salient-frame selector, AV agents, and meta-agent | Evidence tuples and meta-agent synthesis; parallel map-reduce | Precomputed AV database; aggregation context | Audio-visual multi-video QA | AVHaystacks-50; AVHaystacks-Full | BLEU/CIDEr; Text Sim; GPT Eval; R@K; MTGS; STEM; human evaluation | |
| Controller-based | MoReVQA [81] | Event parsing, grounding, reasoning, and prediction; PaLM-2, PaLI-3, OWL-ViT, CLIP | Structured API calls and grounded outputs; staged adaptive execution | Frame IDs, event queue, QA/OCR flags, and grounded evidence | VideoQA; grounded VideoQA | NExT-QA; iVQA; EgoSchema; ActivityNet-QA; NExT-GQA | Accuracy; Acc@GQA; mIoP; mIoU |
| DrVideo [82] | Video-document, retrieval, augmentation, planning, interaction, and answer modules | Binary sufficiency feedback and typed frame requests; iterative | Augmented document and analysis history | Long-video QA | EgoSchema; MovieChat-1K; Video-MME | Accuracy; Gemini-Pro score; ablation | |
| VideoAgent [17] | GPT-4 controller; LaViLa/CogAgent; EVA-CLIP | JSON actions, confidence, and captions; multi-round | Chronological caption history and retrieved frames | Long-form VideoQA | EgoSchema; NExT-QA | Accuracy; frame count; frame efficiency | |
| Video-EM [69] | Query agent, CLIP, TransNetV2, Grounding-DINO, Qwen2.5-VL/Qwen3, downstream Video-LLMs | Query decomposition, narratives, and verification; recursive | Episodic event memory and scene relations | Long-video QA; narrative reasoning | Video-MME; LVBench; HourVideo; EgoSchema | Accuracy; frame count; runtime; FLOPs; ablation | |
| Agent-based | OmAgent [19] | Video2RAG, Conqueror, Divider, Rescuer, Tool Manager, Conclusive Synthesis | Typed results, subtasks, and tool outputs; recursive D&C | Knowledge database, timestamps, and task tree | Complex long-video understanding | OmAgent long-video benchmark; MBPP; FreshQA | Accuracy; event localization; summary-QA accuracy |
| TraveLER [18] | Planner, retriever, extractor, evaluator, and summarizer | Plans, timestamps, QA, and feedback; iterative replanning | Timestamp-keyed caption/QA memory | Zero-shot VideoQA | NExT-QA; EgoSchema; STAR; Perception Test; Causal-VidQA | Accuracy; frame efficiency; inference cost; ablation | |
| VideoChat-M1 [83] | Policy agents, retrieval/perception tools, voting, and lead agent | Shared policy memory and tool updates; multi-round | Policies, tool logs, and shared KV buffer | Long-video QA; spatial/temporal reasoning | LongVideoBench; Video-MME; MLVU; Video-Holmes; VSI-Bench; Charades-STA | Accuracy; M-avg/G-avg; category accuracy; mIoU; latency | |
| HiCrew [84] | Problem/task planning; text, visual, evidence, and answer agents | Workflow specifications, confidence, and tool outputs; ReAct | HybridTree and task state | Long-form VideoQA | EgoSchema; NExT-QA | Accuracy; role/planning/HybridTree ablations |
| Task | Datasets | Size | Sources | Remarks |
|---|---|---|---|---|
| R&L | Kinetics [10] | 240K clips | CVPR 2017 | Broad-scale benchmark evaluating generic human action recognition. |
| Something-Something V2 [133] | 220.8K videos | ICCV 2017 | Fine-grained benchmark emphasizing temporal dynamics and object interaction understanding. | |
| THUMOS14 [134] | 18.4K videos | CVIU 2016 | Benchmark supporting action recognition and temporal action localization | |
| ActivityNet v1.3 [135] | 19,994 videos | CVPR 2015 | Large-scale untrimmed-video benchmark for activity recognition and temporal action localization. | |
| AVA [136] | 437 clips | CVPR 2018 | Dense spatio-temporal benchmark for localized atomic human action recognition. | |
| CMG | MSR-VTT [137] | 10K clips | CVPR 2016 | Open-domain video-language benchmark for cross-modal retrieval and alignment. |
| DiDeMo [138] | 10.5K videos | ICCV 2017 | Benchmark for natural-language moment localization in unedited personal videos. | |
| Charades-STA [139] | 10K videos | ICCV 2017 | Sentence-grounded temporal localization benchmark built on indoor activity videos. | |
| TACoS [140] | 127 videos | TACL 2013 | Early cooking-domain benchmark for language-to-video grounding. | |
| ActivityNet Captions [141] | 20K videos | ICCV 2017 | Dense event benchmark for temporal grounding with natural-language queries. | |
| QVHighlights [142] | 10.1K videos | NeurIPS 2021 | Query-conditioned benchmark with joint moment retrieval and highlight detection. | |
| NExT-GQA [143] | 1.6K videos | CVPR 2024 | Visually grounded VideoQA benchmark extending NExT-QA with temporal grounding annotations. | |
| GD | MSVD [144] | 2.1K clips | ACL 2011 | Early open-domain benchmark for short-video caption generation. |
| MSR-VTT [137] | 10K clips | CVPR 2016 | Large open-domain benchmark for captioning diverse web videos. | |
| VATEX [145] | 41.3K videos | ICCV 2019 | Large-scale multilingual video-language dataset with English and Chinese captions. | |
| ActivityNet Captions [141] | 20K videos | ICCV 2017 | Multi-event captioning benchmark with temporally localized descriptions. | |
| YouCook2 [146] | 2K videos | AAAI 2018 | Procedural video benchmark with step-aligned instructional narrations. | |
| TVSum [147] | 50 videos | CVPR 2015 | Topic-diverse benchmark for supervised shot-level video summarization. | |
| RD | TGIF-QA [148] | 71.7K GIFs | CVPR 2017 | GIF-based VideoQA benchmark for temporal reasoning, counting, and transition understanding. |
| MSR-VTT-QA [149] | 10K videos | PR 2020 | VideoQA benchmark constructed from MSR-VTT videos with large-scale QA annotations. | |
| NExT-QA [150] | 5.4K videos | CVPR 2021 | VideoQA benchmark emphasizing causal and temporal reasoning. | |
| AGQA [151] | 9.6K videos | CVPR 2021 | Compositional reasoning benchmark built on structured action graphs. | |
| STAR [152] | 22K clips | NeurIPS 2021 | Situated reasoning benchmark covering interaction, sequence, prediction, and feasibility. | |
| Video-MME [132] | 900 videos | CVPR 2025 | Long-context multimodal video understanding benchmark with visual, audio, and subtitle information. | |
| EgoSchema [153] | 5K QA pairs | NeurIPS 2023 | Long-form egocentric VideoQA benchmark with five-way multiple-choice questions over three-minute clips. | |
| LongVideoBench [154] | 3.8K videos | NeurIPS 2024 | Long-context video-language benchmark with interleaved video and subtitle inputs for referring reasoning over videos up to one hour. |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Chen, Y.; Zhang, J.; Zhang, L.; Liu, C.; Gao, R.; Lu, Z.; Qi, J.; Lei, Q. A Survey of Multi-Model Collaboration in Video Understanding. Data 2026, 11, 230. https://doi.org/10.3390/data11090230
Chen Y, Zhang J, Zhang L, Liu C, Gao R, Lu Z, Qi J, Lei Q. A Survey of Multi-Model Collaboration in Video Understanding. Data. 2026; 11(9):230. https://doi.org/10.3390/data11090230
Chicago/Turabian StyleChen, Yi, Jianwei Zhang, Lei Zhang, Chang Liu, Rui Gao, Zhixian Lu, Jun Qi, and Qiyu Lei. 2026. "A Survey of Multi-Model Collaboration in Video Understanding" Data 11, no. 9: 230. https://doi.org/10.3390/data11090230
APA StyleChen, Y., Zhang, J., Zhang, L., Liu, C., Gao, R., Lu, Z., Qi, J., & Lei, Q. (2026). A Survey of Multi-Model Collaboration in Video Understanding. Data, 11(9), 230. https://doi.org/10.3390/data11090230

