Next Article in Journal
Predicting Musculoskeletal Injury Risk in Professional Football Using a Supervised Machine Learning Approach Based on Full-Season Multi-Protocol Neuromuscular Assessments
Previous Article in Journal
Chile-ED-Resp: A Curated and Reproducible Weekly Hospital Dataset for Respiratory Emergency-Demand Forecasting in Chile
Previous Article in Special Issue
Learning Adaptive Cross-Modal Interactions for Multimodal Sentiment Analysis
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Review

A Survey of Multi-Model Collaboration in Video Understanding

1
College of Computer Science, Chengdu University, Chengdu 610106, China
2
School of Artificial Intelligence, Sichuan University, Chengdu 610065, China
3
School of Electronic Information and Electrical Engineering, Chengdu University, Chengdu 610106, China
4
OMICTEK Corporation, Shanghai 201203, China
*
Author to whom correspondence should be addressed.
Data 2026, 11(9), 230; https://doi.org/10.3390/data11090230
Submission received: 19 July 2026 / Revised: 1 September 2026 / Accepted: 2 September 2026 / Published: 7 September 2026
(This article belongs to the Special Issue Vision-Based AI in the Real World: Data, Robustness and Deployment)

Abstract

The rapid development of multimodal foundation models has shifted video understanding from perception-centered recognition toward more general semantic interpretation, reasoning, and decision-making over dynamic visual content. As video understanding tasks increasingly require fine-grained perception, long-range temporal modeling, multimodal grounding, and adaptive reasoning, collaboration among heterogeneous functional units, including specialized models, modules, agents, memory systems, and external tools, has emerged as an important system-level paradigm. However, existing surveys mainly organize video understanding methods by architectures, learning strategies, or task categories, leaving the collaborative structure of modern systems insufficiently examined. This survey provides a structured narrative review of multi-model collaboration in video understanding, which we formulate as collaborative video understanding. We introduce a unified analytical framework that characterizes collaborative systems through functional units, inter-unit communication mechanisms, and collaborative state representations, and organize existing methods according to their coordination dynamics into static collaboration and dynamic collaboration, with the latter further distinguished into controller-based and agent-based collaboration. We further review representative benchmarks, evaluation metrics, and empirical analysis, showing that current evaluation protocols mainly capture task-level performance but provide limited insight into collaborative organization, memory use, adaptive execution, and system-level collaborative capability. Finally, we discuss key challenges and future directions, including adaptive task decomposition, semantically aligned inter-unit communication, persistent shared memory, uncertainty-aware error containment, evidence-grounded reasoning, and collaboration-centric evaluation. By reinterpreting video understanding from a collaborative systems perspective, this survey aims to provide a structured foundation for developing more adaptive, reliable, and scalable video understanding systems.

1. Introduction

Video understanding has become a fundamental research topic in computer vision due to its broad applications in surveillance [1], autonomous driving [2], industrial monitoring [3], human–computer interaction [4], and video-language reasoning [5]. Different from static image understanding, video understanding requires effective modeling of temporal dynamics, long-range dependencies, and cross-modal semantic interactions within evolving video content [6]. These characteristics pose substantial challenges for spatiotemporal representation learning, temporal reasoning, and multimodal semantic alignment, making video understanding considerably more complex than static image understanding.
Video understanding methods have evolved through several major technological paradigms. Early approaches primarily relied on handcrafted features [7,8] and conventional statistical models [9], which exhibited limited generalization capability in complex real-world environments. With the rapid advancement of deep learning, Convolutional Neural Networks (CNN) were developed to learn hierarchical spatiotemporal representations from video data, significantly enhancing local spatiotemporal feature extraction and motion modeling capability [10,11]. Nevertheless, these methods still suffered from limited capability in modeling long-range temporal dependencies. Transformer-based architectures further advanced the field by introducing global contextual interaction through self-attention mechanisms [12,13], thereby enabling more effective long-range temporal reasoning.
More recently, video-oriented multimodal foundation models have transformed video understanding from perception-oriented recognition toward more general-purpose semantic understanding and reasoning [14,15]. This evolution has expanded the scope of video understanding, requiring modern systems to support efficient perception, temporal abstraction, multimodal grounding, and adaptive decision-making. However, when video-oriented multimodal foundation models are implemented as monolithic architectures, they still struggle to meet these heterogeneous requirements because a single parameterization and fixed inference pathway must jointly balance perception, temporal abstraction, multimodal grounding, and adaptive reasoning. To address these limitations, multi-model collaboration has emerged as a promising paradigm for video understanding [16,17]. In this survey, we define collaborative video understanding as approaches that coordinate two or more independently callable and replaceable computational units with functionally differentiated roles through explicit information exchange, adaptive routing, shared state maintenance, or decision aggregation [18]. These units may be instantiated as separate neural models, foundation models, retrieval or memory modules, external tools, or autonomous agents, provided that they can be invoked and replaced as distinct computational units during inference [19]. Unlike simple feature fusion, single-model multimodal pretraining, or conventional ensemble averaging, these methods focus on how computational units communicate, divide labor, and adapt their roles during video understanding [20,21]. By exploiting the complementary strengths of specialized components, collaborative video understanding methods can better balance perception, reasoning, efficiency, and scalability than monolithic methods [22].
Despite the rapid progress in this area, existing surveys [23,24,25,26,27,28] have primarily reviewed video understanding from model-centric perspectives, focusing on application tasks, architectural designs, learning paradigms, and video-language models, as summarized in Table 1. Recent multimodal-agent surveys [29,30] have also discussed collaboration-related aspects, such as memory, planning, tools, and multi-agent coordination; however, their discussions are situated in broader multimodal-agent settings rather than focusing specifically on video understanding. Although these surveys provide valuable overviews of existing developments, they leave several collaboration-related questions underexplored: how are specialized units selected, what information is exchanged, how is shared state maintained, and how does coordination evolve from fixed pipelines to adaptive agentic methods [20,31]. To bridge this gap, we present a structured survey of multi-model collaboration in video understanding from the perspective of collaborative components and coordination mechanisms.
To identify relevant studies, we searched Web of Science, Scopus, IEEE Xplore, and ACM Digital Library, supplemented by Google Scholar, using combinations of keywords related to video understanding, including video understanding, video reasoning, video question answering, long-video understanding, and Video-LLM, as well as keywords related to collaborative systems, including multi-model collaboration, multi-agent collaboration, modular, memory-augmented, tool-augmented, retrieval-augmented, routing, and planning. We prioritized studies that (1) address video understanding tasks, (2) involve multiple independently callable and replaceable computational units with functionally differentiated roles, and (3) incorporate explicit collaboration mechanisms, such as coordination, routing, memory, tool use, retrieval, or decision aggregation. Studies focusing only on unimodal image understanding or pure video generation were generally excluded unless they directly informed the proposed taxonomy.
The main contributions of this survey are summarized as follows:
  • We provide a structured analysis of video understanding from a collaborative systems perspective by characterizing task requirements, collaborative components, and coordination challenges across heterogeneous video understanding scenarios.
  • We introduce a unified system-level framework for collaborative video understanding, in which collaborative methods are analyzed through three core components: functional units, inter-unit communication mechanisms, and collaborative state representations.
  • We further organize existing collaborative methods from the perspective of coordination mechanisms into two overarching paradigms, namely static collaboration and dynamic collaboration, revealing a progression from predefined execution structures toward adaptive and increasingly autonomous coordination.
  • We review representative benchmarks, evaluation metrics, and empirical evidence across major video understanding tasks, and discuss why current evaluation protocols remain insufficient for measuring collaborative organization, memory use, adaptive execution, and system-level collaborative capability.
  • We identify key challenges and future directions for collaborative video understanding, including structured task orchestration, semantically aligned inter-unit communication, persistent shared memory, uncertainty-aware error containment, evidence-grounded reasoning, trustworthy collaboration, and collaboration-centric evaluation.
The remainder of this survey is organized as follows. Section 2 formalizes the video understanding problem and summarizes its major task categories. Section 3 presents the collaborative video understanding framework and analyzes representative collaboration paradigms from the perspective of coordination mechanisms. Section 4 discusses benchmarks, evaluation metrics, and empirical evidence. Section 5 analyzes open challenges and future directions, and Section 6 concludes the survey.

2. Video Understanding Problem Definition

2.1. Problem Definition

Video understanding refers to a family of computational tasks that extract temporally grounded representations from video streams to support recognition [32], localization [33], cross-modal grounding [34], description [35], and reasoning [36]. Unlike image understanding, which operates on static visual observations, video understanding must model how entities, actions, and contexts evolve over time, thereby transforming frame-level observations into coherent event-level interpretations [37]. Early studies often treated video understanding as a recognition-oriented problem, focusing mainly on action recognition and event classification [23,24]. With the rise of multimodal foundation models and video-language models, the field has expanded toward open-ended semantic and causal reasoning, where systems are expected to infer what occurs in a video, when it occurs, how it evolves, and why it occurs under specific contextual conditions [5,38,39].
Formally, a video can be represented as V = { I , A * , L * } , where I = { I t } t = 1 T denotes an ordered sequence of T visual frames, A * denotes an optional audio modality, and L * denotes optional language information such as subtitles, transcripts, metadata, or task instructions. Given an optional task query Q, video understanding can be formulated as learning a task-dependent mapping Y = ϕ ( V , Q ) , where Y denotes the desired output. Depending on the application, Y may represent a recognition label, temporal localization result, textual description, question answer, structured prediction, reasoning trace, tool-use plan, or decision output [39,40,41,42]. Following the taxonomy in Figure 1, this survey groups video understanding tasks by their dominant output form and functional objective into four categories:
  • Recognition and Localization Tasks identify entities or events in video content and estimate their spatial or temporal locations [43,44]. The output Y typically takes the form of categorical labels, spatial regions, temporal segments, trajectories, or spatio-temporal event locations. Representative tasks include action recognition, event classification, object tracking, temporal action localization, and spatio-temporal action localization [45]. These tasks primarily rely on video-internal evidence to determine what occurs and where or when it occurs [46].
  • Cross-modal Matching and Grounding Tasks establish semantic correspondences between video content and external modalities, such as language or audio [47]. The output Y typically takes the form of retrieval rankings, similarity scores, matched video-text pairs, or spatially/temporally grounded regions or segments associated with an external query. Representative tasks include video-text retrieval, video-language alignment, video grounding, and audio-visual grounding [48,49]. Their defining characteristic is the explicit alignment of video content with information from another modality [50].
  • Generative Description Tasks convert video content into natural-language descriptions or condensed semantic summaries [51,52]. The output Y is typically a free-form or structured textual representation, such as a caption, a sequence of event descriptions, or a condensed summary. Representative tasks include video captioning, dense video captioning, video summarization, and event description [48,49,53]. These tasks emphasize semantic abstraction and narrative organization by converting observed entities, actions, and temporal relations into coherent linguistic representations [50].
    Reasoning and Decision-making Tasks infer implicit information beyond directly observable video content by integrating evidence across temporal, semantic, and multimodal contexts [26,54]. The output Y may take the form of answers, inferred temporal or causal relations, explanations, decisions, or action plans. Representative tasks include video question answering, temporal reasoning, causal reasoning, commonsense reasoning, embodied video reasoning, and video-agent decision-making [55,56]. These tasks are closely related to collaborative video understanding because they often require coordinated perceptual grounding, evidence retrieval, memory maintenance, and multi-step reasoning [57].
Figure 1. Taxonomy of video understanding tasks organized by their dominant output forms and functional objectives. Recognition and localization tasks predict semantic labels or spatial and temporal locations; cross-modal matching and grounding tasks align video content with language or audio; generative description tasks produce captions, event descriptions, or summaries; and reasoning and decision-making tasks integrate temporal and multimodal evidence to answer questions, infer implicit relations, or select actions.
Figure 1. Taxonomy of video understanding tasks organized by their dominant output forms and functional objectives. Recognition and localization tasks predict semantic labels or spatial and temporal locations; cross-modal matching and grounding tasks align video content with language or audio; generative description tasks produce captions, event descriptions, or summaries; and reasoning and decision-making tasks integrate temporal and multimodal evidence to answer questions, infer implicit relations, or select actions.
Data 11 00230 g001

2.2. Fundamental Characteristics of Video Understanding

Video understanding differs from static visual analysis due to the spatiotemporal and multimodal nature of video data [58,59]. From a fundamental perspective, video understanding can be characterized by three essential properties:
  • Temporal Dependency. Video semantics often emerge from information distributed across multiple temporal segments rather than individual observations [60]. Effective video understanding therefore requires modeling temporal continuity, motion evolution, and long-range temporal interactions [61].
  • Event-Centric Semantics. The semantic meaning of a video is primarily determined by events and interactions that evolve over time rather than isolated visual observations [62]. Consequently, effective video understanding requires identifying participating entities, modeling their relationships, and characterizing event evolution [36]. For example, understanding an interaction such as a person picking up an object and handing it to another person requires coordinated entity perception, relation modeling, and event-level temporal reasoning.
  • Multi-level and Multimodal Reasoning. Effective understanding requires integrating information across multiple semantic levels, ranging from visual appearance and motion patterns to event structures, semantic relationships, and contextual knowledge [35]. In addition, video content is often associated with complementary modalities such as language and audio, which provide information unavailable from visual observations alone [63]. Consequently, modern video understanding increasingly relies on multi-level and multimodal reasoning to integrate heterogeneous information sources and construct comprehensive semantic interpretations of dynamic scenes [64].

3. Collaborative Video Understanding Framework

Modern video understanding requires systems to simultaneously handle temporal dependency, event-centric semantics, and multi-level multimodal reasoning across heterogeneous task outputs. These requirements challenge monolithic video understanding methods, as their single parameterization and fixed inference pathways limit the balance among specialization, adaptability, and computational efficiency. Collaborative video understanding offers a system-level alternative by coordinating functionally specialized units through information exchange, shared state maintenance, and adaptive control, thereby reducing the trade-offs inherent in monolithic architectures [64,65]. To characterize existing collaborative approaches, this section presents a unified framework that analyzes collaborative systems from two complementary perspectives: collaborative components and collaborative mechanisms.

3.1. Collaborative Components

Collaborative video understanding methods can be analyzed through a unified system-level abstraction composed of three elements: (1) collaborative functional units, (2) inter-unit communication mechanisms, and (3) collaborative state representations, as illustrated in Figure 2. This abstraction provides a common basis for analyzing collaborative video understanding across different implementations, including Video-LLM-based systems, retrieval-augmented frameworks, long-video understanding methods, and agent-based architectures [66].
Collaborative functional units constitute the fundamental components that perform specialized functions, such as perception, retrieval, memory, reasoning, planning, or other tools. Formally, the collaborative system is modeled as
U = { u 1 , u 2 , , u N } ,
where u i denotes the i-th collaborative functional unit participating in the overall video understanding process. In practice, a collaborative functional unit may be instantiated as a foundation model [67], a retrieval component [64,68], a memory system [69], a reasoning module [56], a planning controller [17,70], an external tool [16,71], or an autonomous agent [18].
Inter-unit communication mechanisms define how collaborative functional units exchange task-relevant collaborative message during collaboration. Such collaborative message may involve visual features, semantic representations, retrieved evidence, memory contents, reasoning traces, planning instructions, or intermediate predictions [72,73]. The collaborative message transmitted from unit u i to unit u j at execution step t can be abstracted as
m i j t = c i j t ( o i t , S t , Q ) ,
where o i t denotes the intermediate output of unit u i , S t denotes the current collaborative state representation, Q denotes the optional task query, and c i j t ( · ) denotes the inter-unit communication function between two units. In this formulation, collaborative messages serve as the operational interface through which functional units condition one another, transfer evidence, and coordinate intermediate decisions during collaborative inference [74].
Collaborative state representations preserve and organize task-relevant context accumulated throughout the collaborative process. The system maintains a collaborative state representation S t , which denotes task-relevant information accessible globally, locally, or through a controller-mediated interface at execution step t [75]. Depending on the implementation, S t may contain contextual representations, retrieved knowledge, memory records, intermediate results, tool-call traces, routing decisions, uncertainty estimates, or other information required for coordination among collaborative functional units [76]. State evolution can be written as
S t + 1 = g ( S t , { m i j t } i , j , { o i t } i = 1 N ) ,
and the final task output is generated by
Y = h ( S T , Q ) ,
where g ( · ) denotes the state update function and h ( · ) denotes the task-specific output function.

3.2. Coordination Mechanisms

While collaborative functional units, inter-unit communication mechanisms, and collaborative state representations define the structural components of collaborative video understanding, coordination mechanisms determine how these components are organized during execution. From this perspective, existing collaborative systems are first distinguished according to whether their coordination policies remain fixed or are adaptively adjusted during execution, resulting in two primary paradigms: static collaboration and dynamic collaboration, as illustrated in Figure 3 and summarized in Table 2. Dynamic collaboration can be further characterized by its dominant coordination mechanism, including controller-based collaboration and agent-based collaboration. This hierarchical organization separates the primary distinction in coordination behavior from the mechanisms used to realize adaptive coordination, while allowing different implementations to be analyzed under a common framework.
To ground the proposed taxonomy in representative literature, Table 3 maps selected collaborative video understanding systems to their functional units, communication patterns, state or memory designs, coordination mechanisms, tasks, benchmarks, and evaluation criteria. The mapping illustrates how the proposed coordination categories can be consistently applied across systems with different architectures and application settings. In particular, systems are classified according to their dominant coordination behavior: static collaboration follows a predefined execution policy, whereas dynamic collaboration adapts execution online through either an explicit controller or autonomous agents.

3.2.1. Static Collaboration

Static collaboration represents the most fundamental collaboration paradigm for video understanding [85]. Its defining characteristic is that the coordination policy is predetermined before inference and remains invariant throughout execution [86]. Under this paradigm, collaborative functional units, their invocation order, inter-unit communication topology, and collaborative state evolution process are predefined during system design rather than dynamically adjusted according to the input video, intermediate results, or execution context [87,88,89]. Although the information processed by each unit varies across inputs, the collaborative structure itself remains fixed [90].
Under the unified analytical framework introduced in Section 3.1, static collaboration can be characterized by preconfigured unit assignment [91], execution ordering [92], and inter-unit interaction [93]. Individual collaborative functional units assume fixed responsibilities, such as visual perception, temporal modeling, evidence retrieval [94], semantic fusion, reasoning [95], or language generation [96], while collaboration is realized through predefined inter-unit communication among these units.
Representative implementations include cascaded processing pipelines [97,98], modular multi-stage frameworks [99], retrieval-augmented systems with predefined retrieval-generation workflows [100], and multi-component architectures employing fixed routing [101] among specialized modules. Despite substantial architectural differences, these systems share the same coordination principle: the collaborative strategy is specified before execution and remains unchanged throughout inference.
The primary limitation of static collaboration arises from its input-invariant coordination policy. Since unit selection and interaction patterns cannot adapt to the execution context, the system is unable to dynamically adjust computational resources [102], information routing [103], or reasoning depth according to video complexity [104], semantic ambiguity [105], or intermediate uncertainty [106]. Such limitations become increasingly evident in long-video understanding, multimodal reasoning, and open-ended inference [104], where different inputs often require substantially different collaborative behaviors [107]. These limitations have motivated the development of dynamic collaboration, in which coordination policies are adaptively adjusted during execution.

3.2.2. Dynamic Collaboration

Dynamic collaboration extends static collaboration by allowing the coordination policy to be adaptively adjusted during execution according to the input video, intermediate results, and execution context [17]. Its defining characteristic is that the collaborative structure is no longer fully predetermined before inference but can be partially reconfigured online as the task unfolds [18]. In this survey, we further distinguish two representative implementation mechanisms: controller-based collaboration and agent-based collaboration. Controller-based collaboration relies on an explicit controller, router, scheduler, or policy to determine the participation, ordering, interaction, or computational allocation of functional units, whereas agent-based collaboration introduces autonomous agents that directly participate in planning and coordination.
Controller-Based Collaboration
Controller-based collaboration represents a major implementation mechanism for dynamic coordination in video understanding. Under the unified analytical framework introduced in Section 3.1, dynamic collaboration is characterized by adaptive unit selection, execution scheduling, and inter-unit interaction [108]. In controller-based systems, these adaptive decisions are typically mediated by an explicit controller, router, scheduler, or policy that continuously evaluates the current execution state. Instead of following a fixed coordination topology, the system continuously evaluates the current collaborative state to determine which functional units should be activated, what evidence should be retrieved, how information should be routed, and whether additional computation is required [82]. Individual functional units still assume specialized responsibilities, such as temporal search, retrieval, memory access, perception, reasoning, or response generation, but their participation and interactions vary across different inputs [109]. Collaborative state representations further support adaptive coordination by maintaining intermediate evidence, retrieval history, uncertainty signals, or memory contents, thereby enabling subsequent coordination decisions to be conditioned on the evolving execution state [110].
Representative implementations include adaptive temporal search frameworks that progressively refine observation scope according to query relevance [111], dynamic retrieval-augmented systems that adjust retrieval pathways or evidence granularity according to question complexity and intermediate findings [112], memory-adaptive reasoning frameworks that iteratively determine which memory type or temporal scale should be accessed next [113], and uncertainty-aware inference pipelines that decide whether to continue evidence collection or terminate reasoning according to prediction reliability [114]. StreamChat targets streaming video understanding and multi-round interaction by combining hierarchical short-term, long-term, and dialogue memories with parallel scheduling for frame selection, memory formation, and response generation [115]. Under our taxonomy, it is characterized as controller-based collaboration, as its overall processing structure is predefined while component execution and memory updates adapt to the incoming video stream. Thinking-QwenVL decouples evidence-aware reasoning control from progressive memory integration through an Active Thinking Decision Maker and a Hierarchical Progressive Semantic Integration module, enabling adaptive response timing, confidence monitoring, and cross-clip state updates in online video understanding [116]. Under our taxonomy, it is characterized as controller-based collaboration, with ATDM serving as the reasoning controller and HPSI serving as the memory and integration unit. The two units coordinate through the evolving cognitive state, in which HPSI progressively integrates cross-clip visual evidence while ATDM conditions subsequent decisions on the updated state together with progress and confidence signals, enabling adaptive execution and self-triggered reflection. Although these approaches differ substantially in implementation, they share the same coordination principle: collaborative decisions are determined online according to the current execution state rather than being completely specified before inference. In contrast to fixed one-shot retrieval or single-pass modular pipelines [78,117], these methods explicitly modify the collaborative process during execution.
Compared with static collaboration, controller-based collaboration provides substantially greater flexibility, efficiency, and adaptability [118,119]. By tailoring coordination policies to individual inputs, the system can selectively activate specialized units, focus computation on relevant temporal regions, reduce unnecessary retrieval or reasoning steps, and improve robustness under heterogeneous video understanding demands [120]. Such adaptive coordination is particularly beneficial for long-video understanding, retrieval-intensive reasoning, and multimodal question answering, where different queries often require markedly different observation strategies, memory access patterns, evidence aggregation, and reasoning depth [121].
Nevertheless, controller-based collaboration typically operates within coordination policies whose decision space, optimization objective, and interaction patterns are specified implicitly or explicitly during training or system design [122]. Although the system adapts its collaborative behavior online, such policy-bound coordination limits autonomous planning, open-ended role negotiation, and self-directed coordination among heterogeneous computational units [123]. Consequently, controller-based methods remain primarily adaptive execution frameworks rather than fully autonomous collaborative systems [124]. These limitations motivate agent-based collaboration, in which autonomous agents actively plan, coordinate, and invoke heterogeneous computational resources according to task objectives [125].
Agent-Based Collaboration
Agent-based collaboration represents a more autonomous form of dynamic coordination, allowing autonomous agents to participate in goal decomposition, coordination planning, tool invocation, and iterative self-refinement during execution [17,125]. Under the unified analytical framework introduced in Section 3.1, this paradigm is characterized by agentized collaborative functional units, task-driven inter-unit communication, and policy-driven updates of collaborative state representations. Functional units are no longer restricted to predefined roles but are upgraded into agents capable of subgoal formation, strategy refinement, and execution control. These agents dynamically decide when and how to invoke other agents, tools, or memory components [126].
Inter-unit communication in this paradigm extends beyond intermediate features or retrieved evidence to include high-level semantic signals such as plans, critiques, confidence estimates, and tool-selection actions [127]. Collaborative state representations are correspondingly enriched with reasoning traces, interaction histories, agent feedback, and execution plans, thereby supporting more expressive coordination across agents [128]. In many systems, agents also participate in execution control by deciding whether to continue, terminate, backtrack, or replan. Formally, both the message function m i j t and state evolution process S t can be conditioned on agent policies defined over evolving goals and observations.
Representative implementations typically organize video understanding as a task-driven multi-agent process rather than a fixed pipeline. Complex tasks are decomposed into subgoals and assigned to specialized agents responsible for retrieval, grounding, reasoning, memory access, or generation [18]. These agents iteratively exchange evidence, plans, and feedback through explicit interaction protocols, enabling progressive refinement of the collaborative workflow [19]. Some methods further introduce reflection, critique, or self-revision mechanisms to filter unreliable intermediate outputs and improve execution robustness [17]. DoraemonGPT further exemplifies this paradigm by constructing task-related symbolic video memory and using an LLM-driven Monte Carlo Tree Search planner to adaptively schedule specialized sub-task tools, external knowledge tools, and utility tools [70]. Under our taxonomy, it represents agent-based collaboration because its tools are selected and sequenced online according to intermediate observations and planning feedback. Despite diverse implementations, these approaches share a common principle: coordination emerges from agent-mediated interactions over goals, evidence, and decisions rather than being fully predefined.
Compared with controller-based collaboration, agent-based collaboration provides stronger autonomy, flexibility, and generalization capability. By enabling agents to perform subgoal planning, tool selection, strategy revision, and iterative reasoning, this paradigm better supports long-horizon reasoning, multimodal grounding, and decision-oriented video understanding tasks [129]. It also offers a unified framework for heterogeneous tasks such as question answering, retrieval, planning, and embodied reasoning [130]. These properties make it particularly effective for long-video understanding scenarios requiring iterative reasoning and multi-step evidence aggregation. Nevertheless, despite its enhanced autonomy, several challenges still remain in practical deployments. These issues are further analyzed in Section 5.

4. Benchmarks, Metrics, and Empirical Analysis

Benchmarks provide the empirical basis for defining video understanding tasks, specifying annotation protocols, and measuring model performance [26]. This section reviews representative benchmarks, evaluation metrics, and empirical evidence across major task categories, focusing on the capabilities they assess, the protocols they adopt, and the comparability of reported results under heterogeneous settings [131,132].

4.1. Representative Benchmarks

This section analyzes representative datasets for major video understanding tasks from a task-oriented perspective [26]. Representative datasets across four benchmark categories are summarized in Table 4. The selection is intended to be representative rather than exhaustive, emphasizing benchmarks that are widely adopted, influential in the development of their respective task categories, or particularly relevant to recent collaborative and long-context video understanding research. Additional datasets exist across these task categories but are not individually listed because of limited adoption, narrower task scope, substantial overlap with representative benchmarks, or specialized evaluation objectives. The following discussion focuses on their objectives, annotation characteristics, evaluated capabilities, and limitations for system-level assessment.
  • Recognition and Localization Benchmarks evaluate whether systems can identify actions or events and locate their spatial or temporal occurrences. Representative datasets cover action recognition, temporal action localization, and spatio-temporal localization, with annotations ranging from category labels to temporal segments and spatial regions [10,133]. These benchmarks primarily assess visual perception, motion modeling, and temporal aggregation under predefined label spaces and explicit supervision. For collaborative video understanding, they provide evidence for coordinated perceptual and temporal processing, while offering limited evaluation of long-range memory, multi-step reasoning, and adaptive execution [132].
  • Cross-modal Matching and Grounding Benchmarks evaluate whether systems can align video content with language or other modalities. Representative datasets cover video-text retrieval, temporal grounding, query-conditioned moment retrieval and highlight detection, as well as visually grounded video question answering, with annotations including paired instances, ranked results, and temporal spans [137,138]. These benchmarks further differ in whether grounding is evaluated through global retrieval, localized temporal evidence, saliency or highlight identification, or explicit answer-evidence alignment. These benchmarks assess multimodal alignment and retrieval capabilities but exhibit substantial heterogeneity in supervision forms, query formulations, annotation granularity, and modality configurations, which limits direct comparison across datasets [156].
  • Generative Description Benchmarks evaluate whether systems can generate high-level textual representations from video content. Representative datasets cover video captioning, dense captioning, event description, and summarization, with annotations typically provided as sentence-level or event-level references [157,158]. These benchmarks assess semantic abstraction, temporal organization, and content selection, but their performance interpretation is highly sensitive to reference quality, annotation diversity, and evaluation protocols, complicating fair comparison across datasets.
  • Reasoning and Decision-making Benchmarks evaluate whether systems can infer implicit information by integrating temporal, semantic, and multimodal evidence. Representative datasets cover video question answering, causal reasoning, commonsense inference, and decision-oriented tasks, requiring stronger capabilities in evidence integration, memory utilization, and reasoning consistency [131,150,159]. Recent benchmarks further extend this category toward long-form and long-context video understanding, requiring models to integrate evidence over substantially longer temporal horizons and, in some cases, across multiple input modalities such as visual, audio, and subtitle information [132,153,154]. However, variations in task formulations, input configurations, and supervision strategies introduce substantial challenges for fair comparison across benchmarks. In more advanced long-horizon and agentic settings, these tasks further require coherent state maintenance, tool use, planning, and adaptive execution across extended inference processes [160].

4.2. Evaluation Metrics

Evaluation metrics translate benchmark outputs into measurable evidence of model capability. Existing protocols mainly emphasize final task outcomes [28,41], whereas collaborative video understanding further requires process-level evaluation of consistency, efficiency, memory use, adaptive routing, and interaction cost.
  • Task Outcome Metrics. Task outcome metrics remain the predominant evaluation paradigm in existing video understanding benchmarks. Recognition tasks commonly rely on accuracy-based measures, such as Top-1 and Top-5 accuracy, whereas localization tasks use overlap-based measures, including mAP under different temporal IoU thresholds and IoU-based matching criteria [161,162]. Cross-modal retrieval tasks typically employ ranking metrics, such as Recall@K, Mean Rank, and Median Rank, while temporal grounding tasks additionally use IoU-based criteria to assess localization accuracy [163,164]. Generation tasks use reference-based metrics, such as BLEU, METEOR, CIDEr, and SPICE, to measure textual similarity and semantic consistency, whereas dense video captioning additionally employs the task-specific SODAc metric [161,165]. Reasoning-oriented and composite video understanding tasks typically adopt accuracy-based or task-specific measures, including answer accuracy, grounded answer accuracy, category-level accuracy, event localization accuracy, and summary-based question answering accuracy [151]. Decision-oriented settings may additionally use success rate or task completion rate to assess decision effectiveness. These examples are representative rather than exhaustive, as task-specific benchmarks may adopt additional evaluation measures tailored to their particular objectives. Although these metrics effectively characterize task outcomes, they primarily evaluate final predictions under predefined task settings and provide limited insight into system-level behaviors.
  • Beyond Task-level Evaluation. Beyond task-level outcomes, recent evaluations have increasingly considered system-level properties, including output quality, consistency, and efficiency [17,166]. Quality and consistency assessments complement conventional task metrics by examining textual or semantic similarity, cross-modal alignment, and temporal consistency, while human and model-based evaluations provide additional assessment for open-ended outputs [167]. Efficiency-oriented criteria further characterize the computational and interaction costs of collaborative systems, including token cost, frame count, frame efficiency, inference cost, runtime, FLOPs, and latency [168]. These measures are particularly relevant to multi-model and agent-based systems, where adaptive coordination may introduce additional computation, iterative interactions, or redundant evidence processing [169]. However, such system-level evaluations remain less standardized than conventional task metrics, and their adoption varies substantially across benchmarks and studies [170].

4.3. Quantitative Performance Trends and Comparative Analysis

This subsection summarizes representative results across four video understanding regimes. Because these results are obtained on different datasets and under heterogeneous metrics, supervision settings, and inference protocols, they should not be interpreted as a unified ranking of overall video understanding capability. We instead use them to identify task-specific performance trends, compare successive technical paradigms, and determine what current evidence can reveal about collaborative video understanding.
  • Recognition and Localization.  Table 5 summarizes representative quantitative evidence for both non-collaborative and collaborative methods across recognition and localization tasks. Recognition and localization exhibit the clearest gains across successive technical paradigms. On Kinetics-400, Top-1 accuracy increases from 74.7 for TSM to 86.1 for MViTv2, 90.0 for VideoMAE V2, and 92.1 for InternVideo2. SSv2 shows a similar but increasingly saturated trajectory, improving from 64.3 for TSM to 73.3 for MViTv2 and 77.4 for InternVideo2. For temporal localization, InternVideo2 reaches 41.2 average mAP on ActivityNet v1.3 and 72.0 on THUMOS14, exceeding InternVideo by 2.2 and 0.4 points, respectively. On AVA v2.2, frame-mAP increases from 27.4 for X3D to 42.6 for VideoMAE V2 among the representative non-collaborative methods. The collaborative methods show a more heterogeneous pattern: Feature Hallucination VideoMAE V2 and Feature Hallucination InternVideo2 reach 87.5 and 91.6 on Kinetics-400, respectively, while their SSv2 results are 77.4 and 77.3. Other collaborative methods report task-specific results, including 68.3 for TrAction + V-JEPA 2-L on SSv2, 49.8 for MLLM4WTAL on THUMOS14, and 45.1 for LART-Hiera on AVA v2.2. These results show substantial gains from advances in model architectures and large-scale pre-training, while the relatively small gap among the strongest recent models on several established recognition benchmarks suggests an emerging saturation trend. Collaborative methods, however, do not exhibit a uniform performance advantage, partly because they target different objectives and are evaluated under heterogeneous backbone, training, and computational settings. Therefore, controlled comparisons with matched single-model baselines and computational budgets are needed to isolate and quantify the actual contribution of collaboration.
  • Cross-modal Matching and Grounding. Cross-modal matching and grounding show substantial but less uniform gains across the available dataset–metric combinations. On MSR-VTT retrieval, the reported Recall@1 values range from 20.9 for Collaborative Experts and 29.6 for TeachText to 62.8 for InternVideo2 and 78.6 for MAVIS. On Charades-STA, InternVideo2 obtains 70.03 Recall@1, IoU = 0.5, compared with 57.91 for EMTM and 38.4 for Moment-GPT; on QVHighlights, InternVideo2 reaches 49.24 moment-retrieval mAP, compared with 43.63 for UniVTG. As reflected in Table 6, these results demonstrate substantial progress in cross-modal matching and grounding, but the magnitude of improvement remains highly task-dependent. The mixed results of collaborative methods further suggest that the effectiveness of collaboration depends on how the collaboration mechanism aligns with the specific cross-modal objective and evaluation protocol.
Table 5. Representative quantitative evidence across recognition and localization tasks.
Table 5. Representative quantitative evidence across recognition and localization tasks.
Method TypeMethodVideo RecognitionTemporal LocalizationSpatio-Temporal Action Localization
Kinetics-400
Top-1
SSv2
Top-1
ActivityNet v1.3
Avg. mAP
THUMOS14
Avg. mAP
AVA v2.2
Frame-mAP
Non-collaborativeTSM [171]74.764.3
X3D [172]79.127.4
UniFormer [173]83.071.4
Video Swin [174]84.969.6
MViTv2 [175]86.173.334.4
MaskFeat [176]87.075.039.8
VideoPrism [177]87.268.537.8
VideoMAE [178]87.475.439.5
VideoMAE V2 [179]90.077.069.642.6
InternVideo [180]91.177.239.071.641.0
InternVideo2 [181]92.177.441.272.0
CollaborativeTrAction + DINOv2-B [182]62.3
TrAction + V-JEPA 2-L [182]68.3
LART-Hiera [183]45.1
GAP + CLIP [184]31.832.9
MLLM4WTAL [185]49.8
JEDI Student [186]82.08
Feature Hallucination VideoMAE V2 [187]87.577.4
Feature Hallucination InternVideo2 [187]91.677.3
Table 6. Representative quantitative evidence across cross-modal matching and grounding tasks.
Table 6. Representative quantitative evidence across cross-modal matching and grounding tasks.
Method TypeMethodVideo-Text
Retrieval
Temporal GroundingMoment
Retrieval
MSR-VTT
Recall@1
Charades-STA
Recall@1, IoU = 0.5
ActivityNet
Captions
Recall@1, IoU = 0.5
QVHighlights
mAP
Non-collaborative modelsLGI [188]59.4641.51
Moment-DETR [142]55.6536.14
UniVTG [189]60.1943.63
QD-DETR [190]57.3140.19
VideoCLIP [191]30.9
Frozen in Time [192]32.5
CLIP4Clip [193]44.5
InternVideo2 [181]62.870.0349.24
Collaborative modelsMoment-GPT [194]38.431.135.0
EMTM [195]57.9144.73
Collaborative Experts [77]20.9
TeachText [196]29.6
Learning from
Image Captioning [197]
39.2
MAVIS [198]78.6
  • Generative Description. Generative description exhibits substantial progress, although the strongest method varies across datasets and metrics. Table 7 summarizes representative results from both non-collaborative and collaborative methods across global and dense video captioning. On MSVD, CIDEr increases from 49.8 for RETTA and 51.7 for Enc–Dec + Local + Global to 146.2 for Vid2Seq, while Vid2Seq reaches 64.6 on MSR-VTT compared with 53.8 for both SwinBERT and Multiple Networks Consensus + OracleNet. Among the listed methods, no single method consistently dominates all ActivityNet Captions metrics: Masked Transformer records the highest METEOR score of 9.56, while Vid2Seq achieves the highest CIDEr and SODAc scores of 30.10 and 5.80, respectively. The collaborative methods span both global and dense captioning, with APML Ensemble reaching 109.5 CIDEr on MSVD and ECG obtaining 7.06 METEOR on ActivityNet Captions. Overall, the metric- and method-dependent variations indicate that captioning progress cannot be represented by a single performance ranking. The mixed performance of collaborative methods further suggests that collaboration is beneficial only when its design effectively supports the specific generation objective and evaluation criterion.
Table 7. Representative quantitative evidence across generative description tasks categorized by collaboration type.
Table 7. Representative quantitative evidence across generative description tasks categorized by collaboration type.
Method TypeMethodGlobal Video CaptioningDense Video Captioning
MSVDMSR-VTTVATEXActivityNet Captions
CIDErCIDErCIDErMETEORCIDErSODAc
Non-collaborativeDense-Captioning Events [141]4.8217.29
Masked Transformer [199]9.56
PDVC [200]7.5025.875.26
Enc–Dec + Local + Global [201]51.7
SwinBERT [202]120.653.873.0
Vid2Seq [161]146.264.68.5030.105.80
CollaborativeMultiple Networks
Consensus + OracleNet [203]
53.8
Weakly Supervised
Dense Event Captioning [204]
6.3018.77
ECG [205]7.0614.25
RETTA [206]49.824.323.8
APML Ensemble [207]109.552.7
  • Reasoning and Decision-making. Reasoning-oriented tasks show the largest apparent gains but also the greatest sensitivity to evaluation protocols. Table 8 summarizes representative results from non-collaborative and collaborative methods across video question answering, situated reasoning, and long-context multimodal understanding. On STAR, the reported accuracies range from 33.3 for HCRN and 36.8 for ClipBERT to 44.9 for TraveLER, although the latter is obtained under zero-shot evaluation. On Video-MME, VILA-1.5-34B reports 59.0 overall accuracy, while the collaborative methods Video-EM, A4VL, LVAgent, and VideoChat-M1 report 62.0, 77.2, 81.7, and 83.2, respectively; however, the subtitle setting is not explicitly specified for VideoChat-M1. MSR-VTT-QA also requires caution because the available results are reported under different methodological settings, and the table does not support a direct ranking across all methods. More broadly, these results do not isolate whether an improvement arises from visual representation, language-model scale, prompting, memory, or collaboration among functional units. Nevertheless, the stronger results of collaborative methods on Video-MME suggest that explicit coordination may be particularly beneficial for long-context reasoning, where complementary perception, memory, and decision processes must be integrated. This apparent advantage should, however, be interpreted cautiously because the non-collaborative evidence in the current comparison remains limited.
  • Cross-benchmark Comparability. Reliable comparison is restricted to results obtained on the same dataset with the same metric and evaluation protocol. Even within an individual column, differences in pre-training data, supervision resources, input modalities, model scale, inference configurations, and collaboration mechanisms can confound attribution. This limitation is particularly evident for reasoning-oriented tasks, where the reported results include both conventional evaluations and zero-shot settings, and the subtitle condition is not uniformly specified. The four tables should therefore be read as evidence of task-specific capability trends rather than as a global leaderboard of video understanding systems or a uniform ranking of collaborative versus non-collaborative approaches.
  • Implications for Collaborative Video Understanding. Across the four task regimes, performance gains appear most consistent on standardized recognition benchmarks, whereas grounding, generation, and reasoning exhibit stronger dataset, metric, and protocol dependence. The results also document a broad transition from task-specific architectures to large-scale pre-trained and foundation models, alongside an emerging body of collaborative approaches. However, the quantitative evidence does not establish a corresponding improvement attributable to collaboration across tasks. Existing metrics primarily evaluate final task outputs and do not isolate the contributions of functional-unit selection, inter-unit communication, shared-state maintenance, or adaptive execution. Consequently, Table 5, Table 6, Table 7 and Table 8 provide evidence of task-level capability evolution and the emerging quantitative presence of collaborative methods rather than direct evidence of collaboration effectiveness. Collaboration-centric evaluation should complement outcome metrics with matched-cost single-model baselines, unit-level ablations, communication and routing efficiency, memory fidelity, uncertainty calibration, and robustness to component failures.
Table 8. Representative quantitative evidence across reasoning and decision-making tasks categorized by collaboration type.
Table 8. Representative quantitative evidence across reasoning and decision-making tasks categorized by collaboration type.
Method TypeMethodParadigmVideo Question AnsweringSituated ReasoningLong-Context Multimodal Understanding
MSR-VTT-QA
Standard
Accuracy
STAR
Accuracy
Video-MME
Overall Accuracy
Non-collaborativeVILA-1.5-34B [208]Video-capable VLM59.0
HME [209]Heterogeneous multimodal attention33.0
HCRN [210]Conditional relation network35.633.3
ClipBERT [211]Sparse-sampling Transformer37.436.8
HQGA [212]Query-conditioned graph hierarchy38.6
MERLOT [213]Temporal multimodal Transformer43.1
VIOLET [214]Video-language Transformer43.9
CollaborativeVideo-EM [69]Multi-model tool orchestration with episodic memory62.0
A4VL [215]Multi-agent perception-action alliance77.2
LVAgent [169]Multi-round MLLM collaboration81.7
VideoChat-M1 [83]Multi-agent reinforcement learning83.2 *
TraveLER [18]Modular multi-LMM agents44.9
RADI-P [216]Training-stage teacher–student ranking distillation48.2
Note: denotes macro-averaged STAR accuracy over the four question types; denotes zero-shot STAR evaluation; * indicates that the subtitle setting is not explicitly specified for VideoChat-M1. Video-MME values are overall accuracies.

5. Challenges and Future Directions

Collaborative video understanding has been increasingly explored as a system-level paradigm to address the limitations of monolithic video models [16]. However, achieving effective collaboration requires systematic optimization of three fundamental system-level components: collaborative functional units, inter-unit communication mechanisms, and collaborative state representations. From this perspective, we identify six challenges that are critical to advancing collaborative video understanding.

5.1. Adaptive Task Decomposition and Functional Units Orchestration

Existing collaborative systems are often characterized by manually specified functional roles and relatively fixed execution workflows. For example, VideoMultiAgents assigns visual, textual, and scene-graph reasoning to different agents and aggregates their predictions through an organizer agent [79]. LongVideoAgent employs a master agent to coordinate grounding and visual perception agents for long-video understanding [217]. Although these designs provide effective structural priors, predefined decomposition strategies may struggle to accommodate videos and queries with diverse reasoning requirements. Future collaborative systems should explore adaptive task decomposition and functional unit orchestration through hierarchical planners that progressively decompose and refine complex tasks [19], learned routers that select functional units according to query characteristics, and uncertainty-aware mechanisms that invoke additional perception or verification when required [18]. Beyond adaptive routing, self-reconfigurable collaboration strategies are also essential for dynamically adjusting system topology when functional units exhibit failures, generate unreliable evidence, or encounter out-of-distribution inputs.

5.2. Efficient and Semantically Aligned Inter-Unit Communication

Collaborative video understanding requires the exchange of heterogeneous messages among functional units, including visual tokens, video clips, audio segments, subtitles, temporal intervals, scene graphs, natural-language descriptions, tool outputs, and intermediate predictions. These messages vary substantially in modality, temporal resolution, spatial granularity, and semantic abstraction. Without explicit alignment mechanisms, these messages may be difficult to integrate and may correspond to inconsistent temporal or semantic contexts [16]. A promising direction is to develop unified multimodal communication protocols in which each message encodes semantic content together with its temporal support, spatial support, source identity, and confidence. Adaptive-bandwidth communication is another important direction, where low-resolution summaries are exchanged initially and higher-resolution evidence is selectively requested when uncertainty remains high [114]. Multimodal interfaces that preserve visual evidence while maintaining linguistic abstraction may provide a more reliable basis for collaboration than purely language-based communication [66].

5.3. Persistent and Temporally Consistent Shared Memory

Shared or mediated states allow functional units to preserve task-relevant context across multiple interactions. However, constructing such states for long videos remains challenging because systems must determine what information should be retained, organized, updated, or discarded. Early approaches compress dense video tokens into sparse short- and long-term memories [160], while agent-based systems maintain structured temporal and object-centric memories that can be accessed through specialized tools. These mechanisms improve scalability, but aggressive compression may remove brief yet decisive visual evidence [218]. Conversely, retaining all frames or intermediate conclusions increases computational overhead and undermines the efficiency objective of memory mechanisms. Future collaborative systems should employ hierarchical, multimodal, and multi-timescale memory mechanisms [219]. Such designs may combine fine-grained visual memory for object attributes and local interactions, episodic memory for temporally localized events, and semantic memory for persistent relations and high-level knowledge.

5.4. Uncertainty-Aware Collaboration and Error Containment

Collaboration creates additional pathways through which errors can propagate. Errors produced by one functional unit may propagate to subsequent units, resulting in incorrect evidence retrieval or unsupported reasoning outcomes. Because many functional units are built upon similar foundation models or training data, their errors may also be correlated. Agreement among units therefore does not necessarily imply correctness, and a central organizer may inadvertently amplify a shared failure rather than resolve it [220]. Uncertainty-aware collaboration should incorporate disagreement detection, confidence calibration, evidence-level verification, and selective re-observation [120]. When conflicts arise, the system should identify the disputed evidence and acquire additional observations rather than resolve disagreements solely through linguistic priors [221].

5.5. Evidence-Grounded and Verifiable Collaborative Reasoning

A collaborative system may produce a correct answer without grounding it in the underlying visual evidence, particularly when language priors or subtitles provide plausible shortcuts. Conversely, the retrieved video evidence may be relevant while the intermediate reasoning process remains inconsistent with the underlying observations. Final-answer accuracy alone therefore does not establish whether the collaborative reasoning process is faithful [143]. This issue becomes more severe in long videos, where relevant events are often sparse and temporally dispersed, increasing the risk of reasoning from incomplete summaries or incorrectly retrieved evidence. Future systems should interleave retrieval, inspection, reasoning, and verification rather than treating perception and reasoning as isolated stages. Each intermediate claim should be linked to the supporting frame, clip, audio segment, or subtitle from which it was derived [222].

5.6. Collaboration-Centric Evaluation and Scalable Deployment

Existing video benchmarks have substantially expanded the coverage of video duration, task diversity, and multimodal inputs. Recent benchmarks have introduced long-context reasoning, multimodal understanding, online perception, and hallucination evaluation settings, including LongVideoBench [154], Video-MME [132], OVO-Bench [223], and VidHalluc [224]. However, these benchmarks primarily evaluate externally observable task outputs. They do not directly assess whether functional units provide complementary capabilities, whether exchanged messages contribute necessary information, whether shared states preserve decisive evidence, or whether collaboration yields improvements under comparable computational budgets. Collaboration-centric evaluation should therefore consider task performance, evidence grounding, unit contribution, communication efficiency, uncertainty calibration, robustness, and system cost [225]. Specifically, collaboration gains should be evaluated against strong individual or computationally matched baselines, unit contributions through ablations or marginal effects, communication efficiency through interaction volume, transmitted information, and latency, and adaptive execution through comparisons with fixed strategies under matched computational budgets [226]. Collaborative state and memory should further be assessed for evidence retrieval, temporal consistency, stale information, and conflict handling, complemented by oracle and random routing, message perturbation, and failure-robustness analyses [227]. Nevertheless, existing video benchmarks generally lack standardized protocols for these collaboration-specific dimensions, leaving the development of collaboration-oriented benchmarks that jointly evaluate task performance, collaboration effectiveness, and computational efficiency as an open research gap.

6. Conclusions

Recent advances in video understanding have expanded the field from perception-oriented recognition toward semantic interpretation, multimodal reasoning, and decision-making over dynamic visual content. This survey examined this transition from the perspective of collaborative video understanding, where specialized functional units, inter-unit communication mechanisms, and shared state representations are coordinated as a system-level framework for addressing heterogeneous video understanding tasks. Through this perspective, we reviewed existing methods under a unified analytical framework and organized them into two overarching paradigms: static collaboration and dynamic collaboration, with the latter further divided into controller-based and agent-based collaboration. This taxonomy reveals a broader evolution from predefined execution structures toward adaptive and increasingly autonomous coordination. We also analyzed representative benchmarks, metrics, and empirical evidence, showing that current evaluation protocols mainly capture task-level outcomes and remain insufficient for assessing collaboration quality, memory use, adaptive execution, evidence grounding, and system-level robustness. Advancing collaborative video understanding will require progress not only in the capabilities of individual models but also in the principles governing how heterogeneous functional units are selected, coordinated, updated, verified, and evaluated. Key research directions include adaptive task decomposition, semantically aligned inter-unit communication, persistent and temporally consistent memory, uncertainty-aware error containment, evidence-grounded reasoning, and collaboration-centric evaluation under realistic computational constraints. By consolidating existing progress and identifying these unresolved challenges, this survey aims to provide a structured foundation for developing collaborative video understanding systems that are adaptive, reliable, interpretable, and scalable.

Author Contributions

Conceptualization, J.Z. and L.Z.; methodology, Y.C., Z.L. and J.Z.; formal analysis, Y.C.; investigation, Y.C.; data curation, Y.C.; writing—original draft preparation, Y.C.; writing—review and editing, J.Z., L.Z., C.L., Q.L., J.Q., R.G. and Z.L.; visualization, Y.C.; supervision, J.Z., L.Z. and Q.L.; resources, L.Z., C.L., Q.L., R.G. and Z.L.; project administration, J.Z., L.Z. and C.L.; funding acquisition, J.Z. and C.L. All authors have read and agreed to the published version of the manuscript.

Funding

This research was funded by the National Natural Science Foundation of China, grant number 62406043 and the Natural Science Foundation of Sichuan Province, China, grant number 24NSFSC2722.

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

No new data were created or analyzed in this study. Data sharing is not applicable to this article, as this study is a review based on previously published literature and publicly available datasets.

Acknowledgments

Use of AI tools: This manuscript was written entirely by the authors. ChatGPT-5.5 was used only for final language polishing and proofreading.

Conflicts of Interest

Author Qiyu Lei was employed by the company 4OMICTEK Corporation. The remaining authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.

References

  1. Ke, S.R.; Hoang, L.; Lee, Y.J.; Hwang, J.N.; Yoo, J.H.; Choi, K.H. A Review on Video-Based Human Activity Recognition. Computers 2013, 2, 88–131. [Google Scholar] [CrossRef] [Scilit]
  2. AlShami, A.K.; Rabinowitz, R.; Lam, K.N.; Shleibik, Y.; Mersha, M.; Boult, T.E.; Kalita, J. SMART-vision: Survey of modern action recognition techniques in vision. Multimed. Tools Appl. 2025, 84, 32705–32776. [Google Scholar] [CrossRef] [Scilit]
  3. Jyothi, H.; Komala, M.; Mallikarjunaswamy, S. A Comprehensive Survey on Technologies in Video-based Event Detection and Recognition Using Machine Learning and Deep Learning Techniques. In 2024 Second International Conference on Networks, Multimedia and Information Technology (NMITCON); IEEE: New York, NY, USA, 2024. [Google Scholar]
  4. Sargano, A.B.; Angelov, P.; Habib, Z. A Comprehensive Review on Handcrafted and Learning-Based Action Representation Approaches for Human Activity Recognition. Appl. Sci. 2017, 7, 110. [Google Scholar] [CrossRef] [Scilit]
  5. Xiao, J.; Huang, N.; Qin, H.; Li, D.; Li, Y.; Zhu, F.; Tao, Z.; Yu, J.; Lin, L.; Chua, T.S.; et al. VideoQA in the Era of LLMs: An Empirical Study. Int. J. Comput. Vis. 2025, 133, 3970–3993. [Google Scholar] [CrossRef] [Scilit]
  6. Jenni, S.; Jin, H. Time-Equivariant Contrastive Video Representation Learning. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Montreal, QC, Canada, 10–17 October 2021; pp. 9950–9960. [Google Scholar]
  7. Wang, H.; Kläser, A.; Schmid, C.; Liu, C.L. Dense trajectories and motion boundary descriptors for action recognition. Int. J. Comput. Vis. 2013, 103, 60–79. [Google Scholar] [CrossRef] [Scilit]
  8. Wang, H.; Schmid, C. Action Recognition with Improved Trajectories. In Proceedings of the 2013 IEEE International Conference on Computer Vision, Sydney, NSW, Australia, 1–8 December 2013; pp. 3551–3558. [Google Scholar] [CrossRef] [Scilit]
  9. Sanchez, J.; Perronnin, F.; Mensink, T.; Verbeek, J. Image Classification with the Fisher Vector: Theory and Practice. Int. J. Comput. Vis. 2013, 105, 222–245. [Google Scholar] [CrossRef] [Scilit]
  10. Carreira, J.; Zisserman, A. Quo vadis, action recognition? A new model and the kinetics dataset. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA, 21–26 July 2017; pp. 6299–6308. [Google Scholar]
  11. Feichtenhofer, C.; Fan, H.; Malik, J.; He, K. Slowfast networks for video recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Seoul, Republic of Korea, 27 October–2 November 2019; pp. 6202–6211. [Google Scholar]
  12. Bertasius, G.; Wang, H.; Torresani, L. Is Space-Time Attention All You Need for Video Understanding? arXiv 2021, arXiv:2102.05095. [Google Scholar]
  13. Arnab, A.; Dehghani, M.; Heigold, G.; Sun, C.; Lučić, M.; Schmid, C. ViViT: A Video Vision Transformer. arXiv 2021, arXiv:2103.15691. [Google Scholar]
  14. Maaz, M.; Rasheed, H.; Khan, S.; Khan, F. Video-chatgpt: Towards detailed video understanding via large vision and language models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics, Volume 1: Long Papers, Bangkok, Thailand, 11–16 August 2024; pp. 12585–12602. [Google Scholar]
  15. Lin, B.; Ye, Y.; Zhu, B.; Cui, J.; Ning, M.; Jin, P.; Yuan, L. Video-LLaVA: Learning United Visual Representation by Alignment Before Projection. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Miami, FL, USA; Al-Onaizan, Y., Bansal, M., Chen, Y.N., Eds.; Association for Computational Linguistics: Kerrville, TX, USA, 2024; pp. 5971–5984. [Google Scholar] [CrossRef] [Scilit]
  16. Fan, Y.; Ma, X.; Wu, R.; Du, Y.; Li, J.; Gao, Z.; Li, Q. Videoagent: A memory-augmented multimodal agent for video understanding. In Proceedings of the European Conference on Computer Vision; Springer: Berlin/Heidelberg, Germany, 2024; pp. 75–92. [Google Scholar]
  17. Wang, X.; Zhang, Y.; Zohar, O.; Yeung-Levy, S. Videoagent: Long-form video understanding with large language model as agent. In Proceedings of the European Conference on Computer Vision; Springer: Berlin/Heidelberg, Germany, 2024; pp. 58–76. [Google Scholar]
  18. Shang, C.; You, A.; Subramanian, S.; Darrell, T.; Herzig, R. Traveler: A modular multi-lmm agent framework for video question-answering. arXiv 2024, arXiv:2404.01476. [Google Scholar]
  19. Zhang, L.; Zhao, T.; Ying, H.; Ma, Y.; Lee, K. Omagent: A multi-modal agent framework for complex video understanding with task divide-and-conquer. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Miami, FL, USA, 12–16 November 2024; pp. 10031–10045. [Google Scholar]
  20. Sleeman, W.C., IV; Kapoor, R.; Ghosh, P. Multimodal classification: Current landscape, taxonomy and future directions. ACM Comput. Surv. 2022, 55, 1–31. [Google Scholar] [CrossRef] [Scilit]
  21. Zeng, Y.; Zhang, X.; Li, H.; Wang, J.; Zhang, J.; Zhou, W. X2-VLM: All-in-One Pre-Trained Model for Vision-Language Tasks. IEEE Trans. Pattern Anal. Mach. Intell. 2023, 46, 3156–3168. [Google Scholar] [CrossRef] [Scilit]
  22. Liu, Z.; He, Y.; Wang, W.; Wang, W.; Wang, Y.; Chen, S.; Zhang, Q.; Lai, Z.; Yang, Y.; Li, Q.; et al. InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language. arXiv 2023, arXiv:2305.05662. [Google Scholar]
  23. Ding, G.; Sener, F.; Yao, A. Temporal Action Segmentation: An Analysis of Modern Techniques. IEEE Trans. Pattern Anal. Mach. Intell. 2024, 46, 1011–1030. [Google Scholar] [CrossRef] [Scilit]
  24. Selva, J.; Johansen, A.S.; Escalera, S.; Nasrollahi, K.; Moeslund, T.B.; Clapés, A. Video transformers: A survey. IEEE Trans. Pattern Anal. Mach. Intell. 2023, 45, 12922–12943. [Google Scholar] [CrossRef] [Scilit]
  25. Schiappa, M.C.; Rawat, Y.S.; Shah, M. Self-supervised learning for videos: A survey. ACM Comput. Surv. 2023, 55, 1–37. [Google Scholar] [CrossRef] [Scilit]
  26. Tang, Y.; Bi, J.; Xu, S.; Song, L.; Liang, S.; Wang, T.; Zhang, D.; An, J.; Lin, J.; Zhu, R.; et al. Video understanding with large language models: A survey. IEEE Trans. Circuits Syst. Video Technol. 2025, 36, 1355–1376. [Google Scholar] [CrossRef] [Scilit]
  27. Shabaninia, E.; Nezamabadi-pour, H.; Shafizadegan, F. Multimodal action recognition: A comprehensive survey on temporal modeling. Multimed. Tools Appl. 2024, 83, 59439–59489. [Google Scholar] [CrossRef] [Scilit]
  28. Chen, F.L.; Zhang, D.Z.; Han, M.L.; Chen, X.Y.; Shi, J.; Xu, S.; Xu, B. Vlp: A survey on vision-language pre-training. Mach. Intell. Res. 2023, 20, 38–56. [Google Scholar] [CrossRef] [Scilit]
  29. Xie, J.; Chen, Z.; Zhang, R.; Li, G. Large multimodal agents: A survey. Vis. Intell. 2025, 3, 24. [Google Scholar] [CrossRef] [Scilit]
  30. Durante, Z.; Huang, Q.; Wake, N.; Gong, R.; Park, J.S.; Sarkar, B.; Taori, R.; Noda, Y.; Terzopoulos, D.; Choi, Y.; et al. Agent AI: Surveying the Horizons of Multimodal Interaction. arXiv 2024, arXiv:2401.03568. [Google Scholar]
  31. Zhao, F.; Zhang, C.; Geng, B. Deep multimodal data fusion. ACM Comput. Surv. 2024, 56, 1–36. [Google Scholar] [CrossRef] [Scilit]
  32. Wang, M.; Xing, J.; Su, J.; Chen, J.; Liu, Y. Learning spatiotemporal and motion features in a unified 2d network for action recognition. IEEE Trans. Pattern Anal. Mach. Intell. 2022, 45, 3347–3362. [Google Scholar]
  33. Zhai, Y.; Wang, L.; Tang, W.; Zhang, Q.; Zheng, N.; Doermann, D.; Yuan, J.; Hua, G. Adaptive two-stream consensus network for weakly-supervised temporal action localization. IEEE Trans. Pattern Anal. Mach. Intell. 2022, 45, 4136–4151. [Google Scholar] [CrossRef] [Scilit]
  34. Li, J.; Tang, S.; Zhu, L.; Zhang, W.; Yang, Y.; Chua, T.S.; Wu, F.; Zhuang, Y. Variational cross-graph reasoning and adaptive structured semantics learning for compositional temporal grounding. IEEE Trans. Pattern Anal. Mach. Intell. 2023, 45, 12601–12617. [Google Scholar] [CrossRef] [Scilit]
  35. Li, G.; Ye, H.; Qi, Y.; Wang, S.; Qing, L.; Huang, Q.; Yang, M.H. Learning hierarchical modular networks for video captioning. IEEE Trans. Pattern Anal. Mach. Intell. 2023, 46, 1049–1064. [Google Scholar] [CrossRef] [Scilit]
  36. Bai, Z.; Wang, R.; Gao, D.; Chen, X. Event graph guided compositional spatial–temporal reasoning for video question answering. IEEE Trans. Image Process. 2024, 33, 1109–1121. [Google Scholar] [CrossRef] [Scilit]
  37. Huang, Q.; Xiong, Y.; Rao, A.; Wang, J.; Lin, D. Movienet: A holistic dataset for movie understanding. In Proceedings of the European Conference on Computer Vision; Springer: Berlin/Heidelberg, Germany, 2020; pp. 709–727. [Google Scholar]
  38. Geng, T.; Wang, T.; Duan, J.; Zhang, Y.; Guan, W.; Zheng, F.; Shao, L. UniAV: Unified Audio-Visual Perception for Multi-Task Video Event Localization. IEEE Trans. Pattern Anal. Mach. Intell. 2025, 47, 10280–10294. [Google Scholar] [CrossRef] [Scilit]
  39. Zhang, H.; Sun, A.; Jing, W.; Zhou, J.T. Temporal sentence grounding in videos: A survey and future directions. IEEE Trans. Pattern Anal. Mach. Intell. 2023, 45, 10443–10465. [Google Scholar] [CrossRef] [Scilit]
  40. Vahdani, E.; Tian, Y. Deep learning-based action detection in untrimmed videos: A survey. IEEE Trans. Pattern Anal. Mach. Intell. 2022, 45, 4302–4320. [Google Scholar] [CrossRef] [Scilit]
  41. Abdar, M.; Kollati, M.; Kuraparthi, S.; Pourpanah, F.; McDuff, D.; Ghavamzadeh, M.; Yan, S.; Mohamed, A.; Khosravi, A.; Cambria, E.; et al. A review of deep learning for video captioning. IEEE Trans. Pattern Anal. Mach. Intell. 2024, Early Access. [Google Scholar]
  42. Cong, Y.; Liao, W.; Ackermann, H.; Rosenhahn, B.; Yang, M.Y. Spatial-temporal transformer for dynamic scene graph generation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Montreal, QC, Canada, 10–17 October 2021; pp. 16372–16382. [Google Scholar]
  43. Stergiou, A.; Poppe, R. About time: Advances, challenges, and outlooks of action understanding. Int. J. Comput. Vis. 2025, 133, 6251–6315. [Google Scholar] [CrossRef] [Scilit]
  44. Hutchinson, M.S.; Gadepally, V.N. Video action understanding. IEEE Access 2021, 9, 134611–134637. [Google Scholar] [CrossRef] [Scilit]
  45. Karim, M.; Khalid, S.; Aleryani, A.; Khan, J.; Ullah, I.; Ali, Z. Human action recognition systems: A review of the trends and state-of-the-art. IEEE Access 2024, 12, 36372–36390. [Google Scholar] [CrossRef] [Scilit]
  46. Huang, T.E.; Liu, Y.; Van Gool, L.; Yu, F. Video task decathlon: Unifying image and video tasks in autonomous driving. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Paris, France, 1–6 October 2023; pp. 8647–8657. [Google Scholar]
  47. Nguyen, T.; Bin, Y.; Xiao, J.; Qu, L.; Li, Y.; Wu, J.Z.; Nguyen, C.D.; Ng, S.K.; Tuan, L.A. Video-language understanding: A survey from model architecture, model training, and data perspectives. In Proceedings of the Findings of the Association for Computational Linguistics: ACL 2024, Bangkok, Thailand, 11–16 August 2024; pp. 3636–3657. [Google Scholar]
  48. Zhang, Z.; Xu, D.; Ouyang, W.; Zhou, L. Dense video captioning using graph-based sentence summarization. IEEE Trans. Multimed. 2020, 23, 1799–1810. [Google Scholar] [CrossRef] [Scilit]
  49. Song, X.; Chen, J.; Wu, Z.; Jiang, Y.G. Spatial-temporal graphs for cross-modal text2video retrieval. IEEE Trans. Multimed. 2021, 24, 2914–2923. [Google Scholar] [CrossRef] [Scilit]
  50. Wray, M.; Doughty, H.; Damen, D. On semantic similarity in video retrieval. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 19–25 June 2021; pp. 3650–3660. [Google Scholar]
  51. Aafaq, N.; Mian, A.; Liu, W.; Gilani, S.Z.; Shah, M. Video description: A survey of methods, datasets, and evaluation metrics. ACM Comput. Surv. (CSUR) 2019, 52, 1–37. [Google Scholar]
  52. Qasim, I.; Horsch, A.; Prasad, D. Dense video captioning: A survey of techniques, datasets and evaluation protocols. ACM Comput. Surv. 2025, 57, 1–36. [Google Scholar] [CrossRef] [Scilit]
  53. Huang, J.H.; Huck Yang, C.H.; Chen, P.Y.; Chen, M.H.; Worring, M. Conditional Modeling-Based Automatic Video Summarization. ACM Trans. Multimed. Comput. Commun. Appl. 2026, 22, 1–21. [Google Scholar] [CrossRef] [Scilit]
  54. Zhong, Y.; Ji, W.; Xiao, J.; Li, Y.; Deng, W.; Chua, T.S. Video question answering: Datasets, algorithms and challenges. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, Abu Dhabi, United Arab Emirates, 7–11 December 2022; pp. 6439–6455. [Google Scholar]
  55. Liang, T.; Li, L.; Hu, J.F.; Yu, X.; Zheng, W.S.; Lai, J. Rethinking Temporal Context in Video-QA: A Comprehensive Study of Single-Frame Static Bias. IEEE Trans. Multimed. 2025, 27, 5077–5091. [Google Scholar] [CrossRef] [Scilit]
  56. Li, J.; Niu, L.; Zhang, L. From representation to reasoning: Towards both evidence and commonsense reasoning for video question-answering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 18–24 June 2022; pp. 21273–21282. [Google Scholar]
  57. Sun, Y.Z.; Sun, H.L.; Ma, J.C.; Zhang, P.; Huang, X.Y. Multimodal agent AI: A survey of recent advances and future directions. J. Comput. Sci. Technol. 2025, 40, 1046–1063. [Google Scholar] [CrossRef] [Scilit]
  58. Lan, X.; Yuan, Y.; Wang, X.; Wang, Z.; Zhu, W. A survey on temporal sentence grounding in videos. ACM Trans. Multimed. Comput. Commun. Appl. 2023, 19, 1–33. [Google Scholar] [CrossRef] [Scilit]
  59. Hui, T.; Liu, S.; Ding, Z.; Huang, S.; Li, G.; Wang, W.; Liu, L.; Han, J. Language-aware spatial-temporal collaboration for referring video segmentation. IEEE Trans. Pattern Anal. Mach. Intell. 2023, 45, 8646–8659. [Google Scholar] [CrossRef] [Scilit]
  60. Kwon, H.; Kim, M.; Kwak, S.; Cho, M. Learning self-similarity in space and time as generalized motion for video action recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Montreal, QC, Canada, 10–17 October 2021; pp. 13065–13075. [Google Scholar]
  61. Liu, X.; Pintea, S.L.; Nejadasl, F.K.; Booij, O.; Van Gemert, J.C. No frame left behind: Full video action recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 19–25 June 2021; pp. 14892–14901. [Google Scholar]
  62. Cao, Q.; Huang, H. VSRN: Visual-semantic relation network for video visual relation inference. IEEE Trans. Circuits Syst. Video Technol. 2021, 32, 768–777. [Google Scholar] [CrossRef] [Scilit]
  63. Zhao, B.; Gong, M.; Li, X. Audiovisual video summarization. IEEE Trans. Neural Netw. Learn. Syst. 2021, 34, 5181–5188. [Google Scholar] [CrossRef] [Scilit]
  64. Yu, T.; Yu, J.; Yu, Z.; Huang, Q.; Tian, Q. Long-term video question answering via multimodal hierarchical memory attentive networks. IEEE Trans. Circuits Syst. Video Technol. 2020, 31, 931–944. [Google Scholar] [CrossRef] [Scilit]
  65. Li, G.; Wang, X.; Zhu, W. AV-Unified: A Unified Framework for Audio-visual Scene Understanding. IEEE Trans. Multimed. 2026, Early Access. [Google Scholar]
  66. Zhao, H.; Ji, G.P.; Yan, R.; Xiong, H.; Li, Z. VideoExpert: Augmented LLM for Temporal-Sensitive Video Understanding. IEEE Trans. Circuits Syst. Video Technol. 2026, 36, 5893–5908. [Google Scholar] [CrossRef] [Scilit]
  67. Madan, N.; Møgelmose, A.; Modi, R.; Rawat, Y.S.; Moeslund, T.B. Foundation models for video understanding: A survey. arXiv 2024, arXiv:2405.03770. [Google Scholar]
  68. Xu, J.; Lan, C.; Xie, W.; Chen, X.; Lu, Y. Long video understanding with learnable retrieval in video-language models. IEEE Trans. Multimed. 2026, 28, 5254–5261. [Google Scholar] [CrossRef] [Scilit]
  69. Wang, Q.; Wang, Y.; Chen, H.; Wang, S.; Du, J.; Lee, C.H. Video Segmentation and Tokenization for Model-Based Video Scene Classification. IEEE Trans. Multimed. 2025, 27, 6489–6502. [Google Scholar] [CrossRef] [Scilit]
  70. Yang, Z.; Chen, G.; Li, X.; Wang, W.; Yang, Y. DoraemonGPT: Toward Understanding Dynamic Scenes with Large Language Models (Exemplified as A Video Agent). In Proceedings of the Forty-first International Conference on Machine Learning, Vienna, Austria, 21–27 July 2024; pp. 55976–55997. [Google Scholar]
  71. Li, J.; Liao, Z.; Xiao, F.; Li, T.; Zhang, Q.; Zhao, H.; Niu, L.; Chen, G.; Zhang, L.; Jiang, C. Parse, Align and Aggregate: Graph-driven Compositional Reasoning for Video Question Answering. IEEE Trans. Pattern Anal. Mach. Intell. 2026, 48, 5586–5603. [Google Scholar] [CrossRef] [Scilit]
  72. Li, L.; Jin, T.; Lin, W.; Jiang, H.; Pan, W.; Wang, J.; Xiao, S.; Xia, Y.; Jiang, W.; Zhao, Z. Multi-granularity relational attention network for audio-visual question answering. IEEE Trans. Circuits Syst. Video Technol. 2023, 34, 7080–7094. [Google Scholar] [CrossRef] [Scilit]
  73. Pei, B.; Huang, Y.; Chen, G.; Xu, J.; Wang, Y.; Wang, L.; Lu, T.; Qiao, Y.; Wu, F. Guiding audio-visual question answering with collective question reasoning. Int. J. Comput. Vis. 2025, 133, 6912–6929. [Google Scholar] [CrossRef] [Scilit]
  74. Chen, S.; Xu, Q.; Ma, Y.; Qiao, Y.; Wang, Y. Attentive snippet prompting for video retrieval. IEEE Trans. Multimed. 2023, 26, 4348–4359. [Google Scholar] [CrossRef] [Scilit]
  75. Chang, Z.; Zhang, X.; Wang, S.; Ma, S.; Gao, W. Stam: A spatiotemporal attention based memory for video prediction. IEEE Trans. Multimed. 2022, 25, 2354–2367. [Google Scholar] [CrossRef] [Scilit]
  76. Miao, B.; Bennamoun, M.; Gao, Y.; Shah, M.; Mian, A. Temporally consistent referring video object segmentation with hybrid memory. IEEE Trans. Circuits Syst. Video Technol. 2024, 34, 11373–11385. [Google Scholar] [CrossRef] [Scilit]
  77. Liu, Y.; Albanie, S.; Nagrani, A.; Zisserman, A. Use what you have: Video retrieval using representations from collaborative experts. arXiv 2019, arXiv:1907.13487. [Google Scholar]
  78. Luo, Y.; Zheng, X.; Li, G.; Yin, S.; Lin, H.; Fu, C.; Huang, J.; Ji, J.; Chao, F.; Luo, J.; et al. Video-rag: Visually-aligned retrieval-augmented long video comprehension. Adv. Neural Inf. Process. Syst. 2025, 38, 168008–168033. [Google Scholar]
  79. Kugo, N.; Li, X.; Li, Z.; Gupta, A.; Khatua, A.; Jain, N.; Patel, C.; Kyuragi, Y.; Ishii, Y.; Tanabiki, M.; et al. VideoMultiAgents: A Multi-Agent Framework for Video Question Answering. arXiv 2025, arXiv:2504.20091. [Google Scholar]
  80. Chowdhury, S.; Elmoghany, M.; Abeysinghe, Y.; Fei, J.; Nag, S.; Khan, S.; Elhoseiny, M.; Manocha, D. MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks. In Proceedings of the Advances in Neural Information Processing Systems, San Diego, CA, USA, 2–7 December 2025; pp. 49255–49291. [Google Scholar]
  81. Min, J.; Buch, S.; Nagrani, A.; Cho, M.; Schmid, C. Morevqa: Exploring modular reasoning models for video question answering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 17–21 June 2024; pp. 13235–13245. [Google Scholar]
  82. Ma, Z.; Gou, C.; Shi, H.; Sun, B.; Li, S.; Rezatofighi, H.; Cai, J. Drvideo: Document retrieval based long video understanding. In Proceedings of the Computer Vision and Pattern Recognition Conference, Nashville, TN, USA, 11–15 June 2025; pp. 18936–18946. [Google Scholar]
  83. Chen, B.; Wang, Z.; Yue, Z.; Yan, K.; Yu, C.; Huang, Y.; Liu, Z.; Wen, Y.; Chen, X.; Liu, Y.; et al. Videochat-m1: Collaborative policy planning for video understanding via multi-agent reinforcement learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Denver, CO, USA, 3–7 June 2026; pp. 33772–33783. [Google Scholar]
  84. Zhu, Y.; Zhao, J.; Zhao, J.; Mao, X.; Zhao, B. HiCrew: Hierarchical Reasoning for Long-Form Video Understanding via Question-Aware Multi-Agent Collaboration. arXiv 2026, arXiv:2604.21444. [Google Scholar]
  85. Yang, A.; Miech, A.; Sivic, J.; Laptev, I.; Schmid, C. Learning to answer visual questions from web videos. IEEE Trans. Pattern Anal. Mach. Intell. 2022, 47, 3202–3218. [Google Scholar] [CrossRef] [Scilit]
  86. Jiang, Y.; Yin, J. Clip-powered tass: Target-aware single-stream network for audio-visual question answering. Int. J. Comput. Vis. 2025, 133, 2581–2598. [Google Scholar] [CrossRef] [Scilit]
  87. Wang, R.; Luo, Y.; Zhang, F.; Liu, M.; Luo, X. HSSHG: Heuristic Semantics-Constrained Spatio-Temporal Heterogeneous Graph for VideoQA. IEEE Trans. Multimed. 2024, 26, 11176–11190. [Google Scholar] [CrossRef] [Scilit]
  88. Cheng, Y.; Fan, H.; Lin, D.; Sun, Y.; Kankanhalli, M.; Lim, J.H. Keyword-aware relative spatio-temporal graph networks for video question answering. IEEE Trans. Multimed. 2023, 26, 6131–6141. [Google Scholar] [CrossRef] [Scilit]
  89. Xu, W.; Yu, J.; Miao, Z.; Wan, L.; Tian, Y.; Ji, Q. Deep reinforcement polishing network for video captioning. IEEE Trans. Multimed. 2020, 23, 1772–1784. [Google Scholar] [CrossRef] [Scilit]
  90. Liu, F.; Liu, J.; Wang, W.; Lu, H. Hair: Hierarchical visual-semantic relational reasoning for video question answering. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Montreal, QC, Canada, 10–17 October 2021; pp. 1698–1707. [Google Scholar]
  91. Lu, J.; You, S.; Bao, B.K. Question Understanding and Temporality Guiding for Video Question Answering. IEEE Trans. Multimed. 2026, 28, 2772–2783. [Google Scholar] [CrossRef] [Scilit]
  92. Xu, W.; Miao, Z.; Yu, J.; Tian, Y.; Wan, L.; Ji, Q. Bridging video and text: A two-step polishing transformer for video captioning. IEEE Trans. Circuits Syst. Video Technol. 2022, 32, 6293–6307. [Google Scholar] [CrossRef] [Scilit]
  93. Qian, T.; Cui, R.; Chen, J.; Peng, P.; Guo, X.; Jiang, Y.G. Locate before answering: Answer guided question localization for video question answering. IEEE Trans. Multimed. 2023, 26, 4554–4563. [Google Scholar] [CrossRef] [Scilit]
  94. Zhou, S.; Xiao, J.; Yang, X.; Song, P.; Guo, D.; Yao, A.; Wang, M.; Chua, T.S. Scene-text grounding for text-based video question answering. IEEE Trans. Multimed. 2025, 28, 1417–1430. [Google Scholar] [CrossRef] [Scilit]
  95. Xiao, J.; Zhou, P.; Yao, A.; Li, Y.; Hong, R.; Yan, S.; Chua, T.S. Contrastive video question answering via video graph transformer. IEEE Trans. Pattern Anal. Mach. Intell. 2023, 45, 13265–13280. [Google Scholar] [CrossRef] [Scilit]
  96. Wu, K.; Li, X.; Li, X.; Zuo, K.; Lv, Z. AVQACL++: Toward a robust framework and benchmark for audio-visual question answering continual learning. IEEE Trans. Circuits Syst. Video Technol. 2026, 36, 8232–8245. [Google Scholar] [CrossRef] [Scilit]
  97. Gao, J.; Sun, X.; Ghanem, B.; Zhou, X.; Ge, S. Efficient video grounding with which-where reading comprehension. IEEE Trans. Circuits Syst. Video Technol. 2022, 32, 6900–6913. [Google Scholar] [CrossRef] [Scilit]
  98. Su, H.T.; Chang, C.H.; Shen, P.W.; Wang, Y.S.; Chang, Y.L.; Chang, Y.C.; Cheng, P.J.; Hsu, W.H. End-to-end video question-answer generation with generator-pretester network. IEEE Trans. Circuits Syst. Video Technol. 2021, 31, 4497–4507. [Google Scholar] [CrossRef] [Scilit]
  99. You, Z.; Wen, Z.; Chen, Y.; Li, X.; Zeng, R.; Wang, Y.; Tan, M. Toward long video understanding via fine-detailed video story generation. IEEE Trans. Circuits Syst. Video Technol. 2024, 35, 4592–4607. [Google Scholar] [CrossRef] [Scilit]
  100. Sun, X.; Dai, Y.; Wang, Y.; Ma, W.; Lin, X. Video question answering via traffic knowledge database and question classification. Multimed. Syst. 2024, 30, 14. [Google Scholar] [CrossRef] [Scilit]
  101. Lee, J.; Chang, J.; Lee, D.; Choi, J. CA2ST: Cross-Attention in Audio, Space, and Time for Holistic Video Recognition. IEEE Trans. Pattern Anal. Mach. Intell. 2025, 48, 2803–2819. [Google Scholar] [CrossRef] [Scilit]
  102. Song, P.; Zhang, L.; Lan, L.; Chen, W.; Guo, D.; Yang, X.; Wang, M. Towards efficient partially relevant video retrieval with active moment discovering. IEEE Trans. Multimed. 2025, 27, 6740–6751. [Google Scholar] [CrossRef] [Scilit]
  103. Li, X.; He, H.; Yang, Y.; Ding, H.; Yang, K.; Cheng, G.; Tong, Y.; Tao, D. Improving video instance segmentation via temporal pyramid routing. IEEE Trans. Pattern Anal. Mach. Intell. 2022, 45, 6594–6601. [Google Scholar] [CrossRef] [Scilit]
  104. Lai, C.; Ge, W.; Xue, X. Cross-Modal Complementary Learning and Template-Based Reasoning Chains for Future Event Prediction in Videos. IEEE Trans. Multimed. 2025, 27, 7497–7509. [Google Scholar] [CrossRef] [Scilit]
  105. Chen, L.; Lu, J.; Song, Z.; Zhou, J. Ambiguousness-aware state evolution for action prediction. IEEE Trans. Circuits Syst. Video Technol. 2022, 32, 6058–6072. [Google Scholar] [CrossRef] [Scilit]
  106. Xie, Z.; Luo, J.; Wu, K.; Kan, Z.; Guo, D. Learning Confidence-aware Prototypes for Weakly-supervised Video Anomaly Detection. IEEE Trans. Circuits Syst. Video Technol. 2025, 36, 5714–5728. [Google Scholar] [CrossRef] [Scilit]
  107. Huo, S.; Zhou, Y.; Wang, R.; Xiang, W.; Kung, S.Y. Semantic relevance learning for video-query based video moment retrieval. IEEE Trans. Multimed. 2023, 25, 9290–9301. [Google Scholar] [CrossRef] [Scilit]
  108. Yang, Z.; Chen, D.; Yu, X.; Shen, M.; Gan, C. Vca: Video curious agent for long video understanding. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Honolulu, HI, USA, 19–25 October 2025; pp. 20168–20179. [Google Scholar]
  109. Ye, J.; Wang, Z.; Sun, H.; Chandrasegaran, K.; Durante, Z.; Eyzaguirre, C.; Bisk, Y.; Niebles, J.C.; Adeli, E.; Li, F.F.; et al. Re-thinking temporal search for long-form video understanding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 11–15 June 2025; pp. 8579–8591. [Google Scholar]
  110. Kahatapitiya, K.; Ranasinghe, K.; Park, J.; Ryoo, M.S. Language repository for long video understanding. Proc. Find. Assoc. Comput. Linguist. ACL 2025, 2025, 5627–5646. [Google Scholar] [CrossRef] [Scilit]
  111. Wang, Z.; Yu, S.; Stengel-Eskin, E.; Yoon, J.; Cheng, F.; Bertasius, G.; Bansal, M. Videotree: Adaptive tree-based video representation for llm reasoning on long videos. In Proceedings of the Computer Vision and Pattern Recognition Conference, Nashville, TN, USA, 11–15 June 2025; pp. 3272–3283. [Google Scholar]
  112. Xue, Z.; Zhang, J.; Xie, X.; Cai, Y.; Liu, Y.; Li, X.; Tao, D. Adavideorag: Omni-contextual adaptive retrieval-augmented efficient long video understanding. arXiv 2025, arXiv:2506.13589. [Google Scholar]
  113. Yeo, W.; Kim, K.; Yoon, J.; Hwang, S.J. Worldmm: Dynamic multimodal memory agent for long video reasoning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Denver, CO, USA, 3–7 June 2026; pp. 25599–25609. [Google Scholar]
  114. Zhi, Z.; Wu, Q.; Li, W.; Li, Y.; Shao, K.; Zhou, K. Videoagent2: Enhancing the llm-based agent system for long-form video understanding by uncertainty-aware cot. arXiv 2025, arXiv:2504.04471. [Google Scholar]
  115. Xiong, H.; Yang, Z.; Yu, J.; Zhuge, Y.; Zhang, L.; Zhu, J.; Lu, H. Streaming video understanding and multi-round interaction with memory-enhanced knowledge. Proc. Int. Conf. Learn. Represent. 2025, 2025, 69332–69351. [Google Scholar]
  116. Zhang, K.; Yang, Z.; Han, M.; Hao, H.; Zhuge, Y.; Li, C.; Li, Z.; Chang, X. Progressive online video understanding with evidence-aligned timing and transparent decisions. Proc. Int. Conf. Learn. Represent. 2026, 2026, 5092–5121. [Google Scholar]
  117. Hu, K.; Gao, F.; Nie, X.; Zhou, P.; Tran, S.; Neiman, T.; Wang, L.; Shah, M.; Hamid, R.; Yin, B.; et al. M-llm based video frame selection for efficient video understanding. In Proceedings of the Computer Vision and Pattern Recognition Conference, Nashville, TN, USA, 11–15 June 2025; pp. 13702–13712. [Google Scholar]
  118. Tang, X.; Qiu, J.; Xie, L.; Tian, Y.; Jiao, J.; Ye, Q. Adaptive keyframe sampling for long video understanding. In Proceedings of the Computer Vision and Pattern Recognition Conference, Nashville, TN, USA, 11–15 June 2025; pp. 29118–29128. [Google Scholar]
  119. Zhang, S.; Yang, J.; Yin, J.; Luo, Z.; Luan, J. Q-frame: Query-aware frame selection and multi-resolution adaptation for video-llms. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Honolulu, HI, USA, 19–23 October 2025; pp. 22056–22065. [Google Scholar]
  120. Wang, Z.; Zhou, H.; Wang, S.; Li, J.; Xiong, C.; Savarese, S.; Bansal, M.; Ryoo, M.S.; Niebles, J.C. Active video perception: Iterative evidence seeking for agentic long video understanding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Denver, CO, USA, 3–7 June 2026; pp. 9088–9099. [Google Scholar]
  121. Diko, A.; Wang, T.; Swaileh, W.; Sun, S.; Patras, I. Rewind: Understanding long videos with instructed learnable memory. In Proceedings of the Computer Vision and Pattern Recognition Conference, Nashville, TN, USA, 11–15 June 2025; pp. 13734–13743. [Google Scholar]
  122. Yu, S.; Yoon, J.; Bansal, M. Crema: Generalizable and efficient video-language reasoning via multimodal modular fusion. Proc. Int. Conf. Learn. Represent. 2025, 2025, 74382–74406. [Google Scholar]
  123. Endo, M.; Hsu, J.; Li, J.; Wu, J. Motion question answering via modular motion programs. In Proceedings of the International Conference on Machine Learning, Honolulu, HI, USA, 23–29 July 2023; pp. 9312–9328. [Google Scholar]
  124. Choi, M.; Goel, H.; Omama, M.; Yang, Y.; Shah, S.; Chinchali, S. Towards neuro-symbolic video understanding. In Proceedings of the European Conference on Computer Vision; Springer: Berlin/Heidelberg, Germany, 2024; pp. 220–236. [Google Scholar]
  125. Liu, Y.; Qinghong Lin, K.; Chen, C.W.; Shou, M.Z. Videomind: A chain-of-lora agent for long video reasoning. arXiv 2025, arXiv:2503.13444. [Google Scholar]
  126. Shinn, N.; Cassano, F.; Gopinath, A.; Narasimhan, K.; Yao, S. Reflexion: Language agents with verbal reinforcement learning. Adv. Neural Inf. Process. Syst. 2023, 36, 8634–8652. [Google Scholar] [CrossRef] [Scilit]
  127. Du, Y.; Li, S.; Torralba, A.; Tenenbaum, J.B.; Mordatch, I. Improving Factuality and Reasoning in Language Models through Multiagent Debate. arXiv 2023, arXiv:2305.14325. [Google Scholar]
  128. Park, J.S.; O’Brien, J.; Cai, C.J.; Morris, M.R.; Liang, P.; Bernstein, M.S. Generative Agents: Interactive Simulacra of Human Behavior. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology, San Francisco, CA, USA, 29 October–1 November 2023. [Google Scholar] [CrossRef] [Scilit]
  129. Wang, Z.; Cai, S.; Liu, A.; Jin, Y.; Hou, J.; Zhang, B.; Lin, H.; He, Z.; Zheng, Z.; Yang, Y.; et al. Jarvis-1: Open-world multi-task agents with memory-augmented multimodal language models. IEEE Trans. Pattern Anal. Mach. Intell. 2024, 47, 1894–1907. [Google Scholar] [CrossRef] [Scilit]
  130. Szot, A.; Mazoure, B.; Attia, O.; Timofeev, A.; Agrawal, H.; Hjelm, D.; Gan, Z.; Kira, Z.; Toshev, A. From multimodal llms to generalist embodied agents: Methods and lessons. In Proceedings of the Computer Vision and Pattern Recognition Conference, Nashville, TN, USA, 11–15 June 2025; pp. 10644–10655. [Google Scholar]
  131. Li, K.; Wang, Y.; He, Y.; Li, Y.; Wang, Y.; Liu, Y.; Wang, Z.; Xu, J.; Chen, G.; Luo, P.; et al. Mvbench: A comprehensive multi-modal video understanding benchmark. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 17–21 June 2024; pp. 22195–22206. [Google Scholar]
  132. Fu, C.; Dai, Y.; Luo, Y.; Li, L.; Ren, S.; Zhang, R.; Wang, Z.; Zhou, C.; Shen, Y.; Zhang, M.; et al. Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-Modal LLMs in Video Analysis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 11–15 June 2025; pp. 24108–24118. [Google Scholar]
  133. Goyal, R.; Ebrahimi Kahou, S.; Michalski, V.; Materzynska, J.; Westphal, S.; Kim, H.; Haenel, V.; Fruend, I.; Yianilos, P.; Mueller-Freitag, M.; et al. The “something something” video database for learning and evaluating visual common sense. In Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy, 22–29 October 2017; pp. 5842–5850. [Google Scholar]
  134. Idrees, H.; Zamir, A.R.; Jiang, Y.G.; Gorban, A.; Laptev, I.; Sukthankar, R.; Shah, M. The thumos challenge on action recognition for videos “in the wild”. Comput. Vis. Image Underst. 2017, 155, 1–23. [Google Scholar] [CrossRef] [Scilit]
  135. Caba Heilbron, F.; Escorcia, V.; Ghanem, B.; Carlos Niebles, J. Activitynet: A large-scale video benchmark for human activity understanding. In Proceedings of the Ieee Conference on Computer Vision and Pattern Recognition, Boston, MA, USA, 7–12 June 2015; pp. 961–970. [Google Scholar]
  136. Gu, C.; Sun, C.; Ross, D.A.; Vondrick, C.; Pantofaru, C.; Li, Y.; Vijayanarasimhan, S.; Toderici, G.; Ricco, S.; Sukthankar, R.; et al. Ava: A video dataset of spatio-temporally localized atomic visual actions. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 18–22 June 2018; pp. 6047–6056. [Google Scholar]
  137. Xu, J.; Mei, T.; Yao, T.; Rui, Y. Msr-vtt: A large video description dataset for bridging video and language. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA, 26 June–1 July 2016; pp. 5288–5296. [Google Scholar]
  138. Anne Hendricks, L.; Wang, O.; Shechtman, E.; Sivic, J.; Darrell, T.; Russell, B. Localizing moments in video with natural language. In Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy, 22–29 October 2017; pp. 5803–5812. [Google Scholar]
  139. Gao, J.; Sun, C.; Yang, Z.; Nevatia, R. Tall: Temporal activity localization via language query. In Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy, 22–29 October 2017; pp. 5267–5275. [Google Scholar]
  140. Regneri, M.; Rohrbach, M.; Wetzel, D.; Thater, S.; Schiele, B.; Pinkal, M. Grounding action descriptions in videos. Trans. Assoc. Comput. Linguist. 2013, 1, 25–36. [Google Scholar] [CrossRef] [Scilit]
  141. Krishna, R.; Hata, K.; Ren, F.; Fei-Fei, L.; Carlos Niebles, J. Dense-captioning events in videos. In Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy, 22–29 October 2017; pp. 706–715. [Google Scholar]
  142. Lei, J.; Berg, T.L.; Bansal, M. Detecting moments and highlights in videos via natural language queries. Adv. Neural Inf. Process. Syst. 2021, 34, 11846–11858. [Google Scholar]
  143. Xiao, J.; Yao, A.; Li, Y.; Chua, T.S. Can I trust your answer? Visually grounded video question answering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 17–21 June 2024; pp. 13204–13214. [Google Scholar]
  144. Chen, D.; Dolan, W.B. Collecting highly parallel data for paraphrase evaluation. In Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies, Portland, OR, USA, 19–24 June 2011; pp. 190–200. [Google Scholar]
  145. Wang, X.; Wu, J.; Chen, J.; Li, L.; Wang, Y.F.; Wang, W.Y. Vatex: A large-scale, high-quality multilingual dataset for video-and-language research. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Seoul, Republic of Korea, 27 October–2 November 2019; pp. 4581–4591. [Google Scholar]
  146. Zhou, L.; Xu, C.; Corso, J. Towards automatic learning of procedures from web instructional videos. In Proceedings of the AAAI Conference on Artificial Intelligence, New Orleans, LA, USA, 2–7 February 2018; Volume 32. [Google Scholar]
  147. Song, Y.; Vallmitjana, J.; Stent, A.; Jaimes, A. Tvsum: Summarizing web videos using titles. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA, 7–12 June 2015; pp. 5179–5187. [Google Scholar]
  148. Jang, Y.; Song, Y.; Yu, Y.; Kim, Y.; Kim, G. Tgif-qa: Toward spatio-temporal reasoning in visual question answering. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA, 21–26 July 2017; pp. 2758–2766. [Google Scholar]
  149. Wang, W.; Huang, Y.; Wang, L. Long video question answering: A matching-guided attention model. Pattern Recognit. 2020, 102, 107248. [Google Scholar] [CrossRef] [Scilit]
  150. Xiao, J.; Shang, X.; Yao, A.; Chua, T.S. Next-qa: Next phase of question-answering to explaining temporal actions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 19–25 June 2021; pp. 9777–9786. [Google Scholar]
  151. Grunde-McLaughlin, M.; Krishna, R.; Agrawala, M. Agqa: A benchmark for compositional spatio-temporal reasoning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 19–25 June 2021; pp. 11287–11297. [Google Scholar]
  152. Wu, B.; Yu, S.; Chen, Z.; Tenenbaum, J.; Gan, C. STAR: A Benchmark for Situated Reasoning in Real-World Videos. In Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks 1, NeurIPS Datasets and Benchmarks 2021, Virtual, 6–14 December 2021. [Google Scholar]
  153. Mangalam, K.; Akshulakov, R.; Malik, J. Egoschema: A diagnostic benchmark for very long-form video language understanding. Adv. Neural Inf. Process. Syst. 2023, 36, 46212–46244. [Google Scholar] [CrossRef] [Scilit]
  154. Wu, H.; Li, D.; Chen, B.; Li, J. Longvideobench: A benchmark for long-context interleaved video-language understanding. Adv. Neural Inf. Process. Syst. 2024, 37, 28828–28857. [Google Scholar] [CrossRef] [Scilit]
  155. Pehlivan, S.; Laaksonen, J. Temporal teacher with masked transformers for semi-supervised action proposal generation. Mach. Vis. Appl. 2024, 35, 36. [Google Scholar] [CrossRef] [Scilit]
  156. Buch, S.; Eyzaguirre, C.; Gaidon, A.; Wu, J.; Fei-Fei, L.; Niebles, J.C. Revisiting the “video” in video-language understanding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 19–24 June 2022; pp. 2917–2927. [Google Scholar]
  157. Rohrbach, A.; Torabi, A.; Rohrbach, M.; Tandon, N.; Pal, C.; Larochelle, H.; Courville, A.; Schiele, B. Movie Description. Int. J. Comput. Vis. 2017, 123, 94–120. [Google Scholar] [CrossRef] [Scilit]
  158. Apostolidis, E.; Adamantidou, E.; Metsai, A.I.; Mezaris, V.; Patras, I. Video summarization using deep neural networks: A survey. Proc. IEEE 2021, 109, 1838–1863. [Google Scholar] [CrossRef] [Scilit]
  159. Liu, Y.; Li, S.; Liu, Y.; Wang, Y.; Ren, S.; Li, L.; Chen, S.; Sun, X.; Hou, L. Tempcompass: Do video llms really understand videos? Proc. Find. Assoc. Comput. Linguist. ACL 2024, 2024, 8731–8772. [Google Scholar] [CrossRef] [Scilit]
  160. Song, E.; Chai, W.; Wang, G.; Zhang, Y.; Zhou, H.; Wu, F.; Chi, H.; Guo, X.; Ye, T.; Zhang, Y.; et al. Moviechat: From dense token to sparse memory for long video understanding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 17–21 June 2024; pp. 18221–18232. [Google Scholar]
  161. Yang, A.; Nagrani, A.; Seo, P.H.; Miech, A.; Pont-Tuset, J.; Laptev, I.; Sivic, J.; Schmid, C. Vid2seq: Large-scale pretraining of a visual language model for dense video captioning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 18–22 June 2023; pp. 10714–10726. [Google Scholar]
  162. Wang, B.; Zhao, Y.; Yang, L.; Long, T.; Li, X. Temporal action localization in the deep learning era: A survey. IEEE Trans. Pattern Anal. Mach. Intell. 2023, 46, 2171–2190. [Google Scholar] [CrossRef] [Scilit]
  163. Wang, J.; Sun, G.; Wang, P.; Liu, D.; Dianat, S.; Rabbani, M.; Rao, R.; Tao, Z. Text is mass: Modeling as stochastic embedding for text-video retrieval. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 17–21 June 2024; pp. 16551–16560. [Google Scholar]
  164. Xu, H.; He, K.; Plummer, B.A.; Sigal, L.; Sclaroff, S.; Saenko, K. Multilevel language and vision integration for text-to-clip retrieval. Proc. AAAI Conf. Artif. Intell. 2019, 33, 9062–9069. [Google Scholar] [CrossRef] [Scilit]
  165. Rafiq, M.; Rafiq, G.; Choi, G.S. Video description: Datasets & evaluation metrics. IEEE Access 2021, 9, 121665–121685. [Google Scholar] [CrossRef] [Scilit]
  166. Jung, W.; Kim, J. QEVA: A Reference-Free Evaluation Metric for Narrative Video Summarization with Multimodal Question Answering. Proc. Find. Assoc. Comput. Linguist. EMNLP 2025, 2025, 24632–24642. [Google Scholar] [CrossRef] [Scilit]
  167. Shi, Y.; Yang, X.; Xu, H.; Yuan, C.; Li, B.; Hu, W.; Zha, Z.J. Emscore: Evaluating video captioning via coarse-grained and fine-grained embedding matching. In Proceedings of the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: New York, NY, USA, 2022; pp. 17908–17917. [Google Scholar]
  168. Jeoung, S.; Huybrechts, G.; Ganesh, B.; Galstyan, A.; Bodapati, S. Adaptive video understanding agent: Enhancing efficiency with dynamic frame sampling and feedback-driven reasoning. arXiv 2024, arXiv:2410.20252. [Google Scholar]
  169. Chen, B.; Yue, Z.; Chen, S.; Wang, Z.; Liu, Y.; Li, P.; Wang, Y. Lvagent: Long video understanding by multi-round dynamical collaboration of mllm agents. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Honolulu, HI, USA, 19–23 October 2025; pp. 20237–20246. [Google Scholar]
  170. Kumar, Y. VideoLLM Benchmarks and Evaluation: A Survey. arXiv 2025, arXiv:2505.03829. [Google Scholar]
  171. Lin, J.; Gan, C.; Han, S. Tsm: Temporal shift module for efficient video understanding. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Seoul, Republic of Korea, 27 October–2 November 2019; pp. 7083–7093. [Google Scholar]
  172. Feichtenhofer, C. X3d: Expanding architectures for efficient video recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Virtual, 14–19 June 2020; pp. 203–213. [Google Scholar]
  173. Li, K.; Wang, Y.; Gao, P.; Song, G.; Liu, Y.; Li, H.; Qiao, Y. Uniformer: Unified transformer for efficient spatiotemporal representation learning. arXiv 2022, arXiv:2201.04676. [Google Scholar]
  174. Liu, Z.; Ning, J.; Cao, Y.; Wei, Y.; Zhang, Z.; Lin, S.; Hu, H. Video swin transformer. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 19–24 June 2022; pp. 3202–3211. [Google Scholar]
  175. Li, Y.; Wu, C.Y.; Fan, H.; Mangalam, K.; Xiong, B.; Malik, J.; Feichtenhofer, C. Mvitv2: Improved multiscale vision transformers for classification and detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 19–24 June 2022; pp. 4804–4814. [Google Scholar]
  176. Wei, C.; Fan, H.; Xie, S.; Wu, C.Y.; Yuille, A.; Feichtenhofer, C. Masked feature prediction for self-supervised visual pre-training. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 19–24 June 2022; pp. 14668–14678. [Google Scholar]
  177. Zhao, L.; Gundavarapu, N.B.; Yuan, L.; Zhou, H.; Yan, S.; Sun, J.J.; Friedman, L.; Qian, R.; Weyand, T.; Zhao, Y.; et al. Videoprism: A foundational visual encoder for video understanding. arXiv 2024, arXiv:2402.13217. [Google Scholar]
  178. Tong, Z.; Song, Y.; Wang, J.; Wang, L. Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training. Adv. Neural Inf. Process. Syst. 2022, 35, 10078–10093. [Google Scholar] [CrossRef] [Scilit]
  179. Wang, L.; Huang, B.; Zhao, Z.; Tong, Z.; He, Y.; Wang, Y.; Wang, Y.; Qiao, Y. Videomae v2: Scaling video masked autoencoders with dual masking. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 18–22 June 2023; pp. 14549–14560. [Google Scholar]
  180. Wang, Y.; Li, K.; Li, Y.; He, Y.; Huang, B.; Zhao, Z.; Zhang, H.; Xu, J.; Liu, Y.; Wang, Z.; et al. Internvideo: General video foundation models via generative and discriminative learning. arXiv 2022, arXiv:2212.03191. [Google Scholar]
  181. Wang, Y.; Li, K.; Li, X.; Yu, J.; He, Y.; Chen, G.; Pei, B.; Zheng, R.; Xu, J.; Wang, Z.; et al. Internvideo2: Scaling video foundation models for multimodal video understanding. arXiv 2024, arXiv:2403.15377. [Google Scholar]
  182. Meier, J.F.; Mueller, F.B.; Ecker, A.; Lüddecke, T. TrAction: Action Recognition with Sparse Trajectories. arXiv 2026, arXiv:2606.03490. [Google Scholar]
  183. Rajasegaran, J.; Pavlakos, G.; Kanazawa, A.; Feichtenhofer, C.; Malik, J. On the benefits of 3d pose and tracking for human action recognition. In Proceedings of the 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: New York, NY, USA, 2023; pp. 640–649. [Google Scholar]
  184. Du, J.R.; Lin, K.Y.; Meng, J.; Zheng, W.S. Towards completeness: A generalizable action proposal generator for zero-shot temporal action localization. In Proceedings of the International Conference on Pattern Recognition; Springer: Berlin/Heidelberg, Germany, 2024; pp. 252–267. [Google Scholar]
  185. Zhang, Q.; Fang, J.; Yuan, R.; Tang, X.; Qi, Y.; Zhang, K.; Yuan, C. Weakly supervised temporal action localization via dual-prior collaborative learning guided by multimodal large language models. In Proceedings of the 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: New York, NY, USA, 2025; pp. 24139–24148. [Google Scholar]
  186. Bicsi, L.; Alexe, B.; Ionescu, R.T.; Leordeanu, M. JEDI: Joint expert distillation in a semi-supervised multi-dataset student-teacher scenario for video action recognition. In Proceedings of the 2023 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW); IEEE: New York, NY, USA, 2023; pp. 953–962. [Google Scholar]
  187. Wang, L.; Koniusz, P. Feature Hallucination for Self-supervised Action Recognition: L. Wang, P. Koniusz. Int. J. Comput. Vis. 2025, 133, 7612–7646. [Google Scholar]
  188. Mun, J.; Cho, M.; Han, B. Local-global video-text interactions for temporal grounding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Virtual, 14–19 June 2020; pp. 10810–10819. [Google Scholar]
  189. Lin, K.Q.; Zhang, P.; Chen, J.; Pramanick, S.; Gao, D.; Wang, A.J.; Yan, R.; Shou, M.Z. Univtg: Towards unified video-language temporal grounding. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Paris, France, 2–6 October 2023; pp. 2794–2804. [Google Scholar]
  190. Moon, W.; Hyun, S.; Park, S.; Park, D.; Heo, J.P. Query-dependent video representation for moment retrieval and highlight detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 18–22 June 2023; pp. 23023–23033. [Google Scholar]
  191. Xu, H.; Ghosh, G.; Huang, P.Y.; Okhonko, D.; Aghajanyan, A.; Metze, F.; Zettlemoyer, L.; Feichtenhofer, C. Videoclip: Contrastive pre-training for zero-shot video-text understanding. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, Punta Cana, Dominican Republic, 7–11 November 2021; pp. 6787–6800. [Google Scholar]
  192. Bain, M.; Nagrani, A.; Varol, G.; Zisserman, A. Frozen in time: A joint video and image encoder for end-to-end retrieval. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Montreal, QC, Canada, 10–17 October 2021; pp. 1728–1738. [Google Scholar]
  193. Luo, H.; Ji, L.; Zhong, M.; Chen, Y.; Lei, W.; Duan, N.; Li, T. Clip4clip: An empirical study of clip for end to end video clip retrieval and captioning. Neurocomputing 2022, 508, 293–304. [Google Scholar] [CrossRef] [Scilit]
  194. Xu, Y.; Sun, Y.; Zhai, B.; Li, M.; Liang, W.; Li, Y.; Du, S. Zero-shot video moment retrieval via off-the-shelf multimodal large language models. Proc. AAAI Conf. Artif. Intell. 2025, 39, 8978–8986. [Google Scholar] [CrossRef] [Scilit]
  195. Liang, R.; Yang, Y.; Lu, H.; Li, L. Efficient temporal sentence grounding in videos with multi-teacher knowledge distillation. arXiv 2023, arXiv:2308.03725. [Google Scholar]
  196. Croitoru, I.; Bogolin, S.V.; Leordeanu, M.; Jin, H.; Zisserman, A.; Albanie, S.; Liu, Y. Teachtext: Crossmodal generalized distillation for text-video retrieval. In Proceedings of the 2021 IEEE/CVF International Conference on Computer Vision (ICCV); IEEE: New York, NY, USA, 2021; pp. 11563–11573. [Google Scholar]
  197. Ventura, L.; Schmid, C.; Varol, G. Learning text-to-video retrieval from image captioning. Int. J. Comput. Vis. 2025, 133, 1834–1854. [Google Scholar] [CrossRef] [Scilit]
  198. Zhang, J.; Ye, Q.; Zhou, H.; Liang, H.; Luo, F. MAVIS: Multi-Agent Video Retrieval via Structured Video Understanding. In Proceedings of the Findings of the Association for Computational Linguistics: ACL 2026, San Diego, CA, USA, 2–7 July 2026; pp. 21751–21764. [Google Scholar]
  199. Zhou, L.; Zhou, Y.; Corso, J.J.; Socher, R.; Xiong, C. End-to-end dense video captioning with masked transformer. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 18–22 June 2018; pp. 8739–8748. [Google Scholar]
  200. Wang, T.; Zhang, R.; Lu, Z.; Zheng, F.; Cheng, R.; Luo, P. End-to-end dense video captioning with parallel decoding. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Montreal, QC, Canada, 10–17 October 2021; pp. 6847–6857. [Google Scholar]
  201. Yao, L.; Torabi, A.; Cho, K.; Ballas, N.; Pal, C.; Larochelle, H.; Courville, A. Describing Videos by Exploiting Temporal Structure. In Proceedings of the 2015 IEEE International Conference on Computer Vision (ICCV); IEEE: New York, NY, USA, 2015; pp. 4507–4515. [Google Scholar] [CrossRef] [Scilit]
  202. Lin, K.; Li, L.; Lin, C.C.; Ahmed, F.; Gan, Z.; Liu, Z.; Lu, Y.; Wang, L. SwinBERT: End-to-End Transformers with Sparse Attention for Video Captioning. In Proceedings of the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: New York, NY, USA, 2022; pp. 17928–17937. [Google Scholar] [CrossRef] [Scilit]
  203. Duta, I.; Nicolicioiu, A.L.; Bogolin, S.V.; Leordeanu, M. Mining for meaning: From vision to language through multiple networks consensus. arXiv 2018, arXiv:1806.01954. [Google Scholar]
  204. Duan, X.; Huang, W.; Gan, C.; Wang, J.; Zhu, W.; Huang, J. Weakly supervised dense event captioning in videos. Adv. Neural Inf. Process. Syst. 2018, 31, 3063–3073. [Google Scholar]
  205. Wu, B.; Niu, G.; Yu, J.; Xiao, X.; Zhang, J.; Wu, H. Weakly supervised dense video captioning via jointly usage of knowledge distillation and cross-modal matching. arXiv 2021, arXiv:2105.08252. [Google Scholar]
  206. Ma, Y.; Qing, L.; Li, G.; Qi, Y.; Beheshti, A.; Sheng, Q.Z.; Huang, Q. RETTA: Retrieval-enhanced test-time adaptation for zero-shot video captioning. Pattern Recognit. 2025, 171, 112170. [Google Scholar] [CrossRef] [Scilit]
  207. Lin, K.; Gan, Z.; Wang, L. Augmented partial mutual learning with frame masking for video captioning. Proc. AAAI Conf. Artif. Intell. 2021, 35, 2047–2055. [Google Scholar] [CrossRef] [Scilit]
  208. Lin, J.; Yin, H.; Ping, W.; Molchanov, P.; Shoeybi, M.; Han, S. Vila: On pre-training for visual language models. In Proceedings of the 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: New York, NY, USA, 2024; pp. 26679–26689. [Google Scholar]
  209. Fan, C.; Zhang, X.; Zhang, S.; Wang, W.; Zhang, C.; Huang, H. Heterogeneous memory enhanced multimodal attention model for video question answering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA, 15–20 June 2019; pp. 1999–2007. [Google Scholar]
  210. Le, T.M.; Le, V.; Venkatesh, S.; Tran, T. Hierarchical conditional relation networks for video question answering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Virtual, 14–19 June 2020; pp. 9972–9981. [Google Scholar]
  211. Lei, J.; Li, L.; Zhou, L.; Gan, Z.; Berg, T.L.; Bansal, M.; Liu, J. Less is more: Clipbert for video-and-language learning via sparse sampling. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 19–25 June 2021; pp. 7331–7341. [Google Scholar]
  212. Xiao, J.; Yao, A.; Liu, Z.; Li, Y.; Ji, W.; Chua, T.S. Video as conditional graph hierarchy for multi-granular question answering. Proc. AAAI Conf. Artif. Intell. 2022, 36, 2804–2812. [Google Scholar] [CrossRef] [Scilit]
  213. Zellers, R.; Lu, X.; Hessel, J.; Yu, Y.; Park, J.S.; Cao, J.; Farhadi, A.; Choi, Y. Merlot: Multimodal neural script knowledge models. Adv. Neural Inf. Process. Syst. 2021, 34, 23634–23651. [Google Scholar]
  214. Fu, T.J.; Li, L.; Gan, Z.; Lin, K.; Wang, W.Y.; Wang, L.; Liu, Z. Violet: End-to-end video-language transformers with masked visual-token modeling. arXiv 2021, arXiv:2111.12681. [Google Scholar]
  215. Xu, Y.; Liu, G.; Kompella, R.R.; Huang, T.; Hu, S.; Ilhan, F.; Tekin, S.F.; Yahn, Z.; Liu, L. A Multi-Agent Perception-Action Alliance for Efficient Long Video Reasoning. arXiv 2026, arXiv:2603.14052. [Google Scholar]
  216. Liang, T.; Tan, C.; Xia, B.; Zheng, W.S.; Hu, J.F. Ranking distillation for open-ended video question answering with insufficient labels. In Proceedings of the 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: New York, NY, USA, 2024; pp. 13161–13170. [Google Scholar]
  217. Liu, R.; Liu, Z.; Tang, J.; Ma, Y.; Pi, R.; Zhang, J.; Chen, Q. LongVideoAgent: Multi-Agent Reasoning with Long Videos. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics, San Diego, CA, USA, 2–7 July 2026; pp. 40404–40416. [Google Scholar] [CrossRef] [Scilit]
  218. Cheng, D.; Li, M.; Liu, J.; Guo, Y.; Jiang, B.; Liu, Q.; Chen, X.; Zhao, B. Enhancing long video understanding via hierarchical event-based memory. In Proceedings of the 2025 IEEE International Conference on Multimedia and Expo (ICME); IEEE: New York, NY, USA, 2025; pp. 1–6. [Google Scholar]
  219. Balažević, I.; Shi, Y.; Papalampidi, P.; Chaabouni, R.; Koppula, S.; Hénaff, O.J. Memory consolidation enables long-context video understanding. arXiv 2024, arXiv:2402.05861. [Google Scholar]
  220. Fan, Y.; Lin, W.; Chen, R.; Jiang, H.; Chai, W.; Wang, J.; Wang, K. GAM-agent: Game-theoretic and uncertainty-aware collaboration for complex visual reasoning. Adv. Neural Inf. Process. Syst. 2025, 38, 143569–143629. [Google Scholar]
  221. Menon, S.; Iscen, A.; Nagrani, A.; Weyand, T.; Vondrick, C.; Schmid, C. CAViAR: Critic-Augmented Video Agentic Reasoning. arXiv 2025, arXiv:2509.07680. [Google Scholar]
  222. Bhattacharyya, A.; Panchal, S.; Pourreza, R.; Lee, M.; Madan, P.; Memisevic, R. Look, remember and reason: Grounded reasoning in videos with language models. Proc. Int. Conf. Learn. Represent. 2024, 2024, 7339–7357. [Google Scholar]
  223. Niu, J.; Li, Y.; Miao, Z.; Ge, C.; Zhou, Y.; He, Q.; Dong, X.; Duan, H.; Ding, S.; Qian, R.; et al. OVO-Bench: How Far Are Your Video-LLMs from Real-World Online Video Understanding? In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 11–15 June 2025; pp. 18902–18913. [Google Scholar]
  224. Li, C.; Im, E.W.; Fazli, P. VidHalluc: Evaluating Temporal Hallucinations in Multimodal Large Language Models for Video Understanding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 11–15 June 2025; pp. 13723–13733. [Google Scholar]
  225. Ossowski, T.; Maqbool, D.; Chen, J.; Cai, Z.; Bradshaw, T.; Hu, J. Comma: A communicative multimodal multi-agent benchmark. arXiv 2024, arXiv:2410.07553. [Google Scholar]
  226. Davidson, T.R.; Fourney, A.; Amershi, S.; West, R.; Horvitz, E.; Kamar, E. The Collaboration Gap. arXiv 2025, arXiv:2511.02687. [Google Scholar]
  227. Lowe, R.; Foerster, J.; Boureau, Y.L.; Pineau, J.; Dauphin, Y. On the pitfalls of measuring emergent communication. arXiv 2019, arXiv:1903.05168. [Google Scholar]
Figure 2. Overview of the collaborative video understanding framework. The framework consists of three core components: collaborative functional units, inter-unit communication mechanisms, and collaborative state representations. Given a video input V and an optional task query Q, functional units exchange task-relevant messages m i j t and update the collaborative state S t to produce the final task output.
Figure 2. Overview of the collaborative video understanding framework. The framework consists of three core components: collaborative functional units, inter-unit communication mechanisms, and collaborative state representations. Given a video input V and an optional task query Q, functional units exchange task-relevant messages m i j t and update the collaborative state S t to produce the final task output.
Data 11 00230 g002
Figure 3. Taxonomy of coordination mechanisms for collaborative video understanding, distinguishing static and dynamic collaboration according to whether the coordination policy remains fixed or is adaptively adjusted during execution. Blue boxes denote collaborative functional units.
Figure 3. Taxonomy of coordination mechanisms for collaborative video understanding, distinguishing static and dynamic collaboration according to whether the coordination policy remains fixed or is adaptively adjusted during execution. Blue boxes denote collaborative functional units.
Data 11 00230 g003
Table 1. Comparison of representative video understanding and multimodal-agent surveys and their analytical perspectives.
Table 1. Comparison of representative video understanding and multimodal-agent surveys and their analytical perspectives.
SurveyPrimary ScopeAnalytical PerspectiveCollaboration-Centric
[23]Temporal action segmentationCore techniques and supervision regimes×
[24]Video TransformersPipeline design and architectural adaptations×
[25]Self-supervised learning for videosLearning objectives and training paradigms×
[28]Vision-language pre-trainingVLP design dimensions×
[26]Video-LLMsFramework architectures and LLM functional rolesPartial
[29]Large multimodal agentsAgent components, taxonomy, collaboration, and evaluationPartial
[30]Multimodal Agent AIAgent integration, learning, categorization, and applicationsPartial
OursMulti-model collaborationCollaboration paradigms and mechanisms
Note: “✓”, “Partial”, and “×” indicate primary, partial, and non-central collaboration focus, respectively.
Table 2. Characteristics of representative collaboration paradigms.
Table 2. Characteristics of representative collaboration paradigms.
ParadigmCoordination TypePolicyExecution BehaviorMain Limitation
Static CollaborationPredefinedFixed execution pipelineLimited adaptability
Dynamic CollaborationController-basedAdaptiveRuntime routing, scheduling, or resource allocationPolicy-bound coordination
Agent-basedAgent-drivenPlanning and iterative executionHigh complexity
Table 3. Evidence mapping of representative collaborative video understanding systems to the proposed framework. Four representative studies are selected for each coordination category. Selected metrics are reported for compactness; not all metrics from the original studies are listed.
Table 3. Evidence mapping of representative collaborative video understanding systems to the proposed framework. Four representative studies are selected for each coordination category. Selected metrics are reported for compactness; not all metrics from the original studies are listed.
CoordinationStudyFunctional UnitsCommunication/LoadState/MemoryTasksBenchmarksEvaluation
StaticCollaborative Experts [77]Object, scene, motion, face, audio, ASR, and OCR experts; gatingPairwise feature gating; one-passFused representation; no persistent memoryVideo retrievalMSR-VTT; MSVD; LSMDC; DiDeMo; ActivityNet-CaptionsR@K; MdR; MnR
Video-RAG [78]LVLM, Whisper, EasyOCR, APE, CLIP, Contriever, FAISS; scene graphJSON retrieval requests and evidence; one-passOCR, ASR, and scene-graph databasesLong-video QA; retrievalVideo-MME; MLVU; LongVideoBench; VNBenchAccuracy; token cost; ablation
VideoMulti–Agents [79]Text, video, and scene-graph agents; OrganizerIndependent reports and synthesis; parallel inter-agent executionAgent memory and Organizer contextVideoQANExT-QA; IntentQA; EgoSchemaAccuracy; architecture ablation
MAGNET [80]AV-RAG, salient-frame selector, AV agents, and meta-agentEvidence tuples and meta-agent synthesis; parallel map-reducePrecomputed AV database; aggregation contextAudio-visual multi-video QAAVHaystacks-50; AVHaystacks-FullBLEU/CIDEr; Text Sim; GPT Eval; R@K; MTGS; STEM; human evaluation
Controller-basedMoReVQA [81]Event parsing, grounding, reasoning, and prediction; PaLM-2, PaLI-3, OWL-ViT, CLIPStructured API calls and grounded outputs; staged adaptive executionFrame IDs, event queue, QA/OCR flags, and grounded evidenceVideoQA; grounded VideoQANExT-QA; iVQA; EgoSchema; ActivityNet-QA; NExT-GQAAccuracy; Acc@GQA; mIoP; mIoU
DrVideo [82]Video-document, retrieval, augmentation, planning, interaction, and answer modulesBinary sufficiency feedback and typed frame requests; iterativeAugmented document and analysis historyLong-video QAEgoSchema; MovieChat-1K; Video-MMEAccuracy; Gemini-Pro score; ablation
VideoAgent [17]GPT-4 controller; LaViLa/CogAgent; EVA-CLIPJSON actions, confidence, and captions; multi-roundChronological caption history and retrieved framesLong-form VideoQAEgoSchema; NExT-QAAccuracy; frame count; frame efficiency
Video-EM [69]Query agent, CLIP, TransNetV2, Grounding-DINO, Qwen2.5-VL/Qwen3, downstream Video-LLMsQuery decomposition, narratives, and verification; recursiveEpisodic event memory and scene relationsLong-video QA; narrative reasoningVideo-MME; LVBench; HourVideo; EgoSchemaAccuracy; frame count; runtime; FLOPs; ablation
Agent-basedOmAgent [19]Video2RAG, Conqueror, Divider, Rescuer, Tool Manager, Conclusive SynthesisTyped results, subtasks, and tool outputs; recursive D&CKnowledge database, timestamps, and task treeComplex long-video understandingOmAgent long-video benchmark; MBPP; FreshQAAccuracy; event localization; summary-QA accuracy
TraveLER [18]Planner, retriever, extractor, evaluator, and summarizerPlans, timestamps, QA, and feedback; iterative replanningTimestamp-keyed caption/QA memoryZero-shot VideoQANExT-QA; EgoSchema; STAR; Perception Test; Causal-VidQAAccuracy; frame efficiency; inference cost; ablation
VideoChat-M1 [83]Policy agents, retrieval/perception tools, voting, and lead agentShared policy memory and tool updates; multi-roundPolicies, tool logs, and shared KV bufferLong-video QA; spatial/temporal reasoningLongVideoBench; Video-MME; MLVU; Video-Holmes; VSI-Bench; Charades-STAAccuracy; M-avg/G-avg; category accuracy; mIoU; latency
HiCrew [84]Problem/task planning; text, visual, evidence, and answer agentsWorkflow specifications, confidence, and tool outputs; ReActHybridTree and task stateLong-form VideoQAEgoSchema; NExT-QAAccuracy; role/planning/HybridTree ablations
Table 4. Representative datasets for major video understanding tasks.
Table 4. Representative datasets for major video understanding tasks.
TaskDatasetsSizeSourcesRemarks
R&LKinetics [10]240K clipsCVPR 2017Broad-scale benchmark evaluating generic human action recognition.
Something-Something V2 [133]220.8K videosICCV 2017Fine-grained benchmark emphasizing temporal dynamics and object interaction understanding.
THUMOS14 [134]18.4K videosCVIU 2016Benchmark supporting action recognition and temporal action localization
ActivityNet v1.3 [135]19,994 videosCVPR 2015Large-scale untrimmed-video benchmark for activity recognition and temporal action localization.
AVA [136]437 clipsCVPR 2018Dense spatio-temporal benchmark for localized atomic human action recognition.
CMGMSR-VTT [137]10K clipsCVPR 2016Open-domain video-language benchmark for cross-modal retrieval and alignment.
DiDeMo [138]10.5K videosICCV 2017Benchmark for natural-language moment localization in unedited personal videos.
Charades-STA [139]10K videosICCV 2017Sentence-grounded temporal localization benchmark built on indoor activity videos.
TACoS [140]127 videosTACL 2013Early cooking-domain benchmark for language-to-video grounding.
ActivityNet Captions [141]20K videosICCV 2017Dense event benchmark for temporal grounding with natural-language queries.
QVHighlights [142]10.1K videosNeurIPS 2021Query-conditioned benchmark with joint moment retrieval and highlight detection.
NExT-GQA [143]1.6K videosCVPR 2024Visually grounded VideoQA benchmark extending NExT-QA with temporal grounding annotations.
GDMSVD [144]2.1K clipsACL 2011Early open-domain benchmark for short-video caption generation.
MSR-VTT [137]10K clipsCVPR 2016Large open-domain benchmark for captioning diverse web videos.
VATEX [145]41.3K videosICCV 2019Large-scale multilingual video-language dataset with English and Chinese captions.
ActivityNet Captions [141]20K videosICCV 2017Multi-event captioning benchmark with temporally localized descriptions.
YouCook2 [146]2K videosAAAI 2018Procedural video benchmark with step-aligned instructional narrations.
TVSum [147]50 videosCVPR 2015Topic-diverse benchmark for supervised shot-level video summarization.
RDTGIF-QA [148]71.7K GIFsCVPR 2017GIF-based VideoQA benchmark for temporal reasoning, counting, and transition understanding.
MSR-VTT-QA [149]10K videosPR 2020VideoQA benchmark constructed from MSR-VTT videos with large-scale QA annotations.
NExT-QA [150]5.4K videosCVPR 2021VideoQA benchmark emphasizing causal and temporal reasoning.
AGQA [151]9.6K videosCVPR 2021Compositional reasoning benchmark built on structured action graphs.
STAR [152]22K clipsNeurIPS 2021Situated reasoning benchmark covering interaction, sequence, prediction, and feasibility.
Video-MME [132]900 videosCVPR 2025Long-context multimodal video understanding benchmark with visual, audio, and subtitle information.
EgoSchema [153]5K QA pairsNeurIPS 2023Long-form egocentric VideoQA benchmark with five-way multiple-choice questions over three-minute clips.
LongVideoBench [154]3.8K videosNeurIPS 2024Long-context video-language benchmark with interleaved video and subtitle inputs for referring reasoning over videos up to one hour.
Note: R&L, CMG, GD, and RD denote Recognition and Localization, Cross-modal Matching and Grounding, Generative Description, and Reasoning and Decision-making, respectively. For THUMOS14, the reported 18.4K videos refer to the complete challenge collection, while the standard temporal action localization setting uses 413 temporally annotated untrimmed videos, including 200 validation videos for training and 213 test videos for evaluation [155].
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Chen, Y.; Zhang, J.; Zhang, L.; Liu, C.; Gao, R.; Lu, Z.; Qi, J.; Lei, Q. A Survey of Multi-Model Collaboration in Video Understanding. Data 2026, 11, 230. https://doi.org/10.3390/data11090230

AMA Style

Chen Y, Zhang J, Zhang L, Liu C, Gao R, Lu Z, Qi J, Lei Q. A Survey of Multi-Model Collaboration in Video Understanding. Data. 2026; 11(9):230. https://doi.org/10.3390/data11090230

Chicago/Turabian Style

Chen, Yi, Jianwei Zhang, Lei Zhang, Chang Liu, Rui Gao, Zhixian Lu, Jun Qi, and Qiyu Lei. 2026. "A Survey of Multi-Model Collaboration in Video Understanding" Data 11, no. 9: 230. https://doi.org/10.3390/data11090230

APA Style

Chen, Y., Zhang, J., Zhang, L., Liu, C., Gao, R., Lu, Z., Qi, J., & Lei, Q. (2026). A Survey of Multi-Model Collaboration in Video Understanding. Data, 11(9), 230. https://doi.org/10.3390/data11090230

Article Metrics

Back to TopTop