Next Article in Journal
Feature Extraction Approach for Distributed Wind Power Generation Based on Power System Flexibility Planning Analysis
Previous Article in Journal
WhistleGAN for Biomimetic Underwater Acoustic Covert Communication
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

RCAT: Retentive CLIP Adapter Tuning for Improved Video Recognition

1
Information Engineering College, Capital Normal University, 56 West 3rd Ring North Road, Beijing 100048, China
2
School of Cyberspace Security, Hainan University, 58 Renmin Avenue, Haikou 570228, China
*
Author to whom correspondence should be addressed.
Electronics 2024, 13(5), 965; https://doi.org/10.3390/electronics13050965
Submission received: 21 January 2024 / Revised: 27 February 2024 / Accepted: 29 February 2024 / Published: 2 March 2024

Abstract

The advent of Contrastive Language-Image Pre-training (CLIP) models has revolutionized the integration of textual and visual representations, significantly enhancing the interpretation of static images. However, their application to video recognition poses unique challenges due to the inherent dynamism and multimodal nature of video content, which includes temporal changes and spatial details beyond the capabilities of traditional CLIP models. These challenges necessitate an advanced approach capable of comprehending the complex interplay between the spatial and temporal dimensions of video data. To this end, this study introduces an innovative approach, Retentive CLIP Adapter Tuning (RCAT), which synergizes the foundational strengths of CLIP with the dynamic processing prowess of a Retentive Network (RetNet). Specifically designed to refine CLIP’s applicability to video recognition, RCAT facilitates a nuanced understanding of video sequences by leveraging temporal analysis. At the core of RCAT is its specialized adapter tuning mechanism, which modifies the CLIP model to better align with the temporal intricacies and spatial details of video content, thereby enhancing the model’s predictive accuracy and interpretive depth. Our comprehensive evaluations on benchmark datasets, including UCF101, HMDB51, and MSR-VTT, underscore the effectiveness of RCAT. Our proposed approach achieves notable accuracy improvements of 1.4% on UCF101, 2.6% on HMDB51, and 1.1% on MSR-VTT compared to existing models, illustrating its superior performance and adaptability in the context of video recognition tasks.
Keywords: video recognition; clip adapter tuning; retentive network; video temporal analysis; hard prompt video recognition; clip adapter tuning; retentive network; video temporal analysis; hard prompt

Share and Cite

MDPI and ACS Style

Xie, Z.; Xu, M.; Zhang, S.; Zhou, L. RCAT: Retentive CLIP Adapter Tuning for Improved Video Recognition. Electronics 2024, 13, 965. https://doi.org/10.3390/electronics13050965

AMA Style

Xie Z, Xu M, Zhang S, Zhou L. RCAT: Retentive CLIP Adapter Tuning for Improved Video Recognition. Electronics. 2024; 13(5):965. https://doi.org/10.3390/electronics13050965

Chicago/Turabian Style

Xie, Zexun, Min Xu, Shudong Zhang, and Lijuan Zhou. 2024. "RCAT: Retentive CLIP Adapter Tuning for Improved Video Recognition" Electronics 13, no. 5: 965. https://doi.org/10.3390/electronics13050965

APA Style

Xie, Z., Xu, M., Zhang, S., & Zhou, L. (2024). RCAT: Retentive CLIP Adapter Tuning for Improved Video Recognition. Electronics, 13(5), 965. https://doi.org/10.3390/electronics13050965

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop