Next Article in Journal
An Investigation of Three-Dimensional Void Changes and Top-Down Microcrack Formation of AC-16 in Rutted and Non-Rutted Zones Under Extremely High Temperature and Heavy Load
Previous Article in Journal
Low-Resourced Alphabet-Level Pivot-Based Neural Machine Translation for Translating Korean Dialects
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

V-PRUNE: Semantic-Aware Patch Pruning Before Tokenization in Vision–Language Model Inference

Department of Computer Science, Hanyang University, Seoul 04763, Republic of Korea
*
Author to whom correspondence should be addressed.
Appl. Sci. 2025, 15(17), 9463; https://doi.org/10.3390/app15179463
Submission received: 1 August 2025 / Revised: 25 August 2025 / Accepted: 27 August 2025 / Published: 28 August 2025

Abstract

Recent vision–language models (VLMs) achieve strong performance across multimodal benchmarks but suffer from high inference costs due to the large number of visual tokens. Prior studies have shown that many image tokens receive consistently low attention scores during inference, indicating that a substantial portion of visual content contributes little to final predictions. These observations raise questions about the efficiency of conventional token pruning strategies, which are typically applied after all attention operations and depend on late-emerging attention scores. To address this, we propose V-PRUNE, a semantic-aware patch-level pruning framework for vision–language models that removes redundant content before tokenization. By evaluating local similarity via color and histogram statistics, our method enables lightweight and interpretable pruning without architectural changes. Applied to CLIP-based models, our approach reduces FLOPs and inference time across vision–language understanding tasks, while maintaining or improving accuracy. Qualitative results further confirm that essential regions are preserved and the pruning behavior is human-aligned, making our method a practical solution for efficient VLM inference.
Keywords: vision–language models; efficient vision transformers; feature pruning; visual question answering vision–language models; efficient vision transformers; feature pruning; visual question answering

Share and Cite

MDPI and ACS Style

Seo, H.; Choi, Y.S. V-PRUNE: Semantic-Aware Patch Pruning Before Tokenization in Vision–Language Model Inference. Appl. Sci. 2025, 15, 9463. https://doi.org/10.3390/app15179463

AMA Style

Seo H, Choi YS. V-PRUNE: Semantic-Aware Patch Pruning Before Tokenization in Vision–Language Model Inference. Applied Sciences. 2025; 15(17):9463. https://doi.org/10.3390/app15179463

Chicago/Turabian Style

Seo, Hyein, and Yong Suk Choi. 2025. "V-PRUNE: Semantic-Aware Patch Pruning Before Tokenization in Vision–Language Model Inference" Applied Sciences 15, no. 17: 9463. https://doi.org/10.3390/app15179463

APA Style

Seo, H., & Choi, Y. S. (2025). V-PRUNE: Semantic-Aware Patch Pruning Before Tokenization in Vision–Language Model Inference. Applied Sciences, 15(17), 9463. https://doi.org/10.3390/app15179463

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop