- Article
21 Pages
Fine-grained visual classification (FGVC) requires models to distinguish subtle inter-class differences while remaining robust to substantial intra-class variations in pose and background. Hierarchical visual backbones improve semantic abstraction but progressively compress local textures and part boundaries, while global pooling may further weaken discriminative evidence distributed across multiple regions. To address these issues, we propose AMR-VMamba, a lightweight enhancement framework built on VMamba-Tiny. First, adaptive multi-level refinement (AMR) aligns stage 3 and stage 4 features and selectively injects cross-level differences through position-channel-dependent gating, thereby complementing high-level semantics with mid-level details. Second, multi-query discriminative pooling (MQP) uses a small set of learnable queries to extract multiple local descriptors from the refined 7 × 7 feature grid and combines them with the native pooled representation. Zero-initialized learnable residual coefficients reduce the initial perturbation of the pretrained representation. Across five independent runs on CUB-200-2011, Stanford Dogs, and Oxford Flowers-102, AMR-VMamba achieves Top-1 accuracies of 86.98 ± 0.18%, 88.51 ± 0.26%, and 96.32 ± 0.23%, improving the VMamba-Tiny baseline by 3.05, 0.91, and 1.60 percentage points, respectively, with only 0.5 M additional parameters and 0.02 GMACs.
J. Imaging
22 September 2026









