Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

Article Types

Countries / Regions

Search Results (85)

Search Parameters:
Keywords = keyframes extraction

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
35 pages, 25559 KB  
Article
Training-Free Skeleton-Semantic Keyframe Extraction for Fixed-View Industrial Assembly Video: Method Design and Case Study Evaluation
by Qianhui Li, Hua Xiang, Tongxi Wang and Wei Zhan
Sensors 2026, 26(15), 4956; https://doi.org/10.3390/s26154956 - 5 Aug 2026
Viewed by 315
Abstract
Industrial assembly videos contain substantial temporal redundancy, creating a need for preprocessing methods that remain usable when task-specific labels, retraining, or opaque learned selection rules are undesirable. This study designs a deterministic, training-free skeleton-semantic keyframe extraction framework for fixed-view monocular assembly video preprocessing. [...] Read more.
Industrial assembly videos contain substantial temporal redundancy, creating a need for preprocessing methods that remain usable when task-specific labels, retraining, or opaque learned selection rules are undesirable. This study designs a deterministic, training-free skeleton-semantic keyframe extraction framework for fixed-view monocular assembly video preprocessing. OpenPose BODY-25 keypoints are confidence-filtered, normalized, interpolated, and represented by eight upper-body joints. Neck-relative bilateral-wrist activity controls an inverse threshold, while sequential reference updates and a maximum-gap rule provide an explicit output budget. At 271 frames per analyzed file, the method obtains CVsd = 0.3340 ± 0.1117 and MSD = 0.7781 ± 0.1003, compared with 1.0681 ± 0.0754 and 0.2961 ± 0.0305 for pretrained MViT-V2-S feature clustering. Thus, under the shared normalized skeleton evaluation, the method reduces semantic-distance variation by 68.7% and increases adjacent pose-state separation by 162.8% relative to MViT. Matched ablation identifies the skeleton representation as the primary source of this organization; adaptive thresholding provides smaller, budget-dependent density adjustment. A second independent human annotation and adjudication establish a 34-transition event reference. Event results reveal a complementary trade-off: the proposed method retains 79.41% of transitions within ±0.5 s, whereas several temporal and appearance-based baselines retain 97.06–100%, showing that pose-state diversity and process-boundary coverage are different sampling objectives. Matched-baseline observed-joint-only and confidence-weighted analyses show that interpolation contributes to numerical regularity, but the proposed method retains substantially lower semantic-distance variation and greater adjacent pose-state separation than all evaluated baselines under both missing-data treatments. These results establish the method as an interpretable, budget-controllable preprocessing strategy for pose-oriented review and data reduction in the evaluated fixed-view setting. Validation across independently recorded production environments and task-specific downstream systems is the next step toward broader deployment. Full article
(This article belongs to the Section Industrial Sensors)
Show Figures

Figure 1

25 pages, 24261 KB  
Article
Lightweight 2.5D SLAM with Dynamic Map Refinement and Height-Aware Encoding for Resource-Constrained Indoor Robots
by Guitao Yu, Yuping Zhang, Zhiao Qi, Kui Yang, Yang He and Dongtai Liang
Sensors 2026, 26(15), 4765; https://doi.org/10.3390/s26154765 - 27 Jul 2026
Viewed by 384
Abstract
Indoor mobile robots equipped with low-cost and sparse sensors often suffer from limited vertical perception and dynamic residual artifacts in the final map. This paper presents a lightweight 2.5D simultaneous localization and mapping (SLAM) framework using a single-line laser distance sensor (LDS), time-of-flight [...] Read more.
Indoor mobile robots equipped with low-cost and sparse sensors often suffer from limited vertical perception and dynamic residual artifacts in the final map. This paper presents a lightweight 2.5D simultaneous localization and mapping (SLAM) framework using a single-line laser distance sensor (LDS), time-of-flight (ToF) sensing, wheel odometry, and an inertial measurement unit (IMU). In this work, 2.5D refers to a 2D grid map with discretized vertical occupancy bins for each grid cell, rather than a full continuous 3D reconstruction. The system integrates multi-sensor synchronization, motion correction, error-state Kalman filter (ESKF)-based state estimation, normal distributions transform (NDT) registration, and pose graph optimization to reconstruct a pose-consistent global map. Based on this map, an offline dynamic refinement module estimates temporal voxel support across keyframes, extracts low-support candidate regions, and applies geometric clustering and isolated-point filtering to suppress transient residual artifacts while preserving stable structures. A 24-bit RGB occupancy encoding is further proposed to store the discretized vertical occupancy state in a compact three-channel image format. The proposed framework emphasizes system-level deployment value by combining sparse multi-sensor mapping, conservative offline refinement, and compact height-aware map export on a low-cost indoor robot platform. Experiments on public datasets, embedded hardware, and self-collected indoor sequences evaluate odometry reference performance, resource usage, platform-specific 2.5D mapping, dynamic refinement, and height-aware encoding. Full article
(This article belongs to the Section Sensors and Robotics)
Show Figures

Figure 1

19 pages, 4439 KB  
Article
An Algorithm for Fine-Grained Content Extraction and Understanding in Short Videos
by Yanqi Wan, Shuya Zhang, Yi Xu, Kunfang Zhang, Heyi Wang and Mingzheng Liu
Data 2026, 11(7), 179; https://doi.org/10.3390/data11070179 - 20 Jul 2026
Viewed by 533
Abstract
We integrated communication theory with advanced computer vision techniques to propose a novel approach for fine-grained content extraction from short videos. Unlike methods focused on summarization or subtitle generation for longer videos, our approach emphasizes extracting detailed content and understanding the intricate narrative [...] Read more.
We integrated communication theory with advanced computer vision techniques to propose a novel approach for fine-grained content extraction from short videos. Unlike methods focused on summarization or subtitle generation for longer videos, our approach emphasizes extracting detailed content and understanding the intricate narrative structure of short videos. By employing scene segmentation, similarity-based filtering algorithms, and support vector machines, the method identifies keyframes that capture precise visual details. Further, it generates semantically accurate textual descriptions using the mPLUG model, enabling an in-depth understanding of video content. Using a dataset of short videos from the cultural and tourism domain, we validated the proposed method. Experimental results demonstrate that our approach achieves high precision in identifying and understanding detailed visual elements, effectively bridging the gap between visual representation and semantic meaning. Additionally, the study explores the influence of different video content types, interference factors, and image description models on fine-grained content extraction, highlighting its potential for improving intelligent analysis of short-video data. Full article
(This article belongs to the Special Issue Vision-Based AI in the Real World: Data, Robustness and Deployment)
Show Figures

Figure 1

46 pages, 3026 KB  
Article
Keyframe Selection and Multimodal Fusion for Product Recognition in E-Commerce Live Streaming
by Yichuan Zheng, Jin Shi and Wei Shen
Appl. Sci. 2026, 16(13), 6585; https://doi.org/10.3390/app16136585 - 1 Jul 2026
Viewed by 453
Abstract
Product recognition in e-commerce live streaming is hindered by rapid viewpoint changes, occlusions, motion blur, and inconsistencies between visual and spoken information. Existing approaches typically focus on individual components such as detection, OCR, or speech recognition, which limits their effectiveness in end-to-end structured [...] Read more.
Product recognition in e-commerce live streaming is hindered by rapid viewpoint changes, occlusions, motion blur, and inconsistencies between visual and spoken information. Existing approaches typically focus on individual components such as detection, OCR, or speech recognition, which limits their effectiveness in end-to-end structured product understanding. To address this problem, we propose an integrated framework that combines task-oriented keyframe selection with multimodal semantic fusion. The framework first uses D-FINE to localize product regions and then selects informative frames through two complementary strategies. Strategy A considers both detection confidence and Laplacian-based sharpness, while Strategy B combines detection confidence with a learned quality component estimated by an EfficientNetV2-M regression model. OCR, visual-semantic recognition, and ASR are then applied to extract complementary evidence, and a Qwen3.5-27B large language model is used to structure and fuse multimodal evidence into standardized product outputs, including brand, product name, and category. Experiments on an in-house e-commerce livestreaming dataset demonstrate substantial gains over a last-frame baseline. Strategy B achieves the best overall result, improving the Perfect Match Rate from 0.609 to 0.775 and the Semantic Similarity from 0.697 to 0.802. Ablation studies further show that the full multimodal framework consistently outperforms unimodal and dual-modality variants under both frame selection strategies. In addition, Top-K analysis indicates that single-frame inference provides a practical balance between OCR evidence completeness and efficiency. Efficiency analysis shows that the per-video API monetary cost remains low under the pricing configuration used in this study, while API latency is mainly limited by Qwen3.5-27B LLM calls for evidence structuring and final fusion. Overall, the proposed framework offers an effective and extensible solution for structured product recognition in complex live-streaming scenarios. Full article
Show Figures

Figure 1

19 pages, 1217 KB  
Article
Talking with Actionbits—A Part-Enhanced VLM for Action and Interaction Recognition in Animals
by Yang Yang, Ren Nakagawa, Risa Shinoda, Hiroaki Santo, Kenji Oyama, Takenao Ohkawa and Fumio Okura
Sensors 2026, 26(6), 1969; https://doi.org/10.3390/s26061969 - 21 Mar 2026
Viewed by 767
Abstract
Understanding animal actions and interactions is essential for behavior analysis and ecological monitoring. Although large-scale in-the-wild datasets have advanced animal action recognition, existing methods still struggle with fine-grained motion, spatial relations, and multi-individual interactions. To address these challenges, we introduce AIRA, a unified [...] Read more.
Understanding animal actions and interactions is essential for behavior analysis and ecological monitoring. Although large-scale in-the-wild datasets have advanced animal action recognition, existing methods still struggle with fine-grained motion, spatial relations, and multi-individual interactions. To address these challenges, we introduce AIRA, a unified framework for Action and Interaction Recognition in Animals. Built upon a vision–language model (VLM), AIRA learns in an action-centered representation space defined by body parts and their corresponding motions, thereby improving robustness to background noise and enabling cross-species generalization via a unified mammal-centric part ontology. To model actions, we treat body parts and motion as primary cues and introduce Actionbit tokens—compact representations for parts and motions generated by a large language model (LLM) that encode which parts move and how. We further propose Part-Enhanced Prompt Fine-tuning (PEPF) to make the VLM explicitly sensitive to part and pose cues. Within PEPF, the Action–actionbit Alignment (AbA) module enriches action representations with fine-grained part–motion semantics, and Part-Vision Prompting (PVP) extracts keyframes through action-aware prompting. Experiments across multiple benchmarks show consistent improvements in both action and interaction recognition, highlighting the importance of action-centered adaptation and relational reasoning for understanding animal behavior in the wild. Full article
(This article belongs to the Special Issue Innovative Sensing Methods for Motion and Behavior Analysis)
Show Figures

Figure 1

30 pages, 4114 KB  
Article
TricP: A Novel Approach for Human Activity Recognition Using Tricky Predator Optimization Based on Inception and LSTM
by Palak Girdhar, Muslem Al-Saidi, Prashant Johri, Deepali Virmani, Hussein Taha and Oday Ali Hassen
Telecom 2026, 7(2), 32; https://doi.org/10.3390/telecom7020032 - 19 Mar 2026
Viewed by 949
Abstract
Human Activity Recognition (HAR) is a pivotal research area for applications such as automated surveillance, smart homes, security, healthcare, and human behavior analysis. Traditional machine-learning approaches often rely on manual feature engineering, which can limit generalization. Although deep learning has improved HAR through [...] Read more.
Human Activity Recognition (HAR) is a pivotal research area for applications such as automated surveillance, smart homes, security, healthcare, and human behavior analysis. Traditional machine-learning approaches often rely on manual feature engineering, which can limit generalization. Although deep learning has improved HAR through automatic representation learning, achieving high detection performance under computational constraints remains challenging. This paper proposes an efficient HAR framework that combines deep learning with hybrid optimization. Surveillance videos are first decomposed into frames, and a keyframe selection stage identifies distinctive frames to reduce redundancy and computational cost while preserving informative content. Motion and appearance features are then extracted using Histogram of Oriented Optical Flow (HOOF) and a ResNet-101 model, respectively, and concatenated into a unified feature representation. Classification is performed using an Inception-based Long Short-Term Memory (Incept-LSTM) network, which is fine-tuned via the proposed Tricky Predator Optimization (TricP) over a restricted, low-dimensional parameter vector. TricP is inspired by predator poaching behavior and the social dynamics of Latrans to enhance exploration and exploitation during search. Experiments on the UCF-Crime dataset show that the proposed method achieves 96.84% specificity, 92.16% sensitivity, and 93.62% accuracy. Full article
Show Figures

Figure 1

30 pages, 3812 KB  
Review
Video-Based 3D Reconstruction: A Review of Photogrammetry and Visual SLAM Approaches
by Ali Javadi Moghadam, Abbas Kiani, Reza Naeimaei, Shirin Malihi and Ioannis Brilakis
J. Imaging 2026, 12(3), 128; https://doi.org/10.3390/jimaging12030128 - 13 Mar 2026
Cited by 2 | Viewed by 3585
Abstract
Three-dimensional (3D) reconstruction using images is one of the most significant topics in computer vision and photogrammetry, with wide-ranging applications in robotics, augmented reality, and mapping. This study investigates methods of 3D reconstruction using video (especially monocular video) data and focuses on techniques [...] Read more.
Three-dimensional (3D) reconstruction using images is one of the most significant topics in computer vision and photogrammetry, with wide-ranging applications in robotics, augmented reality, and mapping. This study investigates methods of 3D reconstruction using video (especially monocular video) data and focuses on techniques such as Structure from Motion (SfM), Multi-View Stereo (MVS), Visual Simultaneous Localization and Mapping (V-SLAM), and videogrammetry. Based on a statistical analysis of SCOPUS records, these methods collectively account for approximately 6863 journal publications up to the end of 2024. Among these, about 80 studies are analyzed in greater detail to identify trends and advancements in the field. The study also shows that the use of video data for real-time 3D reconstruction is commonly addressed through two main approaches: photogrammetry-based methods, which rely on precise geometric principles and offer high accuracy at the cost of greater computational demand; and V-SLAM methods, which emphasize real-time processing and provide higher speed. Furthermore, the application of IMU data and other indicators, such as color quality and keypoint detection, for selecting suitable frames for 3D reconstruction is investigated. Overall, this study compiles and categorizes video-based reconstruction methods, emphasizing the critical step of keyframe extraction. By summarizing and illustrating the general approaches, the study aims to clarify and facilitate the entry path for researchers interested in this area. Finally, the paper offers targeted recommendations for improving keyframe extraction methods to enhance the accuracy and efficiency of real-time video-based 3D reconstruction, while also outlining future research directions in addressing challenges like dynamic scenes, reducing computational costs, and integrating advanced learning-based techniques. Full article
(This article belongs to the Section Computer Vision and Pattern Recognition)
Show Figures

Figure 1

17 pages, 1568 KB  
Article
Traffic-Oriented Three-Dimensional Vehicle Reconstruction Using Fixed Roadside Monocular Camera Sensors
by Chu Zhang, Yuxin Zhang, Liangbin Li and Xianhua Cai
Sensors 2026, 26(4), 1324; https://doi.org/10.3390/s26041324 - 18 Feb 2026
Cited by 1 | Viewed by 760
Abstract
Fixed roadside monocular cameras are widely used as low-cost sensing devices in intelligent transportation systems; however, extracting reliable three-dimensional (3D) information from such sensors remains challenging due to limited baselines, long observation distances, and moving vehicles. This paper presents a traffic-oriented 3D vehicle [...] Read more.
Fixed roadside monocular cameras are widely used as low-cost sensing devices in intelligent transportation systems; however, extracting reliable three-dimensional (3D) information from such sensors remains challenging due to limited baselines, long observation distances, and moving vehicles. This paper presents a traffic-oriented 3D vehicle reconstruction framework based on monocular image sequences captured by fixed roadside camera sensors. Semantic and non-semantic vehicle feature points are jointly exploited to balance structural consistency and surface completeness, and a feature-map-consistency-based optimization strategy is introduced to refine feature point localization and reduce reprojection errors. In addition, an optimized incremental Structure-from-Motion (SfM) pipeline incorporating traffic-aware initialization, keyframe selection, and local bundle adjustment is developed to improve reconstruction efficiency. Experiments on real-world traffic surveillance videos show that the proposed method reduces the mean reprojection error by 13.6% and shortens reconstruction time by 43.9% compared with widely used incremental SfM systems. Full article
(This article belongs to the Collection 3D Imaging and Sensing System)
Show Figures

Figure 1

5 pages, 1305 KB  
Proceeding Paper
Audiovisual Fusion Technique for Detecting Sensitive Content in Videos
by Daniel Povedano Álvarez, Ana Lucila Sandoval Orozco and Luis Javier García Villalba
Eng. Proc. 2026, 123(1), 11; https://doi.org/10.3390/engproc2026123011 - 2 Feb 2026
Viewed by 1026
Abstract
The detection of sensitive content in online videos is a key challenge for ensuring digital safety and effective content moderation. This work proposes the Multimodal Audiovisual Attention (MAV-Att), a multimodal deep learning framework that jointly exploits audio and visual cues to improve detection [...] Read more.
The detection of sensitive content in online videos is a key challenge for ensuring digital safety and effective content moderation. This work proposes the Multimodal Audiovisual Attention (MAV-Att), a multimodal deep learning framework that jointly exploits audio and visual cues to improve detection accuracy. The model was evaluated on the LSPD dataset, comprising 52,427 video segments of 20 s each, with optimized keyframe extraction. MAV-Att consists of dual audio and image branches enhanced by attention mechanisms to capture both temporal and cross-modal dependencies. Trained using a joint optimisation loss, the system achieved F1-scores of 94.9% on segments and 94.5% on entire videos, surpassing previous state-of-the-art models by 6.75%. Full article
(This article belongs to the Proceedings of First Summer School on Artificial Intelligence in Cybersecurity)
Show Figures

Figure 1

49 pages, 6627 KB  
Article
LEARNet: A Learning Entropy-Aware Representation Network for Educational Video Understanding
by Chitrakala S, Nivedha V V and Niranjana S R
Entropy 2026, 28(1), 3; https://doi.org/10.3390/e28010003 - 19 Dec 2025
Viewed by 1588
Abstract
Educational videos contain long periods of visual redundancy, where only a few frames convey meaningful instructional information. Conventional video models, which are designed for dynamic scenes, often fail to capture these subtle pedagogical transitions. We introduce LEARNet, an entropy-aware framework that models educational [...] Read more.
Educational videos contain long periods of visual redundancy, where only a few frames convey meaningful instructional information. Conventional video models, which are designed for dynamic scenes, often fail to capture these subtle pedagogical transitions. We introduce LEARNet, an entropy-aware framework that models educational video understanding as the extraction of high-information instructional content from low-entropy visual streams. LEARNet combines a Temporal Information Bottleneck (TIB) for selecting pedagogically significant keyframes with a Spatial–Semantic Decoder (SSD) that produces fine-grained annotations refined through a proposed Relational Consistency Verification Network (RCVN). This architecture enables the construction of EVUD-2M, a large-scale benchmark with multi-level semantic labels for diverse instructional formats. LEARNet achieves substantial redundancy reduction (70.2%) while maintaining high annotation fidelity (F1 = 0.89, mAP@50 = 0.88). Grounded in information-theoretic principles, LEARNet provides a scalable foundation for tasks such as lecture indexing, visual content summarization, and multimodal learning analytics. Full article
Show Figures

Figure 1

24 pages, 22793 KB  
Article
GL-VSLAM: A General Lightweight Visual SLAM Approach for RGB-D and Stereo Cameras
by Xu Li, Tuanjie Li, Yulin Zhang, Ziang Li, Lixiang Ban and Yuming Ning
Sensors 2025, 25(24), 7467; https://doi.org/10.3390/s25247467 - 8 Dec 2025
Cited by 1 | Viewed by 1303
Abstract
Feature-based indirect SLAM is more robust than direct SLAM; however, feature extraction and descriptor computation are time-consuming. In this paper, we propose GL-VSLAM, a general lightweight visual SLAM approach designed for RGB-D and stereo cameras. GL-VSLAM utilizes sparse optical flow matching based on [...] Read more.
Feature-based indirect SLAM is more robust than direct SLAM; however, feature extraction and descriptor computation are time-consuming. In this paper, we propose GL-VSLAM, a general lightweight visual SLAM approach designed for RGB-D and stereo cameras. GL-VSLAM utilizes sparse optical flow matching based on uniform motion model prediction to establish keypoint correspondences between consecutive frames, rather than relying on descriptor-based feature matching, thereby achieving high real-time performance. To enhance positioning accuracy, we adopt a coarse-to-fine strategy for pose estimation in two stages. In the first stage, the initial camera pose is estimated using RANSAC PnP based on robust keypoint correspondences from sparse optical flow. In the second stage, the camera pose is further refined by minimizing the reprojection error. Keypoints and descriptors are extracted from keyframes for backend optimization and loop closure detection. We evaluate our system on the TUM and KITTI datasets, as well as in a real-world environment, and compare it with several state-of-the-art methods. Experimental results demonstrate that our method achieves comparable positioning accuracy, while its efficiency is up to twice that of ORB-SLAM2. Full article
Show Figures

Figure 1

22 pages, 1773 KB  
Article
ACE-Net: A Fine-Grained Deepfake Detection Model with Multimodal Emotional Consistency
by Shaoqian Yu, Xingyu Chen, Yuzhe Sheng, Han Zhang, Xinlong Li and Sijia Yu
Electronics 2025, 14(22), 4420; https://doi.org/10.3390/electronics14224420 - 13 Nov 2025
Cited by 1 | Viewed by 1564
Abstract
The alarming realism of Deepfake presents a significant challenge to digital authenticity, yet its inherent difficulty in synchronizing the emotional cues between facial expressions and speech offers a critical opportunity for detection. However, most existing approaches rely on general-purpose backbones for unimodal feature [...] Read more.
The alarming realism of Deepfake presents a significant challenge to digital authenticity, yet its inherent difficulty in synchronizing the emotional cues between facial expressions and speech offers a critical opportunity for detection. However, most existing approaches rely on general-purpose backbones for unimodal feature extraction, resulting in an inadequate representation of fine-grained dynamic emotional expressions. Although a limited number of studies have explored cross-modal emotional consistency of deepfake detection, they typically employ shallow fusion techniques which limit latent expressiveness. To address this, we propose ACE-Net, a novel framework that identifies forgeries via multimodal emotional inconsistency. For the speech modality, we design a bidirectional cross-attention mechanism to fuse acoustic features from a lightweight CNN-based model with textual features, yielding a representation highly sensitive to fine-grained emotional dynamics. For the visual modality, a MobileNetV3-based perception head is proposed to adaptively select keyframes, yielding a representation focused on the most emotionally salient moments. For multimodal emotional consistency discrimination, we develop a multi-dimensional fusion strategy to deeply integrate high-level emotional features from different modalities within a unified latent space. For unimodal emotion recognition, both the audio and visual branches outperform baseline models on the CREMA-D dataset. Building on this, the complete ACE-Net model achieves a state-of-the-art AUC of 0.921 on the challenging DFDC benchmark. Full article
(This article belongs to the Special Issue Computer Vision and Pattern Recognition Based on Machine Learning)
Show Figures

Figure 1

22 pages, 1770 KB  
Article
Key-Frame-Aware Hierarchical Learning for Robust Gait Recognition
by Ke Wang and Hua Huo
J. Imaging 2025, 11(11), 402; https://doi.org/10.3390/jimaging11110402 - 10 Nov 2025
Cited by 1 | Viewed by 932
Abstract
Gait recognition in unconstrained environments is severely hampered by variations in view, clothing, and carrying conditions. To address this, we introduce HierarchGait, a key-frame-aware hierarchical learning framework. Our approach uniquely integrates three complementary modules: a TemplateBlock-based Motion Extraction (TBME) for coarse-to-fine anatomical feature [...] Read more.
Gait recognition in unconstrained environments is severely hampered by variations in view, clothing, and carrying conditions. To address this, we introduce HierarchGait, a key-frame-aware hierarchical learning framework. Our approach uniquely integrates three complementary modules: a TemplateBlock-based Motion Extraction (TBME) for coarse-to-fine anatomical feature learning, a Sequence-Level Spatio-temporal Feature Aggregator (SSFA) to identify and prioritize discriminative key-frames, and a Frame-level Feature Re-segmentation Extractor (FFRE) to capture fine-grained motion details. This synergistic design yields a robust and comprehensive gait representation. We demonstrate the superiority of our method through extensive experiments. On the highly challenging CASIA-B dataset, HierarchGait achieves new state-of-the-art average Rank-1 accuracies of 98.1% under Normal (NM), 95.9% under Bag (BG), and 87.5% under Coat (CL) conditions. Furthermore, on the large-scale OU-MVLP dataset, our model attains a 91.5% average accuracy. These results validate the significant advantage of explicitly modeling anatomical hierarchies and temporal key-moments for robust gait recognition. Full article
(This article belongs to the Section Biometrics, Forensics, and Security)
Show Figures

Figure 1

18 pages, 8879 KB  
Article
Energy-Conscious Lightweight LiDAR SLAM with 2D Range Projection and Multi-Stage Outlier Filtering for Intelligent Driving
by Chun Wei, Tianjing Li and Xuemin Hu
Computation 2025, 13(10), 239; https://doi.org/10.3390/computation13100239 - 10 Oct 2025
Viewed by 1192
Abstract
To meet the increasing demands of energy efficiency and real-time performance in autonomous driving systems, this paper presents a lightweight and robust LiDAR SLAM framework designed with power-aware considerations. The proposed system introduces three core innovations. First, it replaces traditional ordered point cloud [...] Read more.
To meet the increasing demands of energy efficiency and real-time performance in autonomous driving systems, this paper presents a lightweight and robust LiDAR SLAM framework designed with power-aware considerations. The proposed system introduces three core innovations. First, it replaces traditional ordered point cloud indexing with a 2D range image projection, significantly reducing memory usage and enabling efficient feature extraction with curvature-based criteria. Second, a multi-stage outlier rejection mechanism is employed to enhance feature robustness by adaptively filtering occluded and noisy points. Third, we propose a dynamically filtered local mapping strategy that adjusts keyframe density in real time, ensuring geometric constraint sufficiency while minimizing redundant computation. These components collectively contribute to a SLAM system that achieves high localization accuracy with reduced computational load and energy consumption. Experimental results on representative autonomous driving datasets demonstrate that our method outperforms existing approaches in both efficiency and robustness, making it well-suited for deployment in low-power and real-time scenarios within intelligent transportation systems. Full article
(This article belongs to the Special Issue Object Detection Models for Transportation Systems)
Show Figures

Figure 1

24 pages, 4488 KB  
Review
Advances in Facial Micro-Expression Detection and Recognition: A Comprehensive Review
by Tian Shuai, Seng Beng, Fatimah Binti Khalid and Rahmita Wirza Bt O. K. Rahmat
Information 2025, 16(10), 876; https://doi.org/10.3390/info16100876 - 9 Oct 2025
Cited by 11 | Viewed by 9305
Abstract
Micro-expressions are facial movements with extremely short duration and small amplitude, which can reveal an individual’s potential true emotions and have important application value in public safety, medical diagnosis, psychotherapy and business negotiations. Since micro-expressions change rapidly and are difficult to detect, manual [...] Read more.
Micro-expressions are facial movements with extremely short duration and small amplitude, which can reveal an individual’s potential true emotions and have important application value in public safety, medical diagnosis, psychotherapy and business negotiations. Since micro-expressions change rapidly and are difficult to detect, manual recognition is a significant challenge, so the development of automatic recognition systems has become a research hotspot. This paper reviews the development history and research status of micro-expression recognition and systematically analyzes the two main branches of micro-expression analysis: micro-expression detection and micro-expression recognition. In terms of detection, the methods are divided into three categories based on time features, feature changes and deep features according to different feature extraction methods; in terms of recognition, traditional methods based on texture and optical flow features, as well as deep learning-based methods that have emerged in recent years, including motion unit, keyframe and transfer learning strategies, are summarized. This paper also summarizes commonly used micro-expression datasets and facial image preprocessing techniques and evaluates and compares mainstream methods through multiple experimental indicators. Although significant progress has been made in this field in recent years, it still faces challenges such as data scarcity, class imbalance and unstable recognition accuracy. Future research can further combine multimodal emotional information, enhance data generalization capabilities, and optimize deep network structures to promote the widespread application of micro-expression recognition in practical scenarios. Full article
Show Figures

Figure 1

Back to TopTop