Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

Article Types

Countries / Regions

Search Results (37)

Search Parameters:
Keywords = subtask segmentation

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
11 pages, 746 KB  
Article
Validation of a Virtual Reality-Based Timed Up-and-Go Test Using Body-Worn Motion Trackers
by Brooke E. Peters, Sherry Law, Simon Dugré-Rezun, Alana Gullison and Chris A. McGibbon
Sensors 2026, 26(12), 3669; https://doi.org/10.3390/s26123669 - 8 Jun 2026
Viewed by 504
Abstract
Background: Full-body motion capture using commercial virtual reality (VR) systems offers unique opportunities for augmenting common functional assessment such as the timed up-and-go (TUG) test. The purpose of this study was to determine whether task performance (chair, walk, turns) during a VR version [...] Read more.
Background: Full-body motion capture using commercial virtual reality (VR) systems offers unique opportunities for augmenting common functional assessment such as the timed up-and-go (TUG) test. The purpose of this study was to determine whether task performance (chair, walk, turns) during a VR version (vTUG) is parametrically equivalent to the standard test (sTUG). Methods: Twenty healthy adult participants (age 19–71 years) were evaluated with the sTUG followed by the vTUG version of the same test. Body trackers were used to capture kinematics during both tests. TUG time was measured manually with a stopwatch. Tracker data were used to automatically quantify total TUG time and sub-task times for chair, walk and turn portions. Absolute agreement was evaluated using Intraclass Correlation Coefficient (ICC(2,k)) and Bland–Altman analysis. A custom survey was used to evaluate user satisfaction. Results: Very good agreement (ICC > 0.8) was found between sTUG and vTUG for manual and automated measures of total time. ICCs for sub-task times were acceptable (ICC > 0.7) for chair rise, walks and first turn but less so for second turn and sit (ICC < 0.7). User satisfaction was high, and there were no adverse events. Interpretation: The vTUG and sTUG are parametrically equivalent, though sub-task segmentation may require more research. Nevertheless, VR body trackers are a value-added feature whether used with the vTUG or the sTUG and warrant further investigation. Full article
Show Figures

Figure 1

28 pages, 9019 KB  
Article
SAF-SD: Self-Distillation Object Segmentation Method Based on Sequential Three-Way Mask and Attention Fusion
by Biao Wang, Jun Su, Volodymyr Kochan and Lingyu Yan
Sensors 2026, 26(7), 2170; https://doi.org/10.3390/s26072170 - 31 Mar 2026
Viewed by 555
Abstract
Transformer models have achieved powerful performance in various computer vision tasks. However, their black-box nature severely limits model interpretability and the reliability of real-world applications. Most existing interpretation methods generate explanation maps by perturbing masks from the last layer of the Transformer encoder, [...] Read more.
Transformer models have achieved powerful performance in various computer vision tasks. However, their black-box nature severely limits model interpretability and the reliability of real-world applications. Most existing interpretation methods generate explanation maps by perturbing masks from the last layer of the Transformer encoder, but they often overlook uncertain information in masks and detail loss during upsampling and downsampling, resulting in coarse localization, blurred boundaries, and significant background noise in explanations. To address these issues, this paper proposes a self-distillation object segmentation method based on sequential three-way mask and attention fusion (SAF-SD), targeting salient and camouflaged binary object segmentation tasks (sub-tasks of binary pixel-level segmentation). The method consists of two core modules: the sequential three-way mask (S3WM) module and the attention fusion (AF) module. The S3WM module performs strict threshold filtering on masks generated from the final-layer feature maps of the Transformer, aiming to accurately segment foreground objects from backgrounds via binary pixel-level prediction. The AF module aggregates attention matrices across all Transformer encoder layers to construct a cross-layer relation matrix, capturing global semantic dependencies among image patches (e.g., interactions between foreground, background, and edge regions). It then computes the importance score for each patch, refining details and suppressing noise in the initial explanation results. Extensive experimental results demonstrate that SAF-SD significantly outperforms existing baseline methods across key evaluation metrics. Full article
Show Figures

Figure 1

12 pages, 9302 KB  
Article
Robust Vision-Language-Action Models via Object-Centric Learning and Distance-Based Chunk Alignment
by Sung-Gil Park, Yong-Geon Kim, Seuk-Woo Ryu, Byeong Gil Yoo, Sungeun Chung, Jeong-Seop Park, Woo-Jin Ahn and Myo-Taeg Lim
Appl. Sci. 2026, 16(7), 3376; https://doi.org/10.3390/app16073376 - 31 Mar 2026
Viewed by 1543
Abstract
Vision–language–action (VLA) models have shown strong potential for enabling robots to interpret goals and perform complex manipulation tasks by integrating perception, language, and control. However, existing VLAs rely heavily on large-scale, diverse demonstration datasets, which are difficult and expensive to collect. When trained [...] Read more.
Vision–language–action (VLA) models have shown strong potential for enabling robots to interpret goals and perform complex manipulation tasks by integrating perception, language, and control. However, existing VLAs rely heavily on large-scale, diverse demonstration datasets, which are difficult and expensive to collect. When trained with limited data, they often overfit to irrelevant visual cues such as background, lighting, or viewpoint, resulting in weak generalization. To overcome this limitation, we propose a simple yet effective object-centric learning framework for VLA. For each sub-task, the framework leverages an instance segmentation foundation model to identify and track task-relevant objects, and trains the policy on both the original RGB scene and two object-focused representations: (i) a masked image emphasizing the target object and (ii) an object-only crop. These multiple visual inputs share the same action supervision, encouraging the policy to attend to the manipulated object rather than the surrounding context. Furthermore, a distance-based chunk alignment mechanism ensures smooth control transitions between consecutive predicted action segments. Experiments conducted in both simulation and real hardware demonstrate that the proposed method achieves robust performance and stable trajectories across various manipulation tasks, validating its practicality and efficiency in training object-aware robotic behaviors. Full article
(This article belongs to the Special Issue Deep Reinforcement Learning for Multiagent Systems)
Show Figures

Figure 1

46 pages, 33541 KB  
Article
AIFloodSense: A Global Aerial Imagery Dataset for Semantic Segmentation and Understanding of Flooded Environments
by Georgios Simantiris, Konstantinos Bacharidis, Apostolos Papanikolaou, Petros Giannakakis and Costas Panagiotakis
Remote Sens. 2026, 18(6), 938; https://doi.org/10.3390/rs18060938 - 19 Mar 2026
Cited by 3 | Viewed by 1331
Abstract
Accurate flood detection is critical for disaster response, yet the scarcity of diverse annotated datasets hinders robust model development. Existing resources typically suffer from limited geographic scope and insufficient annotation granularity, restricting the generalization capabilities of computer vision methods. To bridge this gap, [...] Read more.
Accurate flood detection is critical for disaster response, yet the scarcity of diverse annotated datasets hinders robust model development. Existing resources typically suffer from limited geographic scope and insufficient annotation granularity, restricting the generalization capabilities of computer vision methods. To bridge this gap, we introduce AIFloodSense, a comprehensive evaluation benchmark designed to advance domain-generalized Artificial Intelligence for climate resilience. The dataset comprises 470 high-resolution aerial images capturing 230 distinct flood events across 64 countries and six continents. Unlike prior benchmarks, AIFloodSense ensures exceptional global diversity and temporal relevance (2022–2024), supporting three complementary tasks: (i) Image Classification, featuring novel sub-tasks for environment type, camera angle, and continent recognition; (ii) Semantic Segmentation, providing precise pixel-level masks for flood, sky, buildings, and background; and (iii) Visual Question Answering (VQA), enabling natural language reasoning for disaster assessment. We provide baseline benchmarks for all tasks using state-of-the-art architectures, demonstrating the dataset’s complexity and its utility in fostering robust AI tools for environmental monitoring. Crucially, we show that despite its compact size, AIFloodSense enables better generalization on external test sets than much larger alternatives, validating the premise that rigorous diversity is more effective than scale for training robust flood detection models, and is made publicly available to accelerate further research in the field. Full article
Show Figures

Figure 1

25 pages, 19621 KB  
Article
Scrap-SAM-CLIP: Assembling Foundation Models for Typical Shape Recognition in Scrap Classification and Rating
by Guangda Bao, Wenzhi Xia, Haichuan Wang, Zhiyou Liao, Ting Wu and Yun Zhou
Sensors 2026, 26(2), 656; https://doi.org/10.3390/s26020656 - 18 Jan 2026
Viewed by 1555
Abstract
To address the limitation of 2D methods in inferring absolute scrap dimensions from images, we propose Scrap-SAM-CLIP (SSC), a vision-language model integrating the segment anything model (SAM) and contrastive language-image pre-training in Chinese (CN-CLIP). The model enables identification of canonical scrap shapes, establishing [...] Read more.
To address the limitation of 2D methods in inferring absolute scrap dimensions from images, we propose Scrap-SAM-CLIP (SSC), a vision-language model integrating the segment anything model (SAM) and contrastive language-image pre-training in Chinese (CN-CLIP). The model enables identification of canonical scrap shapes, establishing a foundational framework for subsequent 3D reconstruction and dimensional extraction within the 3D recognition pipeline. Individual modules of SSC are fine-tuned on the self-constructed scrap dataset. For segmentation, the combined box-and-point prompt yields optimal performance among various prompting strategies. MobileSAM and SAM-HQ-Tiny serve as effective lightweight alternatives for edge deployment. Fine-tuning the SAM decoder significantly enhances robustness under noisy prompts, improving accuracy by at least 5.55% with a five-positive-points prompt and up to 15.00% with a five-positive-points-and-five-negative-points prompt. In classification, SSC achieves 95.3% accuracy, outperforming Swin Transformer V2_base by 2.9%, with t-SNE visualizations confirming superior feature learning capability. The performance advantages of SSC stem from its modular assembly strategy, enabling component-specific optimization through subtask decoupling and enhancing system interpretability. This work refines the scrap 3D identification pipeline and demonstrates the efficacy of adapted foundation models in industrial vision systems. Full article
(This article belongs to the Section Intelligent Sensors)
Show Figures

Figure 1

27 pages, 2129 KB  
Article
Dynamic Task Planning for Heterogeneous Platforms via Spatio-Temporal and Capability Dual-Driven Framework
by Guangxi Zhu, Gang Wang, Wei Fu and Changxing Han
Electronics 2026, 15(1), 202; https://doi.org/10.3390/electronics15010202 - 1 Jan 2026
Cited by 1 | Viewed by 714
Abstract
Dynamic task planning for heterogeneous platforms across land, sea, air, and space is essential for achieving integrated situational awareness, yet current systems suffer from limited spatiotemporal coverage and inefficient resource scheduling. To address these challenges, we propose a novel mission planning method that [...] Read more.
Dynamic task planning for heterogeneous platforms across land, sea, air, and space is essential for achieving integrated situational awareness, yet current systems suffer from limited spatiotemporal coverage and inefficient resource scheduling. To address these challenges, we propose a novel mission planning method that integrates spatiotemporal segmentation with Deep Reinforcement Learning (DRL). The approach establishes a multidimensional spatiotemporal decomposition model to break down complex observation scenarios into manageable subtasks, while incorporating a unified accessibility–visibility computation framework that accounts for Earth curvature, platform dynamics, and sensor constraints. Using a Spatio-Temporal Adaptive Scheduling Network (STAS-Net) algorithm optimized with a multi-objective reward function covering mission completion rate, temporal coordination, and residual detection capacity, the method enables intelligent coordination of heterogeneous platforms. Experimental results across small-, medium-, and large-scale scenarios demonstrate that the proposed framework consistently achieves high target coverage (up to 98.4% in small-scale and 89.7% in large-scale tasks), with a reduction in coverage loss that is only about half of that exhibited by greedy and genetic algorithms as task scale expands. Moreover, STAS-Net maintains low planning time (as low as 9.5 s in small-scale and only 18.3 s in large-scale scenarios) and high resource utilization (reaching 86.8% under large-scale settings), substantially outperforming both baseline methods in scalability and scheduling efficiency. The framework not only establishes a solid theoretical foundation but also provides a practical and feasible solution for enhancing the overall performance of multi-platform cooperative observation systems. Full article
(This article belongs to the Section Artificial Intelligence)
Show Figures

Figure 1

19 pages, 5721 KB  
Article
Efficient Weed Detection in Cabbage Fields Using a Dual-Model Strategy
by Mian Li, Wenpeng Zhu, Xiaoyue Zhang, Ying Jiang, Jialin Yu, Aimin Li and Xiaojun Jin
Agronomy 2026, 16(1), 93; https://doi.org/10.3390/agronomy16010093 - 29 Dec 2025
Cited by 2 | Viewed by 1132
Abstract
Accurate weed detection in crop fields remains a challenging task due to the diversity of weed species and their visual similarity to crops, especially under natural field conditions where lighting and occlusion vary. Traditional methods typically attempt to directly identify various weed species, [...] Read more.
Accurate weed detection in crop fields remains a challenging task due to the diversity of weed species and their visual similarity to crops, especially under natural field conditions where lighting and occlusion vary. Traditional methods typically attempt to directly identify various weed species, which demand large-scale, finely annotated datasets and often suffer from low generalization. To address these challenges, this study proposes a novel dual-model framework that simplifies the task by dividing it into two tractable stages. First, a crop segmentation network is used to identify and remove cabbage (Brassica oleracea L. ssp. pekinensis) regions from field images. Since crop categories are visually consistent and singular, this stage achieves high precision with relatively low complexity. The remaining non-crop areas, which contain only weeds and background, are then subdivided into grid cells. Each cell is classified by a second lightweight classification network as either background, broadleaf weeds, or grass weeds. The classification model achieved F1 scores of 95.1%, 91.1%, and 92.2% for background, broadleaf weeds, and grass weeds, respectively. This two-stage approach transforms a complex multi-class detection task into simpler, more manageable subtasks, improving detection accuracy while reducing annotation burden and enhancing robustness under the tested field conditions. Full article
Show Figures

Figure 1

28 pages, 5373 KB  
Article
Transfer Learning Based on Multi-Branch Architecture Feature Extractor for Airborne LiDAR Point Cloud Semantic Segmentation with Few Samples
by Jialin Yuan, Hongchao Ma, Liang Zhang, Jiwei Deng, Wenjun Luo, Ke Liu and Zhan Cai
Remote Sens. 2025, 17(15), 2618; https://doi.org/10.3390/rs17152618 - 28 Jul 2025
Cited by 3 | Viewed by 1524
Abstract
The existing deep learning-based Airborne Laser Scanning (ALS) point cloud semantic segmentation methods require a large amount of labeled data for training, which is not always feasible in practice. Insufficient training data may lead to over-fitting. To address this issue, we propose a [...] Read more.
The existing deep learning-based Airborne Laser Scanning (ALS) point cloud semantic segmentation methods require a large amount of labeled data for training, which is not always feasible in practice. Insufficient training data may lead to over-fitting. To address this issue, we propose a novel Multi-branch Feature Extractor (MFE) and a three-stage transfer learning strategy that conducts pre-training on multi-source ALS data and transfers the model to another dataset with few samples, thereby improving the model’s generalization ability and reducing the need for manual annotation. The proposed MFE is based on a novel multi-branch architecture integrating Neighborhood Embedding Block (NEB) and Point Transformer Block (PTB); it aims to extract heterogeneous features (e.g., geometric features, reflectance features, and internal structural features) by leveraging the parameters contained in ALS point clouds. To address model transfer, a three-stage strategy was developed: (1) A pre-training subtask was employed to pre-train the proposed MFE if the source domain consisted of multi-source ALS data, overcoming parameter differences. (2) A domain adaptation subtask was employed to align cross-domain feature distributions between source and target domains. (3) An incremental learning subtask was proposed for continuous learning of novel categories in the target domain, avoiding catastrophic forgetting. Experiments conducted on the source domain consisted of DALES and Dublin datasets and the target domain consists of ISPRS benchmark dataset. The experimental results show that the proposed method achieved the highest OA of 85.5% and an average F1 score of 74.0% using only 10% training samples, which means the proposed framework can reduce manual annotation by 90% while keeping competitive classification accuracy. Full article
Show Figures

Figure 1

18 pages, 3167 KB  
Article
Similarity Analysis of Upper Extremity’s Trajectories in Activities of Daily Living for Use in an Intelligent Control System of a Rehabilitation Exoskeleton
by Piotr Falkowski, Maciej Pikuliński, Tomasz Osiak, Kajetan Jeznach, Krzysztof Zawalski, Piotr Kołodziejski, Andrzej Zakręcki, Jan Oleksiuk, Daniel Śliż and Natalia Osiak
Actuators 2025, 14(7), 324; https://doi.org/10.3390/act14070324 - 30 Jun 2025
Cited by 5 | Viewed by 1813
Abstract
Rehabilitation robotic systems have been developed to perform therapy with minimal supervision from a specialist. Hence, they require algorithms to assess and support patients’ motions. Artificial intelligence brings an opportunity to implement new exercises based on previously modelled ones. This study focuses on [...] Read more.
Rehabilitation robotic systems have been developed to perform therapy with minimal supervision from a specialist. Hence, they require algorithms to assess and support patients’ motions. Artificial intelligence brings an opportunity to implement new exercises based on previously modelled ones. This study focuses on analysing the similarities in upper extremity movements during activities of daily living (ADLs). This research aimed to model ADLs by registering and segmenting real-life movements and dividing them into sub-tasks based on joint motions. The investigation used IMU sensors placed on the body to capture upper extremity motion. Angular measurements were converted into joint variables using Matlab computations. Then, these were divided into segments assigned to the sub-functionalities of the tasks. Further analysis involved calculating mathematical measures to evaluate the similarity between the different movements. This approach allows the system to distinguish between similar motions, which is critical for assessing rehabilitation scenarios and anatomical correctness. Twenty-two ADLs were recorded, and their segments were analysed to build a database of typical motion patterns. The results include a discussion on the ranges of motion for different ADLs and gender-related differences. Moreover, the similarities and general trends for different motions are presented. The system’s control algorithm will use these results to improve the effectiveness of robotic-assisted physiotherapy. Full article
Show Figures

Figure 1

16 pages, 7057 KB  
Article
VRBiom: A New Periocular Dataset for Biometric Applications of Head-Mounted Display
by Ketan Kotwal, Ibrahim Ulucan, Gökhan Özbulak, Janani Selliah and Sébastien Marcel
Electronics 2025, 14(9), 1835; https://doi.org/10.3390/electronics14091835 - 30 Apr 2025
Cited by 4 | Viewed by 2448
Abstract
With advancements in hardware, high-quality head-mounted display (HMD) devices are being developed by numerous companies, driving increased consumer interest in AR, VR, and MR applications. This proliferation of HMD devices opens up possibilities for a wide range of applications beyond entertainment. Most commercially [...] Read more.
With advancements in hardware, high-quality head-mounted display (HMD) devices are being developed by numerous companies, driving increased consumer interest in AR, VR, and MR applications. This proliferation of HMD devices opens up possibilities for a wide range of applications beyond entertainment. Most commercially available HMD devices are equipped with internal inward-facing cameras to record the periocular areas. Given the nature of these devices and captured data, many applications such as biometric authentication and gaze analysis become feasible. To effectively explore the potential of HMDs for these diverse use-cases and to enhance the corresponding techniques, it is essential to have an HMD dataset that captures realistic scenarios. In this work, we present a new dataset of periocular videos acquired using a virtual reality headset called VRBiom. The VRBiom, targeted at biometric applications, consists of 900 short videos acquired from 25 individuals recorded in the NIR spectrum. These 10 s long videos have been captured using the internal tracking cameras of Meta Quest Pro at 72 FPS. To encompass real-world variations, the dataset includes recordings under three gaze conditions: steady, moving, and partially closed eyes. We have also ensured an equal split of recordings without and with glasses to facilitate the analysis of eye-wear. These videos, characterized by non-frontal views of the eye and relatively low spatial resolutions (400×400), can be instrumental in advancing state-of-the-art research across various biometric applications. The VRBiom dataset can be utilized to evaluate, train, or adapt models for biometric use-cases such as iris and/or periocular recognition and associated sub-tasks such as detection and semantic segmentation. In addition to data from real individuals, we have included around 1100 presentation attacks constructed from 92 PA instruments. These PAIs fall into six categories constructed through combinations of print attacks (real and synthetic identities), fake 3D eyeballs, plastic eyes, and various types of masks and mannequins. These PA videos, combined with genuine (bona fide) data, can be utilized to address concerns related to spoofing, which is a significant threat if these devices are to be used for authentication. The VRBiom dataset is publicly available for research purposes related to biometric applications only. Full article
Show Figures

Figure 1

28 pages, 3978 KB  
Article
Geographic Named Entity Matching and Evaluation Recommendation Using Multi-Objective Tasks: A Study Integrating a Large Language Model (LLM) and Retrieval-Augmented Generation (RAG)
by Jiajun Zhang, Junjie Fang, Chengkun Zhang, Wei Zhang, Huanbing Ren and Liuchang Xu
ISPRS Int. J. Geo-Inf. 2025, 14(3), 95; https://doi.org/10.3390/ijgi14030095 - 20 Feb 2025
Cited by 11 | Viewed by 5767
Abstract
Geographical named entity matching, a crucial step in address encoding, aims to enhance address resolution accuracy through the precise identification and linkage of geographical named entity data. However, existing approaches tend to ignore the spatial information of entities, leading to misclassification. Drawing on [...] Read more.
Geographical named entity matching, a crucial step in address encoding, aims to enhance address resolution accuracy through the precise identification and linkage of geographical named entity data. However, existing approaches tend to ignore the spatial information of entities, leading to misclassification. Drawing on the human process of searching for addresses, this study proposes a multi-objective learning model named GNEMM that integrates the semantic and spatial information of geographical named entities. To further mimic the human cognitive process during address search, it incorporates the Retrieval-Augmented Generation (RAG) technique. By integrating newly added external address data with an advanced large language model (LLM) like GPT-4, it achieves precise address evaluation and recommendation. The model was tested using a standard geographical named entity dataset from Shandong Province, focusing on three sub-tasks: element segmentation, matching, and spatial similarity score prediction. The experimental results indicate that the method achieves a geographical named entity matching accuracy of up to 99%, with improvements of 10% and 5% in the segmentation and prediction sub-tasks. GNEMM performs best in address-matching tasks of various scales, and the vectors extracted by GNEMM perform best in the downstream retrieval and matching of various address types, which verifies its applicability in geographical named entity recommendation applications. Full article
Show Figures

Figure 1

38 pages, 14791 KB  
Article
Online High-Definition Map Construction for Autonomous Vehicles: A Comprehensive Survey
by Hongyu Lyu, Julie Stephany Berrio Perez, Yaoqi Huang, Kunming Li, Mao Shan and Stewart Worrall
J. Sens. Actuator Netw. 2025, 14(1), 15; https://doi.org/10.3390/jsan14010015 - 2 Feb 2025
Cited by 4 | Viewed by 12041
Abstract
High-definition (HD) maps aim to provide detailed road information with centimeter-level accuracy, essential for enabling precise navigation and safe operation of autonomous vehicles (AVs). Traditional offline construction methods involve several complex steps, such as data collection, point cloud generation, and feature extraction, but [...] Read more.
High-definition (HD) maps aim to provide detailed road information with centimeter-level accuracy, essential for enabling precise navigation and safe operation of autonomous vehicles (AVs). Traditional offline construction methods involve several complex steps, such as data collection, point cloud generation, and feature extraction, but these methods are resource-intensive and struggle to keep pace with the rapidly changing road environments. In contrast, online HD map construction leverages onboard sensor data to dynamically generate local HD maps, offering a bird’s-eye view (BEV) representation of the surrounding road environment. This approach has the potential to improve adaptability to spatial and temporal changes in road conditions while enhancing cost-efficiency by reducing the dependency on frequent map updates and expensive survey fleets. This survey provides a comprehensive analysis of online HD map construction, including the task background, high-level motivations, research methodology, key advancements, existing challenges, and future trends. We systematically review the latest advancements in three key sub-tasks: map segmentation, map element detection, and lane graph construction, aiming to bridge gaps in the current literature. We also discuss existing challenges and future trends, covering standardized map representation design, multitask learning, and multi-modality fusion, while offering suggestions for potential improvements. Full article
(This article belongs to the Special Issue Advances in Intelligent Transportation Systems (ITS))
Show Figures

Figure 1

29 pages, 9718 KB  
Article
Segment, Compare, and Learn: Creating Movement Libraries of Complex Task for Learning from Demonstration
by Adrian Prados, Gonzalo Espinoza, Luis Moreno and Ramon Barber
Biomimetics 2025, 10(1), 64; https://doi.org/10.3390/biomimetics10010064 - 17 Jan 2025
Cited by 8 | Viewed by 3323
Abstract
Motion primitives are a highly useful and widely employed tool in the field of Learning from Demonstration (LfD). However, obtaining a large number of motion primitives can be a tedious process, as they typically need to be generated individually for each task to [...] Read more.
Motion primitives are a highly useful and widely employed tool in the field of Learning from Demonstration (LfD). However, obtaining a large number of motion primitives can be a tedious process, as they typically need to be generated individually for each task to be learned. To address this challenge, this work presents an algorithm for acquiring robotic skills through automatic and unsupervised segmentation. The algorithm divides tasks into simpler subtasks and generates motion primitive libraries that group common subtasks for use in subsequent learning processes. Our algorithm is based on an initial segmentation step using a heuristic method, followed by probabilistic clustering with Gaussian Mixture Models. Once the segments are obtained, they are grouped using Gaussian Optimal Transport on the Gaussian Processes (GPs) of each segment group, comparing their similarities through the energy cost of transforming one GP into another. This process requires no prior knowledge, it is entirely autonomous, and supports multimodal information. The algorithm enables generating trajectories suitable for robotic tasks, establishing simple primitives that encapsulate the structure of the movements to be performed. Its effectiveness has been validated in manipulation tasks with a real robot, as well as through comparisons with state-of-the-art algorithms. Full article
(This article belongs to the Special Issue Bio-Inspired and Biomimetic Intelligence in Robotics: 2nd Edition)
Show Figures

Figure 1

25 pages, 7128 KB  
Article
Comparing Skill Transfer Between Full Demonstrations and Segmented Sub-Tasks for Neural Dynamic Motion Primitives
by Geoffrey Hanks, Gentiane Venture and Yue Hu
Machines 2024, 12(12), 872; https://doi.org/10.3390/machines12120872 - 1 Dec 2024
Cited by 1 | Viewed by 2253
Abstract
Programming by demonstration has shown potential in reducing the technical barriers to teaching complex skills to robots. Dynamic motion primitives (DMPs) are an efficient method of learning trajectories from individual demonstrations using second-order dynamic equations. They can be expanded using neural networks to [...] Read more.
Programming by demonstration has shown potential in reducing the technical barriers to teaching complex skills to robots. Dynamic motion primitives (DMPs) are an efficient method of learning trajectories from individual demonstrations using second-order dynamic equations. They can be expanded using neural networks to learn longer and more complex skills. However, the length and complexity of a skill may come with trade-offs in terms of accuracy, the time required by experts, and task flexibility. This paper compares neural DMPs that learn from a full demonstration to those that learn from simpler sub-tasks for a pouring scenario in a framework that requires few demonstrations. While both methods were successful in completing the task, we find that the models trained using sub-tasks are more accurate and have more task flexibility but can require a larger investment from the human expert. Full article
(This article belongs to the Special Issue Robot Intelligence in Grasping and Manipulation)
Show Figures

Figure 1

11 pages, 483 KB  
Communication
Optimizing the Agricultural Internet of Things (IoT) with Edge Computing and Low-Altitude Platform Stations
by Deshan Yang, Jingwen Wu and Yixin He
Sensors 2024, 24(21), 7094; https://doi.org/10.3390/s24217094 - 4 Nov 2024
Cited by 14 | Viewed by 3604
Abstract
Using low-altitude platform stations (LAPSs) in the agricultural Internet of Things (IoT) enables the efficient and precise monitoring of vast and hard-to-reach areas, thereby enhancing crop management. By integrating edge computing servers into LAPSs, data can be processed directly at the edge in [...] Read more.
Using low-altitude platform stations (LAPSs) in the agricultural Internet of Things (IoT) enables the efficient and precise monitoring of vast and hard-to-reach areas, thereby enhancing crop management. By integrating edge computing servers into LAPSs, data can be processed directly at the edge in real time, significantly reducing latency and dependency on remote cloud servers. Motivated by these advancements, this paper explores the application of LAPSs and edge computing in the agricultural IoT. First, we introduce an LAPS-aided edge computing architecture for the agricultural IoT, in which each task is segmented into several interdependent subtasks for processing. Next, we formulate a total task processing delay minimization problem, taking into account constraints related to task dependency and priority, as well as equipment energy consumption. Then, by treating the task dependencies as directed acyclic graphs, a heuristic task processing algorithm with priority selection is developed to solve the formulated problem. Finally, the numerical results show that the proposed edge computing scheme outperforms state-of-the-art works and the local computing scheme in terms of the total task processing delay. Full article
(This article belongs to the Special Issue Wireless Sensor Networks in Industrial/Agricultural Environments)
Show Figures

Figure 1

Back to TopTop