Learning from Demonstration for Robotic Deburring and Polishing: A Systematic Mapping Study
Abstract
1. Introduction
- Structured Taxonomy: A structured taxonomy of LfD architectures applied in robotic deburring and polishing, categorizing 24 primary studies into five methodological clusters: DMP-based methods, probabilistic and statistical models, deep learning and generative AI architectures, autonomous dynamical systems, and direct adaptive control approaches.
- Mechatronic Analysis: A systematic analysis of sensory and control configurations, documenting how robotic platforms, force/torque sensors, vision systems, and haptic interfaces are integrated to support LfD-based surface finishing.
- Evaluation Metrics Synthesis: A comprehensive characterization of evaluation metrics employed in the field, spanning kinematic trajectory accuracy, dynamic force-tracking performance, and physical surface quality measures.
- Industrial Challenges Roadmap: An evidence-based synthesis of industrial deployment challenges, identifying the key technical barriers that currently impede the adoption of LfD-based surface finishing in production environments.
- RQ1—Methodological Landscape: Which Learning-from-Demonstration (LfD) architectures, trajectory learning algorithms, and mathematical representations are most commonly used in robotic deburring and polishing tasks?
- RQ2—System and Sensory Configuration: How are robotic systems, control architectures, and sensory modalities (e.g., force sensing, vision, multimodal perception) designed and integrated to support LfD-based surface-finishing operations?
- RQ3—Evaluation and Performance Metrics: What performance evaluation criteria and validation metrics are used to assess the effectiveness of LfD-based robotic deburring and polishing methods?
- RQ4—Industrial Deployment Challenges: What are the key technical challenges, limitations, and research gaps that hinder the transition of LfD-based robotic surface-finishing methods from laboratory settings to industrial applications?
2. Methodology
2.1. Protocol and Registration
2.2. Eligibility Criteria
- Thematic Relevance: The study must explicitly investigate robotic surface-finishing operations, specifically deburring or polishing.
- Methodological Approach: The proposed robotic control or programming framework must incorporate an LfD-based method (e.g., kinesthetic teaching, teleoperated demonstration, or haptic imitation learning) where motion or force policies are extracted from human demonstrations.
- Technical Depth: The study must provide experimental validation or detailed theoretical formulations regarding trajectory generation, force regulation, control architectures, or sensor configurations.
- Publication Type and Language: Only peer-reviewed academic journal articles or conference proceedings published in the English language were included.
- Out-of-Scope Operations: Studies addressing generic robotic manipulation tasks (e.g., pick-and-place, assembly, peg-in-hole insertion, or trajectory tracking in free space) without a specific surface-finishing context.
- Non-Demonstration Control: Research utilizing traditional robotic programming (e.g., offline CAD/CAM programming, manual teach-pendant jogging) or model-based adaptive controllers that do not learn or adapt based on human demonstration data.
- Insufficient Quality or Length: Extended Abstracts, editorial prefaces, technical reports, white papers, book reviews, or unpublished Master’s/Doctoral theses.
- Accessibility Barriers: Articles for which the full-text version was unavailable or lacked sufficient data to extract parameters necessary for answering the Research Questions.
2.3. Search Strategy and Information Sources
- IEEE Xplore: Targeted for its high concentration of robotics, automation, and control engineering publications.
- Scopus: Utilized for broad, multidisciplinary coverage of the engineering and physical sciences literature.
- Web of Science (Core Collection): Employed to capture high-impact journals and citation indexes.
- Google Scholar: Used as a supplementary resource to capture early-access articles, conference papers, and preprints, thereby minimizing potential publication bias.
- Block A (Learning Paradigm): “Learning from Demonstration”, “LfD”, “Imitation Learning”, “Programming by Demonstration”, “DMP”, “Dynamic Movement Primitives”, “ProMP”, “Probabilistic Movement Primitives”, “Skill learning”, “Task learning”.
- Block B (Application Context): “Deburring”, “Polishing”, “Surface finishing”, “Surface-finishing”, “Finishing”, “Robotic finish*”.
2.4. Screening and Selection Process
- Identification: The initial database searches yielded a total of 288 records (Scopus: 57; Web of Science: 43; IEEE Xplore: 29; Google Scholar: 159). Following the elimination of 65 duplicate entries, 223 unique records were compiled for screening.
- Screening (Title and Abstract Review): A Title-and-Abstract review was conducted on the 223 unique records against the predefined inclusion and exclusion criteria. This phase resulted in the exclusion of 197 records that lacked relevance to LfD-based robotic deburring or polishing, leaving 26 records for full-text eligibility assessment.
- Eligibility (Full-Text Review): The full texts of the 26 remaining studies were retrieved and analyzed. During this phase, 2 studies were excluded due to out-of-scope applications: one study focused exclusively on grinding-based material removal [2], and another study addressed sanding without a polishing or deburring component [3].
2.5. Methodological Quality Assessment
- QA1—Experimental Validation: Does the study validate the proposed LfD framework on a physical robotic manipulator performing a surface-finishing task?
- QA2—Reproducibility and Statistical Reporting: Are the experiments repeated over multiple trials? Are quantitative statistical measures (e.g., mean and standard deviation) reported?
- QA3—Baseline Comparison: Is the proposed method compared against a baseline (e.g., standard model-based control, human expert performance, or alternative LfD algorithms)?
- QA4—Generalizability: Is the learned skill tested on varying workpiece geometries, orientations, or materials?
- QA5—Quantitative Performance Metrics: Does the study report quantitative evaluation metrics for both kinematic trajectory accuracy and dynamic force tracking?
- QA6—Industrial Relevance: Is the task validated using an industrial workpiece? Are real-world industrial deployment constraints (e.g., tool wear, cycle time, worker safety) discussed?
- QA1—Experimental Validation: Does the study validate the proposed LfD framework on a physical robotic manipulator performing a surface-finishing task?
- Score 2 (Fully Met): The proposed LfD algorithm is validated on physical robotic hardware (i.e., a physical manipulator) performing a real finishing operation (e.g., polishing, deburring) to account for real-world high-frequency contact forces. The majority of the included studies (22 out of 24, 92%; e.g., [4,5]) fully met this criterion.
- Score 1 (Partially Met): The framework is evaluated only in high-fidelity robotic simulators (e.g., CoppeliaSim, Gazebo, MuJoCo, PyBullet) with representative physical models, but it lacks hardware execution.
- Score 0 (Not Met): The framework is evaluated only in basic mathematical simulation (e.g., MATLAB plotting paths without robot dynamics/collision models) or not validated at all.
- QA2—Reproducibility and Statistical Reporting: Are the experiments repeated over multiple trials? Are quantitative statistical measures (e.g., mean and standard deviation) reported?
- Score 0 (Not Met): The study reports only a single test trial or provides no information regarding repetition and statistical metrics.
- QA3—Baseline Comparison: Is the proposed method compared against a baseline (e.g., standard model-based control, human expert performance, or alternative LfD algorithms)?
- Score 0 (Not Met): The proposed method’s performance is reported in isolation with no baseline comparison or ablation study.
- QA4—Generalizability: Is the learned skill tested on varying workpiece geometries, orientations, or materials?
- Score 0 (Not Met): The learned skill is evaluated only on the exact same workpiece geometry and trajectory configuration used in the demonstration phase, without any perturbation.
- QA5—Quantitative Performance Metrics: Does the study report quantitative evaluation metrics for both kinematic trajectory accuracy and dynamic force tracking?
- Score 0 (Not Met): The evaluation is purely qualitative (e.g., visual surface finish inspection) with no quantitative metrics reported for either force or trajectory.
- QA6—Industrial Relevance: Is the task validated using an industrial workpiece? Are real-world industrial deployment constraints (e.g., tool wear, cycle time, worker safety) discussed?
- Score 0 (Not Met): Validation is restricted to simple mock-up workpieces with no discussion of industrial deployment, cycle times, tool wear, or real-world deployment challenges.
- High Quality (H): Cumulative score of 9 to 12.
- Moderate Quality (M): Cumulative score of 5 to 8.
- Low Quality (L): Cumulative score of 0 to 4.
2.6. Data Extraction and Synthesis
- Bibliographic Metadata: Authors, publication year, publication venue, and the geographic location of the corresponding author.
- LfD Algorithmic Architecture: Trajectory representation model, learning algorithms, and mathematical formulations (e.g., DMPs, GMM, GMR, Diffusion Policy, Neural ODEs).
- Sensory Modalities: Sensors integrated during the human demonstration phase and the robot execution phase (e.g., 6-DOF force/torque sensors, joint torque sensors, RGB/RGB-D cameras, haptic interfaces).
- Hardware and Control Platform: The specific robotic manipulator model, degrees of freedom (DOF), and low-level force/position control strategies (e.g., impedance control, admittance control, variable stiffness control).
- Application Environment: The specific surface-finishing task (deburring, polishing, cleaning, rust removal) and the shape/material of the test workpiece.
2.7. Effect Measures and Synthesis Rationale
3. Results
3.1. Study Selection and PRISMA Flow Diagram
- Search Query and Filters: The specific search keywords used are presented in Table 2. The publication date filter was manually set to the 2016–2026 range, and the document type filter was manually restricted to “Articles only”.
- Retrieval Limits and Pages: This search yielded a total of 159 results, which were distributed across nine pages on Google Scholar. Since the total number of results was well below Google Scholar’s standard pagination limit (approx. 1000 results), all nine pages of results were fully retrieved.
- Export Procedure: Due to Google Scholar’s lack of native batch-export features for large datasets, the search results from these nine pages were manually exported into Excel.
- Duplicate Detection and Reference Management: Records retrieved from Scopus, Web of Science, and IEEE Xplore were first automatically imported into Mendeley. Duplicate detection and removal were performed automatically within Mendeley, and the cleaned list was subsequently exported to Excel. Finally, manual cross-checking and duplicate elimination between the database records and the Google Scholar results were performed within Excel.
3.2. Quantitative Study Characteristics
3.3. Bibliometric Analysis
3.3.1. Temporal Distribution
3.3.2. Geographical Distribution
3.3.3. Quantitative Methodological and Sensory Distributions
3.4. Methodological Quality Assessment Results
3.5. RQ1: Methodological Landscape
3.5.1. Dynamic Movement Primitives (DMPs) and Variants
- Forced-Controlled Dynamic Coupling DMPs (FDC-DMP): Introduced by Shen et al. [22], this framework adds a coupling term driven by interaction forces directly into the DMP acceleration equation, enabling the robot to dynamically modify its path (e.g., avoiding obstacles during bus body polishing) without altering the global target.
- B-Spline DMPs (BDMPs): Wang et al. [13] proposed replacing the standard Gaussian basis functions in DMPs with B-splines. This modification significantly improves trajectory modeling accuracy with fewer basis functions. The forcing term is parameterized as , where represents the B-spline basis functions. When optimized using Policy Improvement with Path Integrals (PI2) Reinforcement Learning, BDMPs demonstrate high generalization capability for polishing trajectories and force profiles under unseen workpiece positions.
- Riemannian DMPs: Liao et al. [24] extended DMPs to Riemannian manifolds (e.g., Cartesian space and 2D sphere manifolds) to simultaneously model human motion, 3D endpoint stiffness, and contact forces from a one-shot demonstration, solved via Quadratic Programming (QP). Trajectories on the manifold are generated by mapping the states to the tangent space, , maintaining geometrical properties of robot orientations.
- Neural Network-Augmented DMPs: Wang et al. [9] integrated a Phase-Modulated Diagonal Recurrent Neural Network (PMDRNN) with DMPs to adaptively predict trajectory offsets based on real-time force-tracking deviations, mitigating environmental uncertainties.
3.5.2. Probabilistic and Statistical Models
- Gaussian Mixture Models and Regression (GMM-GMR): Used to model the joint distribution of time, space, and force parameters. Wu et al. [11] utilized GMMs to encode human polishing dynamics, combining GMR with a variable-impedance controller to regulate contact compliance. Zhai et al. [19] integrated GMM-GMR with a vector-valued Gaussian Process (GP) to online-modulate robotic trajectories when subjected to human external forces.
- Probabilistic Movement Primitives (ProMPs): Unlike DMPs, ProMPs capture the statistical variance of demonstrations. ProMPs represent a trajectory as a linear combination of basis functions: , where captures the statistical variance across multiple demonstrations. Wang et al. [14] developed Arc Length ProMPs (AL-ProMP), mapping the probability distribution of contact forces to spatial coordinates (arc-length ) rather than time , formulating the trajectory as . This formulation prevents trajectory distortions during non-linear speed scaling.
- Fourier Movement Primitives (FMPs): Grounded in signal processing, Kulak et al. [15] proposed FMPs using Fourier series as basis functions: FMPs approximate periodic, multi-frequency signals (e.g., circular polishing patterns) from unaligned demonstrations without requiring phase alignment or frequency extraction.
3.5.3. Deep Learning and Generative AI Architectures
- Diffusion Policies: Ke et al. [10] and Li et al. [5] utilized diffusion models to generate continuous, expert-like motion–force trajectories. In the DP-RRL framework [5], the Diffusion Policy generates a trajectory distribution by iteratively denoising a random sequence using a noise predictor conditioned on observation . The residual RL agent then predicts a displacement, , to correct the reference force based on contact dynamics.
- Neural Ordinary Differential Equations (Hyper-NODEs): Xu et al. [7] developed a Hyper-NODE architecture to generate continuous position and orientation (quaternion) trajectories. The system dynamics are modeled as follows:where represents the continuous hidden state, and are the weights generated by the hypernetwork. Paired with Control Barrier Functions (CBFs), the system guarantees obstacle avoidance in local non-polishing areas (LNP areas) while estimating admittance control parameters.
- Neural Network Morphing: Möhl et al. [12] designed a keypoint-driven non-linear morphing network to transfer demonstrated trajectories between 3D point clouds of geometrically similar objects without CAD models.
3.5.4. Autonomous Dynamical Systems (DSs)
- Stable Limit Cycles: Duarte et al. [27] represented periodic human polishing movements (e.g., circular or elliptical motions) using a time-invariant DS with a stable limit cycle attractor. The system is formulated as a second-order non-linear dynamical system, , where the trajectories are forced to converge asymptotically to a closed orbit . This formulation guarantees that the robot converges back to the demonstrated polishing pattern even after being physically displaced.
- DS-based Imitation Learning: Si et al. [8] proposed a dynamically stable energy-field framework to provide virtual haptic guidance during teleoperated human demonstrations, reducing operator physical workload.
3.5.5. Direct Adaptive Control and Parameter Estimation
3.6. RQ2: System and Sensory Configuration
3.6.1. Sensory Modalities and Multimodal Perception
- Force and Torque Sensing: As the primary driver of closed-loop execution, force feedback is integrated in 87.5% of the studies (either as the sole sensor or in multimodal setups). While end-effector 6-DOF F/T sensors are standard, Hamdan et al. [23] proposed a dual-force sensor configuration. One sensor measures the human operator’s guiding force (), while the second measures the tool–workpiece interaction force (). By calculating the environmental reaction force as follows, the system isolates the environmental dynamics (stiffness, friction) from human guidance inputs:
- Vision and Spatial Perception: To handle geometrically complex surfaces, depth sensors (e.g., overhead or wrist-mounted RGB-D cameras) are integrated. These sensors capture raw 3D point clouds, which are processed via PointNet++ or keypoint-based neural networks to reconstruct surface meshes [10] or guide trajectory morphing [12].
- Kinematic and Biometric Tracking: Teleoperation and kinesthetic demonstration interfaces utilize haptic devices (e.g., Geomagic Touch) or wearable inertial measurement units (IMUs). More advanced setups incorporate surface electromyography (sEMG) sensors on the human arm to capture synergistic muscle activity, translating muscle co-contraction directly into robot joint stiffness parameters.
- Instrumented Tools: To facilitate platform-independent demonstrations, Fischer et al. [6] designed custom instrumented tools and mechanical alignment plates to record high-quality contact forces and orientations directly on the workpiece.
3.6.2. Control Architectures
- Variable-Impedance and Admittance Control: These strategies model the robot–workpiece interface as a mass-spring-damper system. The low-level dynamic behavior is governed by the admittance control law:Here, and are the desired mass, damping, and stiffness matrices. is the reference trajectory, is the actual position, is the external interaction force, and is the target reference force. Variable-impedance controllers [11,16] modulate stiffness and damping online. For instance, stiffness is reduced when transitioning onto hard, brittle materials (e.g., iron) to prevent impact chatter and increased when transitioning onto soft materials (e.g., wood) to ensure uniform material removal.
- Iterative Learning Control (ILC): To compensate for repetitive tracking errors, ILC is integrated with impedance control. Zhang et al. [20] combined GMM trajectory models with Dynamic Time Warping ILC (DTW-ILC), iteratively updating the robot’s reference path based on stiffness estimation to achieve fast force convergence over multiple polishing passes.
3.6.3. Robotic Systems and Hardware Integration
3.7. RQ3: Evaluation and Performance Metrics
3.7.1. Kinematic and Trajectory Accuracy Metrics
- Dynamic Time Warping (DTW) Distance: Quantifies the spatiotemporal similarity between the demonstrated human trajectory and the robot’s executed path, especially when feed rates vary.
- Root Mean Square Error (RMSE): Calculates the spatial deviation (in millimeters) between the executed end-effector path and the demonstrated trajectory over samples:
- Pearson Correlation Coefficient (r): Measures the shape similarity of the trajectories:with values closer to 1.0 indicating high imitation fidelity.
- Relative Smoothness ( for position, for orientation): Evaluates the jerk of the generated trajectory to ensure smooth robotic motion.
3.7.2. Force Tracking and Dynamic Interaction Metrics
- Force Root Mean Square Error (Force RMSE): The primary metric to evaluate force-tracking performance, measuring the deviation between the executed contact force and the demonstrated reference force profile :
- Maximum Impact Force (): Evaluates system compliance and safety during initial tool contact or material transitions.
- Mean Force Deviation (): Quantifies the stability of the normal force during continuous polishing.
3.7.3. Surface Quality and Process-Specific Metrics
- Surface Roughness (Ra, Rq, Rz): Measured using contact profilometers or white-light interferometers. A reduction in average roughness () verifies successful surface smoothing.
- Material Removal Rate (MRR): Replicating the expert’s material removal strategy is evaluated based on Preston’s equation:where is the thickness of the removed material, is Preston’s coefficient (depending on tool and material properties), is the contact pressure (directly proportional to normal force ), and is the relative tool speed. Previous studies have estimated these parameters online to adapt forces dynamically on curved surfaces [4].
- Remaining Stain Ratio (RSR): In cleaning applications, image segmentation is used to calculate the percentage of stains remaining on the surface post-execution.
3.8. RQ4: Industrial Deployment Challenges
3.8.1. The Sim-to-Real Gap and Complex Contact Dynamics
3.8.2. Demonstration Quality and Hardware Constraints
- Kinematic Interference: The physical weight and joint limits of the robot arm restrict the operator’s natural movement, leading to distorted demonstrations.
- Sensor Noise: High-speed spindle rotation and pneumatic tool vibrations generate significant mechanical noise, degrading force/torque sensor readings during the teaching phase.
- Cognitive Overload: Controlling the robot’s spatial path, tool orientation, and contact force simultaneously in real time places a high cognitive demand on the human expert.
3.8.3. Generalization to Complex Geometries and LNP Areas
- Local Non-Polishing (LNP) Areas: Industrial components often feature functional geometry (e.g., threaded holes, slots, and ribs) that must remain untouched. Autonomously detecting and avoiding these LNP areas while maintaining a constant normal force on the surrounding freeform surface remains an open control problem.
- Geometric Generalization: Trajectory generalization models (e.g., DMPs) often distort orientations when scaling trajectories to highly curved, non-planar workpieces, risking collision or uneven polishing.
3.8.4. Multimodal Perception and Computational Complexity
3.8.5. Need for Robust Human-in-the-Loop (HITL) Systems
4. Discussion
4.1. Methodological and Academic Perspectives: Temporal Evolution (2016–2026)
4.2. Task-Specific Synthesis: Polishing vs. Deburring
- High-Bandwidth Micro–Macro Control: Ref. [18] implemented a dual-stage macro-micro control scheme where a slow hexapod robot (macro) tracks the tool path, while a high-bandwidth piezoelectric actuator (micro-manipulator with a 1-mm stroke operating at 1000 Hz in a haptic loop) executes micro-adjustments to maintain a constant normal force.
- Force-Coupled Specified DMPs (sDMPs): Ref. [17] developed an sDMP architecture that replaces the standard spatial DMP forcing function with a Gaussian function of the contact force. This force-coupling term enables the robot to adaptively alter its trajectory and feed rate during physical interaction.
- Meta-Heuristic Optimization of Compliance: To tune the coupling parameters of force-sensitive primitives, meta-heuristic algorithms like Particle Swarm Optimization (PSO) are used to minimize trajectory error under oscillatory contact forces [17]. The underrepresentation of deburring represents a critical research gap.
4.3. Economic and Operational Implications for SMEs
- Reduced Reliance on Programming Expertise: LfD has the potential to lower the reliance on highly specialized robotic programming engineers. By using intuitive kinesthetic teaching or teleoperation, existing shop-floor operators can demonstrate finishing strategies. However, specialized engineering oversight is still required for low-level controller tuning, safety configuration, and system integration.
- Setup and Changeover Time Reduction: LfD is projected to shorten setup and changeover times. While traditional CAD/CAM programming requires watertight 3D models and offline path planning, an LfD system can theoretically adapt to new geometries in laboratory trials within minutes through a single expert demonstration or neural morphing. Production-scale validation is necessary to confirm these time savings under factory conditions where part tolerances drift.
- Potential Hardware Cost Reductions: LfD may lower hardware barriers by utilizing consumer-grade sensors (e.g., depth cameras) and avoiding expensive force-accurate simulators. Nevertheless, the capital cost of collaborative robots and industrial force/torque sensors remains a significant financial barrier for many SMEs. By consolidating these hypothesized mechatronic, personnel, and temporal savings, LfD represents a promising and potentially viable pathway for SMEs to automate complex finishing operations under highly variable production requirements. However, empirical studies on cycle time, reliability, and economics are required to confirm these benefits.
4.4. Comparative Synthesis of the Included Studies
4.5. Demonstration Efficiency Across LfD Algorithms
- Dynamic Movement Primitives (DMPs) and Variants: DMPs are highly demonstration-efficient, typically generalizing from a single human demonstration (one-shot learning). By leveraging a second-order spring-damper system as a structural prior, DMPs guarantee global asymptotic convergence to the goal, allowing the learnable forcing term to be fitted using standard regression techniques (e.g., Locally Weighted Regression or Least Squares) on one trajectory. In surface finishing, DMPs can adapt a single learned polishing path to varying coordinate frames (e.g., [22]) or learn force–velocity profiles from a one-shot demonstration (e.g., [17,24]). This makes them highly suitable for small-batch industrial lines where changeover time must be kept under a few minutes.
- Probabilistic and Statistical Models: Unlike DMPs, probabilistic models like Gaussian Mixture Models/Regression (GMM-GMR) or Probabilistic Movement Primitives (ProMPs) require multiple demonstrations (typically 5 to 15) to capture the statistical covariance, temporal correlations, and task variability across different trials. Zhang et al. [20] evaluated path learning errors using 5, 10, and 15 demonstrations, showing that the path error converges to a baseline (under 1.0 mm) after 10 demonstrations, and that additional demonstrations capture human variability without significantly improving spatial precision. This suggests a practical limit of approximately 10 demonstrations for statistical learning in contact-rich tasks. ProMP models like AL-ProMP map these probabilities to spatial arc-lengths rather than time, preserving the statistical correlation of contact forces along the path from a handful of demonstrations. These models are highly feasible for industrial use, requiring less than 30 min of manual demonstration.
- Autonomous Dynamical Systems (DSs): DS-based models guarantee immediate, time-invariant reactivity to physical perturbations. However, learning a stable, high-dimensional vector field over the entire state space requires demonstrations initiated from multiple starting points to define the attraction basin. Refs. [8,11] utilized five physical demonstrations to train stable guidance fields and non-linear dynamics, respectively, while Duarte et al. [27] noted that multiple demonstrations are required to extrapolate periodic polishing limit cycles.
- Deep Learning and Generative AI Architectures: Deep neural networks (e.g., Diffusion Policies, Neural ODEs) have high capacity for visual–haptic integration but suffer from extreme sample inefficiency because they lack physical structural priors (e.g., second-order attractor dynamics). Ref. [5] collected 300 expert demonstrations on physical hardware to train a multimodal Diffusion Policy for basin cleaning, demonstrating the high data-collection overhead. Fischer et al. [6] explicitly addressed this bottleneck in industrial cleaning, noting that traditional imitation learning requires hundreds of demonstrations (e.g., 650 trials), and proposed a few-shot learning approach utilizing approximately 10 demonstrations by using instrumented tool-based segmentation. Thus, deep generative architectures remain restricted to high-volume manufacturing lines due to the days required for data collection and model training.
- Direct Adaptive Control and Parameter Estimation: Decoupling the task into a spatial path (motion skill) and a normal force profile (force skill) allows direct transfer from a single demonstration [4], as the active controller (e.g., adaptive admittance or computed-torque impedance control) compensates for surface irregularities online without requiring a statistical model. This achieves high demonstration efficiency, though it is limited to tasks that can be represented by 1D normal force profiles on simple geometries.
4.6. Identification of Research Gaps in the LfD Literature
- Tactile Complexity and Mechanical Hazards of Deburring:
- 2.
- Generalization to Freeform Geometries without CAD Templates:
- 3.
- The Sim-to-Real Gap in Physical Interaction Physics:
- 4.
- Lack of Intuitive Online Human-in-the-Loop (HITL) Correction:
4.7. How the Current Study Addresses These Research Gaps
- Addressing the Task Imbalance: By exposing the severe deficit in deburring research (only 12.5% of the corpus), this study provides a clear technical analysis of why deburring is exceptionally challenging (high-frequency transient contact, mechanical chattering) and catalogues the haptic and PSO-based control strategies used by the few successful deburring studies. This guides future researchers toward the control and haptic configurations necessary to tackle deburring.
- Providing a Mechatronic and Sensory Roadmap: This study maps the exact mechatronic configurations required to support LfD (such as the dual-force sensor configuration of [23] to isolate human forces from environmental reaction forces). By documenting the frequency of force-only (70.8%) vs. multimodal (16.7%) setups, it provides a design guideline for building hardware platforms capable of perceiving both geometric and tactile environments.
- Synthesizing Quantitative Performance Benchmarks: To assist researchers in bridging the sim-to-real gap, this work extracted and synthesized concrete quantitative baseline performance metrics reported from physical experiments (force-tracking RMSEs of 0.5–1.5 N, trajectory RMSEs under 2.0 mm, and finished surface roughness of 0.1–0.3 µm). These benchmarks provide a standard against which simulated policies can be verified and validated.
- Structuring Algorithmic Trade-offs: By categorizing the LfD algorithms into five methodological families (Table 6) and comparing their characteristics (Table 10), this study maps which algorithms are best suited for specific challenges. For instance, it highlights that while DMPs excel at spatial scaling, probabilistic models (like GMM-GMR) are better suited for online trajectory deformation during human intervention (HITL), and generative AI is best for visual point-cloud mapping, allowing practitioners to select the optimal control architecture.
5. Limitations of This Study
- Database Coverage and Search Strategy: The literature search was restricted to Scopus, Web of Science, IEEE Xplore, and Google Scholar. While these are the primary repositories for engineering and robotics research, some publications indexed in regional or specialized databases may have been omitted. Additionally, keyword-based search queries centered around terms like “Learning from Demonstration” and “DMP” might have missed relevant studies that utilize alternative terminology, such as “skill transfer” or “human–robot co-manipulation,” despite sharing the same underlying architecture.
- Language Bias: In accordance with the screening protocol, only articles published in the English language were included. Given that 62.5% of the included literature originates from China, and that countries like Japan and Germany possess strong academic and industrial backgrounds in robotic manufacturing, excluding non-English publications likely introduced a language bias. Highly innovative papers published in Chinese, Japanese, or German may have been overlooked.
- Exclusion of the Grey Literature: To ensure scientific rigor, only peer-reviewed journal articles and conference proceedings were included, while patents, technical white papers, and corporate reports were excluded. Because real-world industrial implementations of LfD are frequently protected as commercial trade secrets, the exclusion of the grey literature may have limited our ability to map the exact degree of current commercial adoption.
- Selection Bias and Single Screener Limitations: Since the screening, eligibility selection, and data extraction processes were conducted by a single primary researcher—rather than by two independent reviewers as recommended by the PRISMA 2020 guidelines [1]—there is an inherent risk of selection bias. To mitigate this limitation, a retrospective inter-rater reliability check was subsequently conducted by a second independent reviewer on a random 20% subset of records at the Title/Abstract stage and 100% of the full-text articles. The analysis demonstrated strong agreement (Cohen’s Kappa κ = 0.88 for Title/Abstract screening, and κ = 1.00 for full-text eligibility), indicating that the selection process was highly reliable. However, the lack of a second independent auditor during the initial, live extraction phase remains a minor methodological limitation. Although borderline or ambiguous cases were re-evaluated multiple times to ensure coding consistency and adherence to the eligibility protocol, the lack of a second independent auditor represents a methodological limitation.
- Temporal Concentration and Recency Bias: A notable limitation of this systematic mapping study is the high concentration of the literature in the most recent publication years, with 58.3% of the included primary studies (14 out of 24) published between 2024 and 2025. This temporal clustering reflects a surge in research interest driven by the maturation of collaborative robot hardware and the emergence of deep generative control policies. However, this introduces a potential recency bias, over-representing very recent and popular algorithmic paradigms (e.g., Diffusion Policies, Neural ODEs) that have not yet undergone long-term industrial testing. Furthermore, because these studies are very recent, they have had limited time to establish significant citation impact or undergo independent replication and verification by the wider research community. Consequently, some reported performance benchmarks must be treated as early, laboratory-validated indicators rather than long-term, industry-proven baselines.
6. Future Directions
6.1. Multimodal Perception and High-Dimensional Datasets
6.2. Bridging the Sim-to-Real Gap and Safe Exploration
6.3. Interactive Human-in-the-Loop (HITL) Systems
6.4. Control and Mechatronic Architectures for High-Frequency Deburring Dynamics
- Dual-Stage (Macro–Micro) Mechatronics: Standard robotic manipulators are limited by low joint control bandwidth (under 50 Hz) due to gear elasticity and link inertia, which is insufficient to suppress high-frequency chatter. Future systems should integrate high-bandwidth auxiliary actuators—such as piezoelectric end-effectors or active voice-coil-driven compliant tool holders operating at or above 1 kHz haptic loops [18]—to execute fast normal force adjustments, leaving coarse spatial path tracking to the macro-manipulator.
- Force-Coupled Specified Movement Primitives: Classical movement primitives (DMPs, ProMPs) must transition to force-coupled formulations, such as Specified DMPs (sDMPs) [17]. Directly embedding contact force feedback as a coupling term within the canonical and transformation equations allows the robot to dynamically scale feed rates or deform paths when encountering sudden force spikes.
- Adaptive Variable-Impedance Control (AVIC): During rigid workpiece–tool collisions, static impedance parameters lead to excessive force peaks or chatter. Control systems must adaptively modulate stiffness and damping coefficients online (e.g., via evolutionary optimization such as PSO [17] or residual Reinforcement Learning [5]) to increase damping at the moment of burr impact, absorbing kinetic energy and stabilizing the physical interaction.
7. Conclusions
Supplementary Materials
Funding
Data Availability Statement
Conflicts of Interest
References
- Page, M.J.; McKenzie, J.E.; Bossuyt, P.M.; Boutron, I.; Hoffmann, T.C.; Mulrow, C.D.; Shamseer, L.; Tetzlaff, J.M.; Akl, E.A.; Brennan, S.E.; et al. The PRISMA 2020 statement: An updated guideline for reporting systematic reviews. BMJ 2021, 372, n71. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Kim, Y.; Sloth, C.; Kramberger, A. Skill transfer for surface finishing tasks based on estimation of key parameters. In 2022 IEEE 18th International Conference on Automation Science and Engineering (CASE); IEEE: New York, NY, USA, 2022; pp. 2148–2153. [Google Scholar]
- Eiband, T.; Leimbach, L.; Nottensteiner, K.; Albu-Schäffer, A. Extraction of Robotic Surface Processing Strategies from Human Demonstrations. In 2025 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS); IEEE: New York, NY, USA, 2025; pp. 20494–20500. [Google Scholar]
- Min, K.; Ni, F.; Chen, Z.; Liu, H. A Force Control Method Integrating Human Skills for Complex Surface Finishing. Machines 2024, 12, 756. [Google Scholar] [CrossRef] [Scilit]
- Li, Y.; Lyu, Q.; Yang, J.; Salam, Y.; Wang, W. A Hybrid Framework Using Diffusion Policy and Residual RL for Force-Sensitive Robotic Manipulation. IEEE Robot. Autom. Lett. 2025, 10, 10266–10273. [Google Scholar] [CrossRef] [Scilit]
- Fischer, A.; Unger, C.; Kugi, A.; Hartl-Nesic, C. Few-Shot Learning of a Force-Based Industrial Cleaning Process using an Instrumented Tool. IFAC-PapersOnLine 2025, 59, 103–108. [Google Scholar] [CrossRef] [Scilit]
- Xu, X.; Qian, K.; Liu, A.; Yue, Z.; Huang, W. Polishing via ODEs: Adaptive admittance control for robot polishing based on Neural ODEs. J. Manuf. Process. 2025, 155, 428–442. [Google Scholar] [CrossRef] [Scilit]
- Si, W.; Jin, Z.; Lu, Z.; Wang, N.; Yang, C. A Stable Guidance Method for Teleoperation-based Robot Learning from Demonstration. In IEEE International Conference on Automation Science and Engineering (CASE); IEEE: New York, NY, USA, 2024; pp. 2376–2381. [Google Scholar] [CrossRef] [Scilit]
- Wang, Y.; Zheng, Z.; Chen, C.; Wang, Z.; Gao, Z.; Peng, F.; Tang, X.; Yan, R. Adaptive Tuning of Robotic Polishing Skills based on Force Feedback Model. In IEEE International Conference on Robotics and Biomimetics (ROBIO); IEEE: New York, NY, USA, 2023; pp. 1–7. [Google Scholar] [CrossRef] [Scilit]
- Ke, S.; Zhang, J.; Zhao, H.; Guo, Y.; Wei, Z.; Pan, J.; Ding, H. Visual-Guided Diffusion Policy and Mesh-DMP Integration for Robotic Freeform Surface Polishing. In International Conference on Intelligent Robotics and Applications; Springer: Singapore, 2025; pp. 77–90. [Google Scholar] [CrossRef] [Scilit]
- Wu, H.; Zhai, X.; Wu, X.; Gu, S.; Liao, Z.; Xu, Z.; Zhou, X. Learning Stable Nonlinear Dynamics and Interactive Force-Aware Variable Impedance Control for Robotic Contact Tasks. Procedia Comput. Sci. 2023, 226, 127–133. [Google Scholar] [CrossRef] [Scilit]
- Möhl, P.; Pratheepkumar, A.; Ikeda, M.; Pichler, A. Morphing based transfer of demonstrated surface finishing trajectories to point clouds of similar objects. Procedia Comput. Sci. 2025, 253, 1002–1011. [Google Scholar] [CrossRef] [Scilit]
- Wang, Y.; Chen, C.; Hong, Y.; Zheng, Z.; Gao, Z.; Peng, F.; Yan, R.; Tang, X. PI2-BDMPs in combination with contact force model: A robotic polishing skill learning and generalization approach. IEEE/ASME Trans. Mechatron. 2025, 30, 978–988. [Google Scholar] [CrossRef] [Scilit]
- Wang, Y.; Chen, C.; Peng, F.; Zheng, Z.; Gao, Z.; Yan, R.; Tang, X. AL-ProMP: Force-relevant skills learning and generalization method for robotic polishing. Robot. Comput.-Integr. Manuf. 2023, 82, 102538. [Google Scholar] [CrossRef] [Scilit]
- Kulak, T.; Silvério, J.; Calinon, S. Fourier movement primitives: An approach for learning rhythmic robot skills from demonstrations. In Proceedings of the Robotics: Science and Systems (RSS), Virtual, 12–16 July 2020. [Google Scholar] [CrossRef] [Scilit]
- Wu, H.; Zhai, X.; Zheng, H.; Liao, Z.; Xu, Z.; Zhou, X. Learning stability-guaranteed skill and adaptive control strategies from demonstrations for heterogeneous component robotic machining. J. Manuf. Process. 2025, 151, 506–520. [Google Scholar] [CrossRef] [Scilit]
- Parvizi, P.; Ugurlu, M.C.; Acikgoz, K.; Konukseven, E.I. Parametrization of robotic deburring process with motor skills from motion primitives of human skill model. In International Conference on Methods and Models in Automation and Robotics (MMAR); IEEE: New York, NY, USA, 2017; pp. 373–378. [Google Scholar] [CrossRef] [Scilit]
- Acikgoz, K.; Parvizi, P.; Donder, A.; Ugurlu, M.C.; Konukseven, E.I. Dynamic movement primitives and force feedback: Teleoperation in precision grinding process. In 10th International Conference on Electrical and Electronics Engineering (ELECO), Bursa, Turkey, 30 November–2 December 2017; IEEE: New York, NY, USA, 2017; pp. 722–726. ISBN 978-605-01-1134-7. [Google Scholar]
- Zhai, X.; Ou, Y.; Xu, Z.; Jiang, L.; Zhou, X.; Wu, H. Effective learning and online modulation for robotic variable impedance skills. In IEEE International Conference on Robotics and Biomimetics (ROBIO); IEEE: New York, NY, USA, 2022; pp. 1–6. [Google Scholar] [CrossRef] [Scilit]
- Zhang, R.; Xia, J.; Ma, J.; Huang, D.; Zhang, X.; Li, Y. Human-robot interactive skill learning and correction for polishing based on dynamic time warping iterative learning control. IEEE Trans. Control Syst. Technol. 2024, 32, 2310–2320. [Google Scholar] [CrossRef] [Scilit]
- Haninger, K.; Hegeler, C.; Peternel, L. Model predictive impedance control with Gaussian processes for human and environment interaction. Robot. Auton. Syst. 2023, 165, 104431. [Google Scholar] [CrossRef] [Scilit]
- Shen, N.; Mao, J.; Li, J.; Mao, Z. Research on trajectory learning and modification method based on improved dynamic movement primitives. Robot. Comput.-Integr. Manuf. 2024, 89, 102748. [Google Scholar] [CrossRef] [Scilit]
- Hamdan, S.; Aydin, Y.; Oztop, E.; Basdogan, C. Robotic Learning of Haptic Skills from Expert Demonstration for Contact-Rich Manufacturing Tasks. In IEEE International Conference on Automation Science and Engineering (CASE); IEEE: New York, NY, USA, 2024; pp. 2334–2341. [Google Scholar] [CrossRef] [Scilit]
- Liao, Z.; Tassi, F.; Gong, C.; Leonori, M.; Zhao, F.; Jiang, G.; Ajoudani, A. Simultaneously learning of motion, stiffness, and force from human demonstration based on Riemannian DMP and QP optimization. IEEE Trans. Autom. Sci. Eng. 2024, 22, 7773–7785. [Google Scholar] [CrossRef] [Scilit]
- Li, X.; Yang, C.; Feng, Y. The Generalization of Robot Skills Based on Dynamic Movement Primitives. IFAC-Pap. OnLine 2020, 53, 265–270. [Google Scholar] [CrossRef] [Scilit]
- Nemec, B.; Yasuda, K.; Mullennix, N.; Likar, N.; Ude, A. Learning by demonstration and adaptation of finishing operations using virtual mechanism approach. In IEEE International Conference on Robotics and Automation (ICRA); IEEE: New York, NY, USA, 2018; pp. 7219–7225. [Google Scholar] [CrossRef] [Scilit]
- Duarte, N.F.; Santos-Victor, J. Robot Imitation of Polishing Motions by Observing Humans: From Human Non-Verbal Cues to Stable Limit Cycles. In IEEE International Conference on Development and Learning (ICDL); IEEE: New York, NY, USA, 2024; pp. 1–7. [Google Scholar] [CrossRef] [Scilit]
| Criteria Category | Inclusion Criteria | Exclusion Criteria |
|---|---|---|
| Application Domain | Robotic deburring and polishing operations | Generic assembly, pick-and-place, free-space trajectory tracking, or heavy rough grinding/milling (bulk shaping) |
| Methodology | Learning from Demonstration (LfD, PbD, imitation learning, DMP, GMM, etc.) | Standard CNC programming, offline CAD/CAM paths, non-learning adaptive control |
| Evidence Type | Experimental validation or detailed simulation of finishing tasks | Abstract concepts without task validation, technical sheets, patents |
| Publication Medium | Peer-reviewed journal articles and conference papers | Abstract concepts without task validation, technical sheets, patents |
| Language | English language only | Abstract concepts without task validation, technical sheets, patents |
| Database | Search Query |
|---|---|
| Scopus | TITLE-ABS-KEY ((“Learning from Demonstration” OR “LfD” OR “Imitation Learning” OR “Programming by Demonstration” OR “DMP” OR “Dynamic Movement Primitives” OR “ProMP” OR “Probabilistic Movement Primitives” OR “Skill learning” OR “Task learning”) AND (“Deburring” OR “Polishing” OR “Surface finishing” OR “Surface-finishing” OR “Finishing” OR “Robotic finish*”)) AND PUBYEAR > 2015 AND PUBYEAR < 2027 AND (LIMIT-TO(DOCTYPE, “ar”) OR LIMIT-TO (DOCTYPE, “cp”)) AND (LIMIT-TO (LANGUAGE,“English”)) |
| Web of Science | TS=((“Learning from Demonstration” OR “LfD” OR “Imitation Learning” OR “Programming by Demonstration” OR “DMP” OR “Dynamic Movement Primitives” OR “ProMP” OR “Probabilistic Movement Primitives” OR “Skill learning” OR “Task learning”) AND (“Deburring” OR “Polishing” OR “Surface finishing” OR “Surface-finishing” OR “Finishing” OR “Robotic finish*”)) AND PY=(2016-2026) AND DT=(ARTICLE OR PROCEEDINGS PAPER) AND LA=(ENGLISH) |
| IEEE Xplore | (“Learning from Demonstration” OR “LfD” OR “Imitation Learning” OR “Programming by Demonstration” OR “DMP” OR “Dynamic Movement Primitives” OR “ProMP” OR “Probabilistic Movement Primitives” OR “Skill learning” OR “Task learning”) AND (“Deburring” OR “Polishing” OR “Surface finishing” OR “Surface-finishing” OR “Finishing” OR “Robotic finish*”) |
| Google Scholar | (“Learning from Demonstration” OR “Imitation Learning” OR “Programming by Demonstration” OR “Dynamic Movement Primitives” OR “Probabilistic Movement Primitives”) AND (Deburring OR Polishing OR “Surface finishing” OR “Surface-finishing” OR Finishing OR “Robotic finish*”) |
| No. | Authors | Year | Robot Type | Application Area | Algorithm Used |
|---|---|---|---|---|---|
| 1 | Min et al. 2024 [4] | 2024 | Franka Emika Panda | Polishing | Computed-torque impedance control |
| 2 | Li et al. 2025 [5] | 2025 | xArm 7 | Polishing | Diffusion Policy + Residual RL |
| 3 | Fischer et al. 2025 [6] | 2025 | KUKA LBR iiwa 14 | Polishing | ProMPs (Few-Shot) |
| 4 | Xu et al. 2025 [7] | 2025 | Universal Robots UR5 | Polishing | Neural ODEs (Hyper-NODEs) |
| 5 | Si et al. 2024 [8] | 2024 | Franka Emika Panda | Polishing | DS-based imitation learning |
| 6 | Wang et al. 2023a [9] | 2023a | Universal Robots UR16e | Polishing | PMDRNN + DMPs |
| 7 | Ke et al. 2025 [10] | 2025 | KUKA LBR iiwa 14 | Polishing | Vision-Diffusion+ Mesh-DMP |
| 8 | Wu et al. 2023 [11] | 2023 | Franka Emika Panda | Polishing | GMM-GMR, Var. Impedance |
| 9 | Möhl et al. 2025 [12] | 2025 | None (Point cloud morphing) | Polishing | Neural network morphing |
| 10 | Wang et al. 2025 [13] | 2025 | Universal Robots UR16e | Polishing | PI2-BDMPs |
| 11 | Wang et al. 2023b [14] | 2023b | Universal Robots UR16e | Polishing | AL_ProMP |
| 12 | Kulak et al. 2020 [15] | 2020 | Franka Emika Panda | Polishing | Fourier Movement Primitives (FMPs) |
| 13 | Wu et al. 2025 [16] | 2025 | Franka Emika Panda | Polishing | PC-GMM-DS, Var. Impedance |
| 14 | Parvizi et al. 2017 [17] | 2017 | Phantom haptic device | Deburring | Modified DMPs (sDMP) |
| 6 | Acikgoz et al. 2017 [18] | 2017 | PI 6-DOF Hexapod | Deburring | DMPs |
| 7 | Zhai et al. 2022 [19] | 2022 | Franka Emika Panda | Polishing | GMM-GMR, Var. Impedance |
| 10 | Zhang et al. 2024 [20] | 2024 | Rethink Robotics Sawyer | Polishing | DTW-ILC + GMM |
| 13 | Haninger et al. 2023 [21] | 2023 | Franka Emika Panda | Polishing | MPC + Gaussian Processes |
| 18 | Shen et al. 2024 [22] | 2024 | KUKA KR6 R900 | Polishing | FDC-DMP |
| 19 | Hamdan et al. 2024 [23] | 2024 | Universal Robots UR5 | Polishing | MLP-based force learning |
| 20 | Liao et al. 2024 [24] | 2024 | Franka Emika Panda | Polishing | Riemannian DMP + QP |
| 21 | Li et al. 2020 [25] | 2020 | Rethink Robotics Baxter | Polishing | DMPs + Vision |
| 23 | Nemec et al. 2018 [26] | 2018 | Yaskawa Motoman MH-6 | Polishing | Virtual mechanism + ILC |
| 24 | Duarte et al. 2024 [27] | 2024 | KUKA LBR iiwa 7 and Kinova Gen3 | Polishing | Dynamical system (Limit cycle) |
| Publication Year | Frequency | Percentage (%) | Cum. Percentage (%) |
|---|---|---|---|
| 2017 | 2 | 8.3% | 8.3% |
| 2018 | 1 | 4.2% | 12.5% |
| 2019 | 0 | 0% | 12.5% |
| 2020 | 2 | 8.3% | 20.8% |
| 2021 | 0 | 0% | 20.8% |
| 2022 | 1 | 4.2% | 25.0% |
| 2023 | 4 | 16.7% | 41.7% |
| 2024 | 7 | 29.2% | 70.8% |
| 2025 | 7 | 29.2% | 100% |
| Total | 24 | 100% | 100% |
| Country/ Region | Frequency | Percentage (%) | Key Institutions |
|---|---|---|---|
| China | 15 | 62.5% | Huazhong University of Science and Technology, Harbin Institute of Technology |
| Türkiye | 3 | 12.5% | Middle East Technical University, Koç University |
| Austria | 2 | 8.3% | PROFACTOR GmbH, Johannes Kepler University Linz |
| Germany | 1 | 4.2% | Fraunhofer Institute for Manufacturing Engineering and Automation |
| Slovenia | 1 | 4.2% | Jožef Stefan Institute |
| Switzerland | 1 | 4.2% | Idiap Research Institute/EPFL |
| Portugal | 1 | 4.2% | Instituto Superior Técnico, University of Lisbon |
| Total | 24 | 100% | - |
| Algorithmic Cluster | Frequency | Percentage (%) | Representative Methods |
|---|---|---|---|
| Dynamic Movement Primitives (DMPs) and Variants | 9 | 37.5% | FDC-DMP, B-Spline DMP, Riemannian DMP |
| Probabilistic and Statistical Models | 8 | 33.3% | GMM-GMR, AL-ProMP, Fourier MP, Gaussian Processes |
| Deep Learning and Generative AI | 4 | 16.7% | Diffusion Policy, Residual RL, Hyper-NODEs, MLPs |
| Autonomous Dynamical Systems (DSs) | 2 | 8.3% | Stable Limit Cycles, DS-based Imitation |
| Direct Impedance Control and Parameter Estimation | 1 | 4.2% | Computed-Torque Impedance Control |
| Total | 24 | 100% | - |
| Sensory Modality | Frequency | Percentage (%) | Key Hardware Elements |
|---|---|---|---|
| Force/Torque Sensing Only | 17 | 70.8% | 6-DOF F/T Sensors, Joint Torque Sensors |
| Vision Only (RGB/RGB-D) | 3 | 12.5% | Depth Cameras, Point Clouds |
| Multimodal (Force + Vision/Haptic) | 4 | 16.7% | RGB-D + F/T Sensor, Haptic Interface + Force Feedback |
| Total | 24 | 100% | - |
| Source | QA1 | QA2 | QA3 | QA4 | QA5 | QA6 | Total | Rating |
|---|---|---|---|---|---|---|---|---|
| Min et al. [4] | 2 | 1 | 1 | 1 | 2 | 2 | 9 | High |
| Li et al. [5] | 2 | 2 | 2 | 1 | 2 | 1 | 10 | High |
| Fischer et al. [6] | 2 | 2 | 2 | 2 | 1 | 1 | 10 | High |
| Xu et al. [7] | 2 | 2 | 2 | 1 | 2 | 1 | 10 | High |
| Si et al. [8] | 2 | 1 | 1 | 1 | 1 | 1 | 7 | Moderate |
| Wang et al. [9] | 2 | 1 | 1 | 1 | 2 | 1 | 8 | Moderate |
| Ke et al. [10] | 2 | 1 | 2 | 2 | 1 | 2 | 10 | High |
| Wu et al. [11] | 2 | 1 | 1 | 2 | 2 | 1 | 9 | High |
| Möhl et al. [12] | 2 | 1 | 2 | 2 | 1 | 2 | 10 | High |
| Wang et al. [13] | 2 | 1 | 2 | 2 | 2 | 1 | 10 | High |
| Wang et al. [14] | 2 | 1 | 1 | 1 | 2 | 1 | 8 | Moderate |
| Kulak et al. [15] | 2 | 1 | 1 | 1 | 1 | 1 | 7 | Moderate |
| Wu et al. [16] | 2 | 1 | 1 | 1 | 2 | 2 | 9 | High |
| Parvizi et al. [17] | 2 | 1 | 1 | 1 | 1 | 2 | 8 | Moderate |
| Acikgoz et al. [18] | 2 | 1 | 1 | 1 | 1 | 2 | 8 | Moderate |
| Zhai et al. [19] | 2 | 1 | 1 | 1 | 2 | 1 | 8 | Moderate |
| Zhang et al. [20] | 2 | 1 | 1 | 1 | 2 | 1 | 8 | Moderate |
| Haninger et al. [21] | 2 | 1 | 1 | 1 | 2 | 1 | 8 | Moderate |
| Shen et al. [22] | 2 | 1 | 1 | 2 | 2 | 2 | 10 | High |
| Hamdan et al. [23] | 2 | 1 | 1 | 1 | 2 | 1 | 8 | Moderate |
| Liao et al. [24] | 2 | 1 | 1 | 1 | 2 | 1 | 8 | Moderate |
| Li et al. [25] | 2 | 1 | 1 | 2 | 1 | 1 | 8 | Moderate |
| Nemec et al. [26] | 2 | 1 | 1 | 1 | 2 | 1 | 8 | Moderate |
| Duarte et al. [27] | 2 | 1 | 1 | 1 | 1 | 1 | 7 | Moderate |
| Fully Met (score = 2) | 24 | 2 | 6 | 6 | 14 | 6 | - | - |
| % Fully Met | 100% | 8.3% | 25.0% | 25.0% | 58.3% | 25.0% | - | - |
| Study (Ref.) | Robot Platform and Sensors | Workpiece Material and Geometry | Tooling and Normal Force Range | Specific Performance Metrics |
|---|---|---|---|---|
| [4] | Franka Emika Panda (7-DOF), joint torque sensors | Wooden violin body (highly curved) | Pneumatic sanding tool; | Successful skill transfer; roughness reduced below manual craft baseline |
| [5] | 7-DOF robot arm, overhead/wrist RGB-D cameras, F/T sensor | Basin (unstructured, variable geometry) | Cleaning brush / polishing | Cleaning success rate: 93% ± 3%; trajectory RMSE < 2.0 mm; correlation r > 0.95 |
| [7] | Franka Emika Panda (7-DOF), joint torque sensors | Curved workpieces (HT200 Cast Iron) | Polishing disk | Force fluctuation: mean = 0.0171 N; endpoint trajectory RMSE: 0.07046 mm (WP1)/0.0015 mm (sim); roughness : 14.39 μm to 5.29 μm (WP1, 63.2% reduction) and 2.63 μm to 0.459 μm (WP2, 56.1% reduction) |
| [9] | Franka Emika Panda (7-DOF), joint torque sensors | Curved mold surface | Polishing tool with force feedback | Force RMSE: 1.51 N (PMDRNN-DMP) vs. 2.14 N (PMNN) and 1.83 N (baseline) |
| [10] | Collaborative robot, RGB-D camera | Freeform curved workpiece | Polishing pad | High trajectory generalization on freeform meshes via Visual-Guided Diffusion Policy and Mesh-DMP |
| [14] | Collaborative robot, end-effector F/T sensor | Aluminum alloy plate | Polishing pad; variable speed-scaling | Roughness : reduced from 0.909 μm to 0.106 μm after 10 repetitions (AL-ProMP) |
| [20] | Collaborative robot, end-effector F/T sensor | Curved metal plate | Polishing tool; target polishing force | Force-tracking error: mean decreased to 0.18 N (SD: 0.12 N) via DTW-ILC + PBIC |
| [24] | Franka Emika Panda (7-DOF), joint torque + F/T sensor | Aluminum plate (height step) and brick | Polishing wool pad; variable target forces | Force RMSE: 0.8588 N (with QP) vs. 3.6290 N (without QP) on brick; 2.2619 N on aluminum |
| No. | Study | LfD Category | Core Commonalities with Corpus | Unique Differences and Key Contributions |
|---|---|---|---|---|
| 1 | Min et al. [4] | Direct Impedance | Focuses on polishing; utilizes force control and collaborative robot platforms. | Decouples human demonstrations into separate “motion skills” (discrete pose sequences) and “force skills,” validating on complex violin surfaces. |
| 2 | Li et al. [5] | Deep Learning | Focuses on cleaning/polishing; utilizes force control and collaborative robot platforms. | Combines a generative Diffusion Policy (for motion–force generation) with a Residual RL agent (for online force updates) using point clouds. |
| 3 | Fischer et al. [6] | Probabilistic | Focuses on cleaning; utilizes haptic interfaces and movement primitives. | Developed a location-invariant few-shot cleaning framework using an instrumented manual tool to capture expert data independent of the robot platform. |
| 4 | Xu et al. [7] | Deep Learning | Focuses on polishing; utilizes admittance control and collaborative robot platforms. | Uses Hyper-NODEs to generate smooth position–quaternion trajectories, combined with CLF/CBF for obstacle avoidance in LNP areas. |
| 5 | Si et al. [8] | Dynamical Systems | Focuses on polishing; utilizes force control and haptic teleoperation interfaces. | Introduces a dynamically stable energy-field-based virtual haptic guidance force that decays iteratively to reduce operator workload. |
| 6 | Wang et al. [9] | DMP | Focuses on polishing; utilizes force control and collaborative robot platforms. | Integrates a Phase-Modulated Diagonal Recurrent Neural Network (PMDRNN) to adaptively predict trajectory offsets based on force errors. |
| 7 | Ke et al. [10] | DMP | Focuses on polishing; utilizes force control and depth cameras. | Combines a Diffusion Policy to generate continuous spatial actions from RGB-D images and embeds them on freeform meshes using Mesh-DMP |
| 8 | Wu et al. [11] | Probabilistic | Focuses on polishing; utilizes force control and variable-impedance architectures | Learns a non-parametric, globally stable GMM for rhythmic motions, optimizing variable impedance via GMR to minimize control torque. |
| 9 | Möhl et al. [12] | Deep Learning | Focuses on trajectory transfer; utilizes depth cameras and vision data | Direct trajectory transfer between scan point clouds of similar objects using keypoint-driven neural network morphing, without CAD models. |
| 10 | Wang et al. [13] | DMP | Focuses on polishing; utilizes force control and collaborative robot platforms. | Introduces B-spline DMPs (BDMPs) requiring fewer basis functions, optimized via Policy Improvement with Path Integrals (PI2) for generalization. |
| 11 | Wang et al. [14] | Probabilistic | Focuses on polishing; utilizes force control and collaborative robot platforms. | Proposes Arc-Length ProMPs (AL-ProMP) to decouple force scaling and speed scaling in the spatial coordinate (arc-length) domain. |
| 12 | Kulak et al. [15] | Probabilistic | Focuses on polishing; utilizes collaborative robots and joint torque sensing. | Uses Fourier series basis functions (FMP) to learn periodic tasks from unaligned demonstrations without temporal or phase registration. |
| 13 | Wu et al. [16] | Probabilistic | Focuses on polishing; utilizes force control and variable-impedance architectures. | Tailored for Heterogeneous Material Components (HMCs); uses PC-GMM-DS and SMoGP to regulate rapid transitions across wood–iron splicing. |
| 14 | Parvizi et al. [17] | DMP | Focuses on deburring; utilizes haptic teleoperation interfaces | Uses Particle Swarm Optimization (PSO) to parameterize sDMPs to capture expert force responses under sharp corner and circular geometries. |
| 15 | Acikgoz et al. [18] | DMP | Focuses on deburring; utilizes haptic teleoperation interfaces. | Developed for deburring; human guides a 1-DOF haptic knob, and a high-speed piezoelectric actuator executes micro-adjustments on the workpiece. |
| 16 | Zhai et al. [19] | Probabilistic | Focuses on polishing; utilizes force control and collaborative robot platforms. | Couples GMM-GMR with a vector-valued Gaussian Process to enable online trajectory deformation under human physical intervention. |
| 17 | Zhang et al. [20] | Probabilistic | Focuses on polishing; utilizes force control and collaborative robot platforms. | Combines GMM with Dynamic Time Warping Iterative Learning Control (DTW-ILC) to estimate environment stiffness and update reference paths. |
| 18 | Haninger et al. [21] | Probabilistic | Focuses on co-manipulation polishing; utilizes force control and collaborative robot platforms | Captures task uncertainty using Gaussian Processes (GPs) and solves trajectory and impedance planning online using a non-linear MPC. |
| 19 | Shen et al. [22] | DMP | Focuses on polishing; utilizes force control and collaborative robot platforms | Introduces force-controlled dynamic coupling terms (FDC-DMP) using virtual coupling forces to dynamically alter local paths. |
| 20 | Hamdan et al. [23] | Deep Learning | Focuses on polishing; utilizes force control and admittance control | Employs a dual-force sensor configuration to isolate human guide forces (Fh) from environmental reaction forces (Fint), training an MLP. |
| 21 | Liao et al. [24] | DMP | Focuses on polishing; utilizes force control and collaborative robot platforms. | Uses Riemannian DMPs and QP optimization to simultaneously learn motion, 3D endpoint stiffness, and applied forces from a one-shot demonstration. |
| 22 | Li et al. [25] | DMP | Focuses on finishing; utilizes collaborative robot platforms and movement primitives. | Pairs DMPs with machine vision object detection to automatically recognize workpiece locations and generalize trajectory paths. |
| 23 | Nemec et al. [26] | DMP | Focuses on polishing; utilizes force control and haptic teleoperation interfaces | Models the tool as a “virtual mechanism” (augmented kinematic chain) for redundancy resolution, refining trajectories via Iterative Learning Control. |
| 24 | Duarte et al. [27] | Dynamical Systems | Focuses on polishing; utilizes collaborative robot platforms and human data. | Models all circular/ellipse polishing motions as a time-invariant dynamical system with a stable limit cycle attractor, mapping non-verbal human cues. |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the author. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Düzgün, E. Learning from Demonstration for Robotic Deburring and Polishing: A Systematic Mapping Study. J. Manuf. Mater. Process. 2026, 10, 293. https://doi.org/10.3390/jmmp10080293
Düzgün E. Learning from Demonstration for Robotic Deburring and Polishing: A Systematic Mapping Study. Journal of Manufacturing and Materials Processing. 2026; 10(8):293. https://doi.org/10.3390/jmmp10080293
Chicago/Turabian StyleDüzgün, Ercan. 2026. "Learning from Demonstration for Robotic Deburring and Polishing: A Systematic Mapping Study" Journal of Manufacturing and Materials Processing 10, no. 8: 293. https://doi.org/10.3390/jmmp10080293
APA StyleDüzgün, E. (2026). Learning from Demonstration for Robotic Deburring and Polishing: A Systematic Mapping Study. Journal of Manufacturing and Materials Processing, 10(8), 293. https://doi.org/10.3390/jmmp10080293
