2.1. Movement Description Language
Pose-derived angle proxies capture position information during movements. These proxies are rarely applied in machine learning and data science compared to skeleton data despite their advantages; for example, the joint angles are viewpoint- and subject-independent, individual movements are decoupled, and the angles at each joint are independent of the angles elsewhere [
20]. Pose-derived angle proxies are promising for applications where a human expert, such as a coach, has to interact with software and communicate poses and movements in terms of joint angles, as they enable more natural interaction and communication [
20]. Combined with velocity estimates, they provide comprehensive performance measurement.
Prior research demonstrates that strength training outcomes vary systematically with joint position [
21]. For squats, published studies report meaningful differences across knee flexion angles such as
,
, and
, while biomechanical reviews emphasize that squat depth is commonly evaluated using approximate joint angle ranges rather than a single fixed value [
22] and that deadlift knee flexion values are around
and
[
23]. In addition, monocular pose estimation introduces a non-negligible joint angle error, with the reported knee flexion error ranging from 9.3° to 25.8° depending on the model and setting [
24]. Accordingly, our default angle thresholds (15–30° tolerance) are informed by established exercise science ranges. This design allows the system to accommodate both pose estimation uncertainty and inter-individual variability, including users with reduced mobility or movement limitations. The thresholds remain user-adjustable to support different execution styles and subject-specific range of motion.
The proposed MDL formally models each exercise as a finite state machine with explicitly defined states, transitions, and threshold parameters. The MDL adopts best practices from formal modeling (e.g., UML/Statecharts and Modelica state machines) and movement analyses (e.g., Labanotation and MotionScript) to describe human motion phases in a concise, composable way. Each exercise is declared as an MDL module containing the following:
States—distinct movement phases (for example, in a squat, start, descending, and ascending).
Transitions—how the movement changes from one phase to another, based on conditions like specific pose-derived angle proxies.
Parameters or thresholds—numerical limits such as pose-derived angle proxies or durations that trigger these transitions.
The functional movement patterns fall into six categories: squat, lunge, hinge, push, pull, and carry [
5]. To create an MDL, the fundamental functional training exercises chosen were squats, deadlifts, shoulder presses, lunges, planks, push-ups, pull-ups, and bent-over rows, to understand all the principal movement categories. By using these exercises, it is possible to create more complex exercises like a thruster, which is a a squat followed by a shoulder press, or a burpee, which can consist of different sequences depending on how it is performed: it can consist of a squat followed by a plank and a push-up or begin with a deadlift-like motion that transitions into a plank and push-up. The exact structure can vary between individuals and can be adjusted by the user to match their preferred execution style.
The squat is considered a sequence of three states: the starting position (standing upright), where the athlete stands upright with the knee and torso angles typically above
; descending (flexing knees and hips), which is triggered when the knee angle and torso angle decrease below
, representing hip and knee flexion into the squat; and ascending (extending to return to the starting position), which is when the knee and torso angles re-extend above
as the athlete returns to the standing position [
25], shown in
Figure 1.
The deadlift is divided into two states: state 1, the starting position, where the torso leans forward (with a torso angle below
) with straight elbows (above
), and state 2, lifting, which is triggered when the torso is extended above
and the elbows remain straight (above
), shown in
Figure 2.
The lunge has three phases: phase 1, the starting position, where the individual is stood upright, with front and rear knee angles above
; phase 2, stepping, where the individual enters the lunge as both knee angles decrease to below
; and phase 3, returning, which is the recovery phase as the knee angles are extended again above
, shown in
Figure 3.
The shoulder press is broken down into three distinct states: the starting position (with the weights at shoulder level), where the elbow and shoulder angles are generally below
; the pressing phase (extending the arms upward and lifting the weights overhead), where the elbow and shoulder angles are extended above
; and the lowering phase (returning the weights to shoulder level), where the angles return below
. During the pressing phase, the system also verifies that the hands and elbows rise above the head, based on the relative position of key points, ensuring correct overhead extension, shown in
Figure 4.
The push-up is a three-phase cycle: phase 1, the starting position, where the body is in a plank position, and the elbow angle is above
; phase 2, lowering, which is triggered when the elbows are bent to below
; and phase 3, pushing, where the elbows are re-extended above
. In addition, the system verifies that the body remains extended in a horizontal alignment throughout the movement to ensure proper plank posture, shown in
Figure 5.
For the pull-up, there are three phases: phase 1, the starting position, where the arms are fully extended, the elbow angle is above
, and the torso is upright above
; phase 2, pulling, where the elbows are flexed below
, the chin is moved above the bar, and the torso may lean slightly below
; and phase 3, lowering, where the elbows are re-extended above
, and the torso returns to an upright position. In addition, the system verifies a vertical displacement of key points, ensuring that the body moves upward and downward along a consistent vertical axis, confirming a proper pull-up motion, shown in
Figure 6.
The bent-over row can be divided into three phases: phase 1, the starting position, where the torso is bent forward (with a torso angle below
), and the elbows are extended (above
); phase 2, pulling, where the elbows are flexed to below 100°, and the weight is lifted; and phase 3, lowering, where the elbows are extended (above 160°) to return to the starting position, shown in
Figure 7.
For the plank, which represents isometric exercise modeling, the primary state is a holding position (maintaining body alignment), with specific torso angle thresholds defining proper form maintenance. In addition, the system verifies that the body remains extended in a horizontal alignment throughout the movement to ensure proper plank posture, shown in
Figure 8.
Lastly, the box jump can be represented in three phases: phase 1, the starting position, where the individual is stood upright, with knee and hip angles above
; phase 2, jumping, where the knees and hips are flexed below
before propelling upward; and phase 3, stepping down, where the individual returns to the starting position as the joints are re-extended. To ensure that the jump is correctly recognized and counted, the system also checks for a vertical displacement of the body key points, confirming that the center of mass rises during take-off and lowers during landing, distinguishing a true jump from partial or incomplete movements, shown in
Figure 9.
By defining these primitive components, it is possible to combine them to create more complex exercises. For instance, a thruster can be created by sequencing a squat module immediately followed by a shoulder press module, as shown in
Figure 10.
Alternatively, a variant exercise, like an overhead squat, can be created by combining a squat with a shoulder press state, as shown in
Figure 11. The specific grip width (for example, a clean grip or snatch grip) affects the shoulder and elbow joint angles, as well as overall stability requirements. The literature shows that wider grip positions (similar to a snatch grip) increase shoulder abduction while decreasing shoulder flexion, whereas narrower grips (like a clean grip) keep the barbell closer to the body, increasing the requirement for shoulder flexion and stability. Moreover, the thoracic posture plays a critical role in shoulder mechanics; a more extended thoracic spine posture enhances scapular posterior tilt and upward rotation, thereby facilitating greater shoulder mobility in overhead positions [
26]. The MDL accommodates such variations through user-adjustable angle thresholds, enabling parameterization of exercise definitions for different execution styles without requiring new state machine definitions.
A burpee can be considered a complex exercise that can be described in various ways depending on how it is performed. One common description divides the movement into several phases: starting from a standing position, moving into a plank by placing the hands on the ground and jumping the feet back, performing a push-up, then jumping the feet forward toward the chest, and returning to a standing position with a final jump. Alternatively, a burpee can be performed in a lower-impact variation, where the person steps back into a plank position one foot at a time, performs a push-up, and then steps forward again to stand up without jumping.
This movement, as shown in
Figure 12, begins from a standing position, followed by lowering the upper body to the starting position of a deadlift as the hips and knees are bent to bring the hands to the ground. From this crouched posture, the legs are extended backwards into a plank, engaging the core and maintaining body alignment. The movement then transitions into a push-up, where the elbows are flexed and extended to lower and raise the torso. After the push-up phase, the athlete jumps or steps their feet forward toward their hands, returning to an upright position through the starting phase of a deadlift, and often finishes with an explosive vertical jump. Throughout the sequence, the burpee integrates the mechanics of the initial pose of a deadlift, followed by a plank, push-up, and jump, making it a complex exercise that develops strength, coordination, and cardiovascular endurance.
The MDL defines a grammar for functional training exercise descriptions, specifying valid syntax for exercise modules, phases, and conditions based on pose-derived angle proxies.
| exercise | → “exercise” ID “{” characterization + “}” |
| characterization | → “characterization” : STRING |
| phase | → “phase” INTEGER “{” description “}” |
| description | → “description” : STRING |
| conditions | → |
| condition | → ANGLE_ID COMPARISON NUMBER |
| ANGLE_ID | → “knee_angle” | “hip_angle” | “elbow_angle” | “shoulder_angle” |
| COMPARISON | → “>” | “<” | “≥” | “≤” |
| NUMBER | → [0–9]+ (“.” [0–9]+)? |
Each ANGLE_ID represents an angle computed from 3 key points () (where = nose, = left shoulder, etc.) extracted from MediaPipe’s 33 body landmarks or any pose estimator providing equivalent 3-key-point triplets.
- knee_angle
—hip-knee-ankle
- hip_angle
—torso-hip-knee
- elbow_angle
—shoulder-elbow-wrist
- shoulder_angle
—neck-shoulder-elbow
- torso_angle
—neck-hip-torso
where .
An MDL module represents a single exercise defined as a sequence of movement phases. Each phase is characterized by a set of conditions expressed as constraints on pose-derived angle proxies. A phase is considered active when all its conditions are satisfied. Transitions between phases occur when the corresponding conditions are met, following the natural temporal order of the exercise movement. In this way, each exercise can be modeled as a finite state machine in which phases represent states and the satisfaction of these conditions triggers transitions between them. To illustrate the structure of the MDL, Listing 1 presents an example definition for the squat exercise. The exercise is described through a characterization and a set of phases, where each phase contains constraints on the pose-derived angle proxies used to determine the progression of the movement. The angle values shown in the listing represent default thresholds and are user-adjustable within the graphical user interface (GUI) to accommodate different athletes and execution limitations.
| Listing 1. Example MDL definition of the squat exercise. |
exercise squat { characterization: ‘‘Standing upright, descending by flexing knees and hips, then returning to the starting position.’’ phase 1 { description: ‘‘Starting position’’ conditions: knee_angle > 160 # default, adjustable hip_angle > 160 # default, adjustable } phase 2 { description: ‘‘Descending’’ conditions: knee_angle < 100 # default, adjustable hip_angle < 150 # default, adjustable } }
|
The MDL definitions are interpreted during video analysis to determine the current phase of the exercise and to count repetitions. Algorithm 1 summarizes the phase detection procedure.
| Algorithm 1: Phase Detection and Repetition Counting |
![Jfmk 11 00162 i001 Jfmk 11 00162 i001]() |
2.2. System Architecture and Implementation
The proposed MDL framework was tested as a real-time, camera-based exercise analysis using pose landmark extraction with finite state machine modeling to enable accurate exercise recognition and repetition counting. The system architecture follows a modular design consisting of three main components: the pose estimation module, the finite state machine engine, and the user interface layer. The angle thresholds can be changed by the user.
Figure 13 illustrates the complete camera-based pipeline. For each frame, MediaPipe first estimates 33 body landmarks, which are used to compute pose-derived angle proxies (knee, hip, elbow, shoulder, and torso) used by the MDL definitions. The finite state machine then evaluates these angles against the exercise-specific conditions frame by frame, triggering state transitions and increasing the repetition counter upon completion of a cycle. The GUI provides real-time feedback, displaying phases and repetition counts.
The system uses MediaPipe, a pose estimation framework developed by Google, for camera-based real-time human pose detection and pose-derived angle proxy calculation. It should be noted that MediaPipe landmarks correspond to visually estimated key points. Therefore, the resulting joint angles represent approximate kinematic proxies derived from video. The MDL can process key points extracted by any pose estimation model, with MediaPipe used here for demonstration. MediaPipe provides robust performance across varying lighting conditions, backgrounds, and clothing, making it suitable for diverse training environments. The pose estimation module detects 33 key body landmarks from video frames and calculates pose-derived angle proxies using vector mathematics. Although MediaPipe is used in this implementation, the modular design can accept key points from other pose estimation models.
Pose-derived angle proxies are calculated using the interior angle formula between three points, where three landmarks, A, B, and C, form an angle at point B. The core innovation of the system lies in its finite state machine implementation, which models each exercise as a sequence of movement phases.
The finite state machine operates on real-time pose data, continuously evaluating pose-derived angle proxies against predefined conditions to determine the current exercise phase. State transitions are triggered when angle proxies cross specified thresholds, enabling precise tracking of exercise progression and repetition counting.
A key architectural feature is the configurable threshold system that allows users to customize angle parameters for different exercise variations and execution styles. The default threshold tolerance is set to 30 degrees globally, but users can modify specific angle thresholds for each exercise phase. This flexibility addresses the need for personalized exercise analysis while maintaining the precision required for functional movement assessment.
The threshold configuration is implemented through a hierarchical parameter system: global tolerance settings apply across all exercises, exercise-specific thresholds can override global settings, and individual pose-derived angle proxy limits can be modified for optimal performance.
The system processes video input through a pipeline designed for real-time performance. First, raw video frames are captured from the camera and passed to the pose estimation module, which extracts body landmarks for each frame. These landmarks are then normalized and organized into structured representations that feed the MDL system. The MDL specification encodes each exercise as a finite state machine defined by pose-derived angle proxy thresholds, allowing the system to track exercise phases and count repetitions. The evaluation module applies these rules frame by frame to determine whether the user’s movement satisfies the expected execution criteria and to log relevant events. Finally, the system presents the repetition counts in real time.
In real-time pose estimation, transient frame losses or noisy landmark detections often cause intermittent invalid states, leading the exercise state machine to briefly enter a “None” phase. This behavior can interrupt the logical sequence of movement phases, producing false transitions or missed repetitions. To address this, an adaptive phase persistence mechanism is implemented. Instead of immediately accepting a “None” value when a phase cannot be matched, the system temporarily retains the last valid phase for a short duration of N frames (typically N = 3–5). This persistence window allows the finite state machine to tolerate brief interruptions in pose detection while maintaining continuity in phase tracking. This approach effectively filters out transient tracking errors while preserving responsiveness to genuine phase changes. It is particularly beneficial in movements involving rapid transitions or partial occlusions. And, since the dataset includes videos recorded in wild environments, this adjustment is crucial to keep the state transitions stable despite short tracking interruptions caused by lighting variation, occlusion, or camera movement. Empirical tests showed that adaptive phase persistence eliminated nearly all spurious “None” transitions observed in the initial implementation, reducing false negatives and stabilizing repetition counting.
Adaptive phase persistence retains the last valid phase for up to N frames when no current phase satisfies the transition conditions. If a valid phase is detected before the timeout expires, the state machine resumes normally. Otherwise, the system transitions to “None” and re-evaluates the next frame.
This mechanism acts as a lightweight temporal stability filter, but it is not a full hysteresis model or a low-pass temporal smoother.
The system’s modular architecture enables easy addition of new exercises through standardized exercise definition structures. Each exercise is defined as a separate module containing its specific state machine configuration, making the system extensible and maintainable. The current implementation includes nine fundamental functional training exercises: squats, deadlifts, shoulder presses, lunges, planks, push-ups, pull-ups, bent-over rows, and box jumps.
Complex exercises can be created by combining primitive exercise modules, such as creating a thruster by sequencing a squat module followed by a shoulder press module. This compositional approach enables the representation of advanced functional training routines while maintaining the precision of individual movement analysis.
The GUI, illustrated in
Figure 14, represents the first version of the interface and enables users to intuitively configure exercises, adjust phase thresholds, and visualize results in real time. Through the GUI, users can edit exercise parameters; build composite movements (e.g., squat + shoulder press = thruster or squat + shoulder press (phase 2) = overhead squat), as illustrated; select between webcam or video input, in this case, by uploading a video of an athlete performing thruster repetitions; and monitor real-time repetition counting with live video feedback.
2.3. Experimental Evaluation
For experimental validation, a comprehensive dataset was assembled combining publicly available exercise videos and recordings to test the robustness of the Movement Description Language system across diverse exercise scenarios. The dataset was strategically constructed to include multiple exercise types, varying execution qualities, and different environmental conditions to thoroughly evaluate system performance. The evaluation dataset incorporates videos from several established public sources, including exercise demonstration videos from fitness platforms and competition footage. It includes video recordings from extreme conditioning program competitions, which provide high-quality examples of functional training exercises performed under standardized conditions. The links to these videos are available in the shared dataset presented, and they are described later in the article. These competition videos offer valuable ground truth data, as they feature athletes performing exercises with proper form validation by certified judges.
The dataset of extreme conditioning program competition videos includes recordings of fundamental movements such as squats, deadlifts, and other compound exercises that align with the MDL exercise definitions. These videos represent varying skill levels and execution speeds, providing robust test cases for evaluating the system’s ability to recognize exercise phases across different performance qualities and pacing strategies. The dataset includes some of the exercises defined in the MDL framework: squats, deadlifts, shoulder presses, planks, push-ups, and pull-ups. The dataset contains recordings from various settings, including home environments, commercial gyms, and competition venues, testing the system’s robustness to different lighting conditions, backgrounds, and camera angles.
2.3.1. Dataset Collection
The dataset was constructed by collecting videos from public datasets; from extreme conditioning program competitions; and recording videos of 7 participants performing squats, deadlifts, and shoulder presses and 2 participants performing bent-over rows, box jumps, burpees, lunges, push-ups, pull-ups, planks, overhead squats, and thrusters. All participants provided written informed consent to participate in the study and to allow the publication of the research details.
Videos of 9 participants were recorded at the collaborating MAREBox facility following institutional review board approval (CEIC-UA 6-CEIC-UA2026-D). Participants were recruited voluntarily through an open call shared with facility members, and all participants provided written informed consent for participation and publication of the recordings.
Recorded Videos (R): 240 videos of squats, 160 videos of deadlifts, 160 videos of shoulder presses, 2 videos of bent-over rows, 4 videos of box jumps, 6 videos of burpees, 3 videos of lunges, 5 videos of push-ups, 4 videos of pull-ups, and 2 videos of plank, thrusters and overhead squats.
Public Pose Datasets (P): 47 videos of deadlifts, 7 videos of planks, 26 videos of pull-ups, 106 videos of push-ups, 45 videos of shoulder presses, and 7 videos of squats [
27,
28,
29].
Competition and YouTube Videos (C): 24 videos of deadlifts, 49 videos of shoulder presses, 171 videos of squats, 18 videos of lunges, 6 videos of bent-over rows, 10 videos of box jumps, 7 videos of thrusters, 132 videos of overhead squats, and 3 videos of burpees.
The repetition-counting performance across exercises, including 95% Clopper–Pearson confidence intervals, is summarized in
Table 2.
2.3.2. Annotation Procedure
Manual repetition counts served as the ground truth for evaluating the MDL system’s performance on the 689 videos recorded specifically for this study (participant recordings of squats: 240, deadlifts: 160, shoulder presses: 160, bent-over rows: 2, box jumps: 4, burpees: 6, lunges: 3, push-ups: 5, pull-ups: 4, planks: 2, thrusters: 2, and overhead squats: 2; see
Table 2).
Annotations were performed by one annotator using CVAT (
https://app.cvat.ai (accessed on 10 April 2026)), an open-source computer vision annotation tool that enables precise frame-level labeling. The labels were created with the names of the exercises. For each video, the start and end frames of individual repetitions were identified based on exercise-specific movement criteria. For the example shown in
Figure 15, the start and end of the exercise (in this case, a deadlift) were identified. Each frame within this range was labeled as “deadlift”, from the beginning frame to the end frame, ensuring accurate annotation of the exercise represented. This frame-by-frame approach allowed verification of exercise execution quality and accurate repetition boundary detection.
The annotation task was limited to repetition boundary identification based on clearly defined movement criteria, reducing subjectivity and minimizing ambiguity in labeling.
To evaluate annotation reliability, a subset of three videos from the smartphone recordings dataset was independently annotated by a second annotator. Inter-annotator agreement was assessed at the repetition count level. Repetition count agreement for the same videos revealed identical repetition counts for all sequences (exact match rate of 100%, mean absolute error of 0, Pearson correlation , and Cohen’s ). These results further demonstrate the consistency of the annotation protocol and the reliability of the manually annotated repetition intervals used for evaluation.