1. Introduction
Industrial and collaborative robots are increasingly deployed in environments that demand flexibility, safety, and rapid task reconfiguration. As robots move from structured production lines into shared human workspaces, they must support frequent task changes, close-proximity interactions, and fast redeployment, making the efficient and safe adaptation of robot motion a central requirement.
Despite their growing adoption, programming robot motion remains challenging for non-expert users. Conventional tools such as teach pendants and offline programming require training and force users to map abstract coordinates and joint configurations to real workspace geometry, often leading to time-consuming trial-and-error processes and potentially unsafe motions. Augmented reality (AR) addresses these limitations by enabling in-workspace specification of robot intent through overlaid robot models and interactive controls, reducing cognitive load and improving usability compared to traditional interfaces [
1,
2,
3]. In a prior work, we introduced AURaPath, an AR system in which users can place and edit 6-DoF waypoints around a UR10 (Universal Robots, Odense, Denmark)digital twin, preview the resulting motion, and execute reachability-checked trajectories on the physical robot [
4]. AURaPath, like most AR robot-programming systems, relies on hand tracking as its sole input modality. Hand tracking is expressive but sensitive to lighting, occlusion, and the limited tracking volume of the headset, and it occupies the user’s hands during authoring.
To facilitate human–robot interactions, we look beyond hand tracking as AURaPath’s only input channel. The voice is a natural complementary modality: it is hands-free, requires no line of sight, and lets an operator issue commands while attending to the workspace. Simple voice control can be implemented with a fixed keyword vocabulary, but a fixed vocabulary cannot express the parameterized commands that authoring requires, such as which waypoint to move, along which axis, and by how much. Enumerating every combination as a registered phrase is not practical. Recent large language models (LLMs) can instead map free-form speech to structured commands, which removes the need to memorize exact phrasings. Placing an LLM in the control path of a physical robot, however, raises questions of reliability, safety, and deployment that must be answered before such a mapper can be trusted: How accurately does it map language to commands? Does it decline commands that it cannot safely interpret? Can it run under the cost and hardware constraints of real deployment settings?
This paper extends AURaPath with a voice modality that addresses these questions. The voice layer uses a two-tier design. An on-device keyword recognizer handles the safety-critical verbs (run, confirm, cancel, stop) behind a two-step confirmation gate, and an LLM maps the resulting transcript to the system’s structured command set. We evaluate the mapping layer with a benchmark that isolates language-to-command mapping from speech recognition and robot execution, so that the metric reflects the one new learned component. The benchmark then selects the backend that the deployed system uses.
Relative to the conference system [
4], the contributions new to this paper are the voice architecture, the structured command schema with its reject category, the labeled mapping benchmark, the local-versus-cloud deployment analysis, and the modality pilot; the AR authoring and digital-twin foundations are summarized from that prior work, rather than reintroduced. The scope of this work is the authoring layer—mapping language to a discrete, schema-constrained waypoint command—not low-level continuous control, which remains the responsibility of the UR10’s own controller. The two-tier safety design applies established safety-engineering practices (
Section 2) rather than proposing a new technique.
The main contributions of this work are as follows:
A two-tier voice modality for AR robot programming that combines a deterministic keyword layer for safety-critical verbs with an LLM mapper for free-form authoring commands, integrated with the existing AURaPath dispatcher.
A structured command schema, grounded in AURaPath’s 6-DoF waypoint representation, that constrains the LLM to a fixed command set and includes an explicit reject category for out-of-scope or under-specified input.
An objective benchmark of 169 labeled utterances that isolates language-to-command mapping from speech recognition and execution, together with a comparison of a keyword baseline, local models, and free-tier cloud backends on accuracy, false-accept behavior, and latency.
A deployment analysis under a zero-cost, remote-operator constraint, including a controlled comparison of the same model served locally and in the cloud, that selects the backend used in the deployed system.
2. Related Work
This work builds on prior research in augmented reality (AR) and mixed reality (MR) interfaces for robot programming, on the use of language models to map natural language to robot commands, and on safety mechanisms for language-controlled robots [
4]. The AR authoring and digital twin foundations are described in the base system [
4] and summarized in
Section 3; here, we focus on the work most relevant to the voice- and language-mapping extension.
2.1. Augmented and Mixed Reality for Robot Programming
Augmented and mixed reality have been widely explored as intuitive interfaces for robot teaching and end-user programming. Spatially registered AR interfaces reduce cognitive load and improve usability compared to teach pendants and desktop tools [
1,
2,
5], and many systems support waypoint or target pose specification through the manipulation of virtual robot models, spatial anchors, or holographic handles [
1,
2,
3], including a HoloLens 2 (Microsoft Corporation, Redmond, WA, USA) toolkit that generates and manipulates end-effector trajectories through AR waypoints for bi-directional human–robot collaboration [
6]. Several works target industrial programming specifically, using AR to raise the level of abstraction from joint coordinates to tasks and to support programming directly on the shop floor [
7,
8,
9]. A related line communicates robot motion intent to co-located humans through mixed reality, so that the operator can anticipate and verify planned motion [
10,
11]. More immersive systems combine spatial interaction with continuous feedback of robot motion, including AR-based trajectory preview before execution [
12] and end-to-end skill teaching and deployment [
13]. Programming by demonstration through head-mounted displays has also been studied as a way to author robot behavior without a pendant [
14,
15,
16]. Across this body of work, input is predominantly hand tracking or handheld controllers; voice, where present, is generally limited to a fixed vocabulary of command words for discrete actions rather than parameterized authoring.
2.2. Digital Twins for Motion Authoring and Preview
A parallel line of work pairs the AR interface with a digital twin of the robot, so that planned motion can be inspected on a synchronized virtual model before it runs on hardware. Virtual-reality digital twins have been used to program industrial robots by recording operator motions in the twin and reproducing them on the physical arm [
17]. Interactive mixed-reality digital twins have been used specifically for robot motion authoring [
18], and digital-twin representations support human–robot collaborative teaming and shared task understanding [
19]. Twin-based interfaces have also been extended to reconfigurable and soft robots, where the virtual model tracks a changing morphology [
20]. AURaPath follows this preview-before-execute paradigm: authored waypoints are previewed on a co-registered UR10 twin and executed only after reachability checking, and the voice layer in this paper drives the same authoring and preview loop [
4].
2.3. Large Language Models for Robot Command Mapping
Large language models have recently been used to translate natural-language instructions into structured robot commands, plans, or code. The closest system to our work is that of Fang et al. [
21], which combines a HoloLens 2, Unity, and a Universal Robots arm to generate waypoints from natural-language input, previews them in AR, and streams them to the physical robot. Their motivation, freeing users from a memorized command set, is close to ours. The approach differs in three ways that motivate the present work. First, they use the LLM as an open-ended generator that synthesizes waypoints from scene geometry, whereas we constrain the LLM to a fixed authoring schema and treat it as a mapper rather than a planner. Second, their system is validated through a single pick-and-place demonstration, with no labeled dataset, accuracy metric, or baseline, whereas our contribution is an objective benchmark of the mapping layer. Third, they do not compare deployment options or separate the safety-critical verbs from the language model. In short, the individual ingredients exist in prior work; the assembled and measured system we present does not.
Closest to our evaluation methodology, Huo et al. [
22] pair a fine-tuned LLM with a grammar-based canonicalizer to force outputs into a valid symbolic command format, and score the result against a fine-tuned API model and a grammar–NLU baseline on an existing corpus. Our setting differs in that the mapper is used off the shelf without fine-tuning, the schema is grounded in a deployed AR authoring system rather than a general corpus, and the input comprises free-form language commands in an interactive loop. On the deployment question, Gopee et al. [
23] compare a smaller on-robot model against a larger edge-server model for the speech-driven, function-calling control of a construction robot, and report that the server-hosted model is faster and more reliable despite the network hop. Our local-versus-cloud comparison reaches a similar conclusion for AR waypoint authoring on manipulation hardware. Das et al. [
24] benchmark multiple LLMs for natural-language robot navigation with an emphasis on latency percentiles, which supports the latency methodology we adopt.
2.4. Reliability and Safety of Language-Controlled Robots
Because a language model in the control path of a physical robot can produce unsafe or under-specified commands, several systems add an explicit reliability or safety layer. CLARA [
25] classifies user commands as clear, ambiguous, or infeasible and declines or seeks clarification rather than executing uncertain commands, using a labeled dataset of command types. SafeGate [
26] extracts safety-relevant properties from natural-language commands and applies a deterministic decision gate to authorize or reject execution, and reports that such a gate reduces the acceptance of defective commands while preserving the acceptance of benign ones. RoboGuard [
27] grounds predefined safety rules in the robot’s context and repairs unsafe plans before execution. These systems motivate two elements of our design: the explicit reject category in the command schema, scored through false-accept and false-reject rates, and the two-tier arrangement in which the safety-critical verbs are handled by a deterministic layer rather than the language model.
2.5. Positioning of This Work
The base AURaPath system integrates AR-based 6-DoF waypoint determination, spatially co-registered digital twin preview, and networked execution on a physical UR10 within a single workflow [
4]. This paper extends that system with a voice modality and characterizes the one new learned component it introduces. In contrast to open-ended language-to-waypoint generation [
21], we constrain the language model to a fixed authoring schema, evaluate the mapping in isolation from speech recognition and execution against labeled ground truth, compare local and free-tier cloud deployments, and keep the safety-critical verbs on a deterministic layer. We treat the choice of language mapper as a system-design decision answered by an objective benchmark, rather than proposing a new language-mapping technique.
3. System Overview
AURaPath connects a HoloLens 2 AR client, a PC middleware server, and a UR10 collaborative robot to support waypoint authoring, digital twin preview, and execution, as shown in
Figure 1. The base system is described in full in prior work [
4]; we summarize it here to the extent needed for the voice extension.
3.1. Voice-Interaction Layer
The voice layer adds hands-free authoring on top of the base system without altering the underlying workflow (
Figure 2). A microphone stream feeds two parallel recognition paths. The on-device keyword recognizer matches fixed vocabulary directly against the audio stream. For the open-ended authoring and navigation commands, speech is transcribed on-device by the Windows dictation recognizer (via MRTK), and the resulting transcript is passed to the LLM mapper. Both paths produce commands in a single structured format that a dispatcher applies through the same authoring, preview, and execution methods used by the touch interface. The underlying state machine remains the single source of truth, so voice adds a modality rather than a parallel control path. The benchmark in
Section 4 and
Section 5 evaluates only the transcript-to-command mapping stage; dictation and robot execution are held constant, as discussed in
Section 7.
3.1.1. Command Set and Interaction Design
The system exposes a fixed command set in three groups.
Navigation commands move between modes and menus (configure, trajectory, preview, run, exit, and the create, edit, and delete modes).
Authoring commands edit the waypoint list (create a waypoint, delete a waypoint, delete all, and offset a waypoint along an axis by a signed amount).
Execution commands are the safety-critical verbs that drive the robot (run, confirm, cancel, stop). The design goal is that the operator need not memorize exact phrasings: the LLM accepts the natural variants of these commands and confirms its interpretation before anything executes. Voice navigation commands move between the AR canvases shown in
Figure 3, and authoring commands operate within them. As noted in
Section 2, removing a memorized vocabulary is a long-standing goal in voice-based human–robot interaction; the contribution here is delivering it inside the AR authoring loop and measuring the mapping layer objectively.
3.1.2. Two-Tier Design: Deterministic Keywords and LLM Mapping
The voice layer is organized in two tiers that differ in how much they can be trusted. The first tier is an on-device keyword recognizer that runs offline and matches a small fixed vocabulary. It owns the safety-critical verbs—run, confirm, cancel, and stop—and always listens for stop. Because it is deterministic and local, its behavior is predictable and does not depend on a network connection or a model’s interpretation. The second tier is the LLM mapper, which handles the open-ended authoring and navigation commands where natural phrasing varies.
The safety-critical verbs are never delegated to the LLM. This separation is a deliberate design choice: the results in
Section 5 show that a language model can misclassify or over-accept commands, and a model that fabricates an execution or destructive command from an ambiguous phrase is unsafe in the control path of a physical arm. Keeping the safety verbs on the deterministic tier means an error in the LLM cannot, on its own, start or fail to stop the robot.
These are application-level safeguards, not a certified functional-safety system: the deterministic tier governs which verbs the LLM may issue, but it does not certify the reliability of speech recognition, and none of the voice-layer components are safety-rated in the IEC/ISO functional-safety sense. The UR10’s own certified safety functions—the physical emergency stop and the controller’s built-in safety monitoring—remain the safety-rated layer beneath this application logic. The two-tier separation applies established practice from prior safety-gated language-robot systems (
Section 2) rather than introducing a new safety technique.
3.1.3. Structured Output Schema
The LLM maps a transcript to a structured command object defined by the running system rather than by a generic benchmark. A command is one object identified by a type field, with four types.
Authoring commands target the exact 6-DoF waypoint representation used by the base system,
in the UR10 base frame (Equation (
1)); an offset command, for example, becomes {type: authoring, operation: offset, reference:
waypoint_id, axis:
z, offset:
}, with the offset expressed in SI base units (metres for
and radians for
), which is what the dispatcher applies to the selected waypoint.
Navigation commands carry an intent that maps to one of the system’s menu or mode transitions.
Execution commands carry one of the four safety verbs; at run time, these are handled by the deterministic tier, and including them in the schema lets the evaluation check whether the LLM correctly routes them rather than inventing an action.
Reject is a type-only object that the mapper returns for input that is out of scope, unsupported, or under-specified, so that the model declines rather than guesses. This is a label-based definition fixed at dataset construction time; whether a given backend actually produces a reject on these rows is the separate empirical question the false-accept and false-reject rates in
Section 5 answer.
The object scored in
Section 4 is identical to the object the dispatcher consumes at run time. This grounding is what makes the evaluation a measurement of the system’s mapping layer rather than a free-standing benchmark. The offset directions follow a fixed UR10 base-frame convention (up is
, forward is
, left is
, with the corresponding negatives); resolving directions relative to the operator’s viewpoint is out of scope, as discussed in
Section 7. The full command schema and the labeled dataset are provided as a companion specification.
3.1.4. Conversational Confirmation and Safety Gate
Execution is protected by a two-step gate that voice shares with the touch interface. Saying “run” arms the gate and opens a confirmation prompt; saying “confirm” within a short timeout executes the trajectory, and “cancel” or a timeout disarms it. The recognizer always listens for “stop,” which disarms the gate and issues a stop to the robot. For authoring and navigation commands, the LLM confirms its interpretation before applying it, so a misheard or mismapped command can be corrected before it takes effect. Because this mirrors the existing touch-confirmation flow, the voice modality adds hands-free input without weakening the safety model of the base system.
Figure 4 shows the base system authoring and executing a trajectory on the physical UR10; the voice layer described above drives the same dispatcher and safety gate pictured there.
3.2. AR Client (HoloLens 2)
The AR client is implemented in Unity on the HoloLens 2 using the Mixed-Reality Toolkit (MRTK) for hand tracking and spatial interaction (
Figure 5). Through it, an operator aligns a virtual UR10 twin with the workspace, places and edits 6-DoF end-effector waypoints around the twin, and requests a preview or an execution.
At the start of a session, the server queries the UR10’s current joint state and the client initializes the twin from it; the twin’s pose is then under the operator’s control and can be co-registered with the physical robot or placed arbitrarily in the scene. Waypoints are authored in the twin’s local frame and are mapped into the UR10 base frame only when a trajectory is sent for execution. Because Unity uses a left-handed,
y-up coordinate system while the UR controller uses a right-handed,
z-up system, a waypoint
expressed in Unity coordinates is converted to the UR10 base frame as
; orientations are sent as axis-angle rotation vectors, with a small-angle guard applied to avoid degenerate rotations near zero. Each waypoint is serialized as a pose in the UR10 base frame,
where
is the position and
is an axis-angle rotation vector, and the ordered list of waypoints defines the trajectory.
Interaction uses a minimal gesture set: a pinch grasps and repositions the digital twin and places waypoints, and a tap performs discrete UI actions such as switching modes or confirming a step. A canvas-based interface separates twin setup, waypoint determination, and preview/execution into distinct modes reached from a main canvas (
Figure 3).
3.3. PC Server (Middleware)
The PC server validates the waypoints it receives from the AR client, checks their reachability through the UR10 controller, and generates the corresponding executable commands. It communicates with the robot over URSocket using the Real-Time Data Exchange (RTDE) protocol, both to query the robot’s state and to execute reachability-checked trajectories. When a sequence is feasible, the server returns a solution that animates the trajectory on the digital twin so the operator can inspect it before committing, and the twin resets prior to execution (
Figure 6); if any waypoint in the sequence is infeasible, both the preview and the execution are rejected. Centralizing this robot-specific logic on the server keeps the AR client focused on interaction and lets the voice layer described in
Section 3.1 reuse the same authoring, preview, and execution paths without modification.
3.4. UR10 Robot
The UR10 executes the reachability-checked trajectories it receives from the PC server and streams joint-state feedback back to it. Communication is split across two links: the HoloLens 2 talks to the PC server over Wi-Fi, while the server talks to the UR10 over a wired Ethernet connection through a network switch. Execution safeguards include verifying reachability before a trajectory runs, honoring a stop command issued at any point, and applying timeout checks that halt execution if the communication link becomes unstable.
4. Experimental Setup
We evaluate the one learned component the voice pipeline adds, the language-to-command mapper, in isolation. Each backend’s structured output is scored against ground truth on AURaPath’s command schema. Speech recognition and robot execution are held constant by construction (transcript in, structured command out), so the metric reflects the mapping layer alone. As described in
Section 3.1.3, the object scored is the object AURaPath’s dispatcher consumes at run time, so this is a controlled isolation of the deployed mapper rather than a separate LLM benchmark. The system runs on the physical platform shown in
Figure 7.
The evaluation is designed to answer three practical questions about the mapper. The first is reliability: whether a constrained LLM can translate free-form commands into AURaPath’s structured command set accurately enough, per field and overall, to replace memorized commands in the AR authoring loop. The second is deployment: under a zero-cost, remote-operator constraint, whether the system should run a small model locally on a single consumer GPU or call a free-tier cloud API, given the resulting accuracy and latency. The third is value over the existing baseline: whether the LLM mapper meaningfully improves robustness to natural phrasing over the deterministic fixed-keyword layer, and whether it fails safe on input outside its scope. The benchmark, systems, metrics, and protocol below are built to answer these three questions together, and
Section 5 reports the results against each in turn.
4.1. Command Dataset and Ground Truth
The benchmark is a labeled set of 169 command utterances. Each has one unambiguous ground-truth object on AURaPath’s four-type schema (authoring, navigation, execution, reject;
Section 3.1.3). Of the 169 utterances, 61 are authoring commands, 51 are navigation commands, 28 are execution commands, and 29 are reject-worthy. In-scope commands are phrased in several natural ways, and offset magnitudes are varied (2–15 cm, 5–30°) so that a backend cannot memorize a single constant. A reject bucket of 29 rows (≈17%) contains out-of-scope, unsupported, and under-specified utterances, and tests whether the mapper declines rather than guesses. A phrasing-stress slice of polite, hedged, and colloquial wordings tests robustness beyond canonical templates. The dataset is generated deterministically from templated slots, which assigns each utterance’s one ground-truth object by construction, and is used as a test set only. No training or fine-tuning is performed, so every row is scored.
4.2. Systems Compared
Each backend is identified by a single letter, used consistently throughout the text and the tables that follow, so that individual model names do not have to be repeated. The letters group the backends by
how they are hosted and controlled rather than by which specific model they are, because that grouping is what the deployment question in
Section 4 turns on:
A is the deterministic keyword baseline already deployed on the HoloLens, included as the no-LLM reference point;
B groups the open-weight models running locally on the consumer GPU, representing an offline, private deployment;
D is the single cloud-hosted proprietary model (Gemini), representing a managed cloud-API deployment; and
E groups the models served on Groq’s cloud inference hardware, kept separate from
B so that the effect of the serving stack can be isolated from the effect of the model itself. (There is no backend C.)
All model-based backends share one system prompt and the same schema. The prompt is defined once and imported by each backend, so the comparison is not affected by prompt differences. Decoding is deterministic (temperature 0, with a fixed seed where the API exposes one). Reasoning is disabled where the model supports it, through a zero-thinking budget for Gemini and a no-think flag for Qwen3, since the task is slot-filling rather than reasoning.
A—keyword baseline: This is the vocabulary the deployed HoloLens keyword layer recognizes, extracted from the running system (21 literal phrases). It is not tuned to the benchmark and represents what voice authoring achieves today without an LLM.
B—Local LLM (Ollama, RTX 5060): This comprises four instruction-tuned models that fit in 8 GB of VRAM, across two generations: Llama 3.1 8B and Qwen2.5 7B (2024); Qwen3 8B and Granite 4 7B-A1B-H (2026) (the models were served with 4-bit quantization via Ollama). Decoding is schema-constrained.
D—Cloud LLM (Gemini 3.1 Flash-Lite, free tier): Here, decoding is schema-constrained through the API’s response-schema mode.
E—Cloud LLM (Groq): Two models served on Groq’s LPU hardware: Llama 3.1 8B, with the same weights as one local backend, included to separate the effect of the serving stack from the model; and GPT-OSS 20B, a larger open-reasoning model. Llama 3.1 8B is served in JSON-validity mode (a schema constraint is not offered for it), while GPT-OSS 20B is schema-constrained.
Frontier paid models such as GPT-4/5-class, Claude, and Gemini Pro are excluded by the zero-cost constraint rather than by capability. The question is what a zero-budget deployment can achieve.
Section 5 shows that a free-tier model reaches 95.3% exact-match on this task, so a paid tier is not required.
4.3. Metrics
We report exact-match accuracy, where all applicable fields must be correct and a type mismatch is a failure: the fraction of utterances where every applicable field of the predicted command matches the ground-truth command, with no partial credit for getting some fields right. For a benchmark of
N utterances, with predicted command
and ground-truth command
for utterance
i,
We also report per-field accuracy for operation, reference, axis, offset, intent, and verb: the fraction correct for each field, computed over the rows to which that field applies, independent of whether the rest of that same command is also correct. For a field
f, scored over the
rows to which it applies,
Offset is scored with a 1 mm (
) tolerance, and the sign must match. These are standard intent- and slot-filling evaluation measures [
28], applied here to a robot-authoring command schema rather than being newly proposed.
Reject handling is reported as two rates. The false-accept rate is the fraction of the
reject-worthy rows for which the mapper produced a concrete command instead of declining. In plain terms, this measures how often the mapper acts on input it should have refused—a false alarm that lets an out-of-scope or under-specified instruction through as if it were a valid command,
and the false-reject rate is the fraction of the
in-scope rows for which the mapper declined instead of producing the correct command. In plain terms, this measures how often the mapper wrongly refuses a command that was actually valid and in scope, forcing the operator to repeat it,
Latency (mean, median, p95) is measured for the mapping call only, and rate-limit backoff on the cloud APIs is excluded so that throttling is not counted as model latency. For a backend’s
n per-call latencies,
, sorted in ascending order as
,
Malformed output is counted as an exact-match failure and logged. These metrics are reported for every backend in
Section 5: exact-match and per-field accuracy, together with false accepts, under zero-shot prompting in
Table 1; the effect of adding in-prompt examples in
Table 2; and mapping latency in
Table 3.
4.4. Protocol
Three aspects are held constant across all model-based backends: the prompt, the schema, and the dataset. Two aspects differ, and we report them. First, the constraint mode: backends B and D use schema-constrained decoding, as does the GPT-OSS 20B model on Groq (E), which supports it; Groq offers only a JSON-validity mode for Llama 3.1 8B, so that one backend is not schema-constrained. To separate this effect, we also run the local Llama under JSON mode (
Table 4). Second, the prompt was revised once, to define the execution verbs by meaning after an initial version listed only their enum values. The revision was made before per-item scoring was inspected, so it could not have been shaped by observed test failures; it contains no dataset phrasings, and the few-shot example pool is disjointed from the test set.
Every backend, local and cloud alike, was run 10 times over the full 169-utterance benchmark at temperature 0. Reported accuracy, per-field, reject- handling, and latency figures are means over these 10 runs, with standard deviations reported alongside the means in
Section 5. Run-to-run variation is present for every backend, including the local models, and is not confined to the cloud APIs.
Beyond the mapping benchmark, we additionally conducted a small within-subjects user study of the full authoring loop, comparing hand-only and multimodal authoring; its protocol and results are reported in
Section 5.6.
5. Experimental Results
The experiments reported in this section primarily address reliability: determining whether a constrained LLM mapper can translate free-form commands into AURaPath’s structured command set reliably enough to replace memorized commands in the AR authoring loop. The deployment choice and the value of the LLM tier over the keyword baseline are addressed as well, as the accuracy, latency, and reject-handling results below depend on them.
5.1. Accuracy and the Keyword Floor
Table 1 reports zero-shot performance for every backend. The keyword baseline scores 21.9% exact-match and 0.0% on every authoring field. This is a structural limit, not a tuning problem. A fixed keyword grammar cannot carry parameters such as which waypoint, which axis, and how far without enumerating every combination as a pre-registered phrase, so the deployed vocabulary contains no authoring commands. Its false-accept rate is 0% and its false-reject rate is 94.3%: the keyword layer fails safe but authors almost nothing. This is the gap the LLM tier fills, and the reason the LLM sits behind the deterministic layer rather than replacing it. This comparison establishes the size of that gap against this system’s existing keyword layer; it is not a general claim that LLMs outperform deterministic NLU approaches, since a stronger slot-filling or grammar-based baseline was not implemented for this benchmark. Such a baseline is a reasonable extension but is out of scope for this evaluation (
Section 7).
Every LLM backend clears the baseline by a wide margin. The strongest local model, Qwen3 8B, reaches 81.7% exact-match. The cloud backend, Gemini 3.1 Flash-Lite, reaches 91.7%, with reference, offset, and axis all above 99% and operation the comparatively weaker field at 97.0%. Offset is the hardest field for every other backend, since it requires unit conversion, sign inference, and axis mapping together; Gemini is accurate enough across all four fields that this pattern does not hold for it.
The reject bucket separates the backends more clearly than accuracy does, and it reveals two opposite failure modes. Llama 3.1 8B accepts an average of 28.5 of 29 reject-worthy utterances as concrete commands across the 10 runs, in one case mapping an assent phrase to a delete_all. It is fluent but does not decline, which makes it unsuitable as the language-mapping tier in a safety-gated pipeline, independent of its accuracy. GPT-OSS 20B fails in the opposite direction: it over-declines, rejecting an average of 16.7 valid in-scope commands as zero-shot (6.4 with few-shot), while keeping false-accepts low. For a safety-gated pipeline, this is the less dangerous error, a refusal rather than a fabricated action, but it carries a coverage cost, since each false reject is a command the operator must repeat. Gemini sits between these extremes with an average of 3.7 false-accepts and 0.3 false-rejects per run (of 29 reject-worthy and 140 in-scope items, respectively), and is both the most accurate backend tested and has the lowest false-accept rate among them.
5.2. Few-Shot Ablation
Table 2 adds a disjoint 13-example demonstration pool to the prompt. Few-shot improves every model, and the gain is largest for the weaker models: Llama 3.1 8B gains 15.4 points, Granite 4 7.1, Qwen2.5 7.7, and Qwen3 6.5; the two Groq backends gain 10.7 (Llama 3.1 8B) and 10.1 (GPT-OSS 20B), and Gemini gains 3.6. Demonstrations compensate for limited capability, and a capable model gains little from them. The best configuration is Gemini few-shot at 95.3% exact-match, with reference, offset, and axis accuracy all at or above 99%. GPT-OSS 20B (Groq) few-shot is second at 88.8%, narrowly ahead of Qwen3 8B few-shot at 88.2%; the two are within one standard deviation of each other and should be read as tied rather than ordered.
5.3. Latency and Controlled Comparisons
Table 3 reports latency, and
Table 4 examines two variables the headline comparison combines. The latency result runs counter to the common assumption that local inference is faster. Gemini Flash-Lite averages 1.0 s, while the local models average ∼2.6 s on the RTX 5060. The same Llama 3.1 8B weights served on Groq average 0.3 s, about 9× faster than the identical model (and decoding mode) run locally (
Table 4). Because the local copy is 4-bit-quantized and the Groq copy is not (
Section 7), this gap reflects both the serving stack and quantization, not serving location alone.
Two results follow. Full-precision cloud serving of the same weights exceeds the 4-bit-quantized local copy by 5.7 points zero-shot and 1.0 point few-shot, which locates the residual gap in quantization and the serving stack once the model is fixed. Schema-constrained decoding also has little effect on accuracy for this model: exact-match is within 0.4 points of the JSON-only mode in both conditions. Malformed output, averaged over the 10 runs, was 0.0 for every schema-constrained backend (B—Llama 3.1 8B; B—Granite 4 7B; B—Qwen2.5 7B; B—Qwen3 8B; D—Gemini 3.1 Flash-Lite; E—GPT-OSS 20B); only the two backends not run under schema-constrained decoding produced any malformed rows: E—Llama 3.1 8B on Groq averaged 1.8 ± 0.8 per run, and the same model run locally under JSON-only decoding (
Table 4) averaged 0.8 ± 0.4 per run, out of 169 calls each. This is direct evidence that schema-constrained decoding, rather than the underlying model or where it runs, is what prevents malformed outputs; no backend was rendered unusable on this ground. Constrained decoding acts as a safeguard against rare malformed outputs rather than as an accuracy lever.
5.4. From Benchmark to Deployed System
The evaluation selects the backend AURaPath ships with. Across accuracy, false-accept behavior, and latency, Gemini 3.1 Flash-Lite (few-shot) is the strongest option: 95.3% exact-match, the lowest average false-accept count, a near-zero false-reject rate (0.2 on average), and 1.0 s mean latency, which is about 7 points more accurate than the best local model and more than twice as fast. The choice is not uniform across axes. Groq serves the same task at 0.3 s but at lower accuracy, and the local models offer offline, private operation at a modest cost in accuracy and latency. The decision follows the operating constraint. For a networked authoring loop in which a wrong command moves a physical arm, accuracy and false-accept behavior matter most, and the free-tier cloud backend is preferred. The selected mapper is integrated into the deployed pipeline shown in
Figure 4, and the base-system demonstration of authoring and executing trajectories on the physical UR10 that the voice layer now drives is provided in the project repository, closing the loop from the isolated benchmark back to the integrated system.
5.5. Comparison with Related Systems
Because the systems closest to ours use different datasets, robots, and command sets, a direct comparison of accuracy figures would not be meaningful. Instead,
Table 5 compares AURaPath with the most closely related systems on high-level design parameters: the interaction modality, whether the interface is AR-based, how the language model is used, the deployment setting, the evaluation method, and the presence of an explicit safety mechanism. The comparison shows that, while the individual capabilities appear in prior work, AURaPath is distinguished by the combination it assembles: a schema-constrained mapper evaluated against a labeled benchmark, a local-versus-cloud deployment comparison, and a two-tier design that keeps the safety-critical verbs on a deterministic layer.
5.6. Preliminary Modality Study: Hand vs. Multimodal Authoring
To complement the mapping-layer benchmark, which isolates the language model, we ran a small within-subjects study of the complete authoring loop to ask a practical question: does adding the voice modality make waypoint authoring easier and faster for a user, or does it get in the way? Five participants each authored trajectories at three, five, and seven waypoints under two conditions: hand-only AR authoring (the pinch-and-tap interface of the base system) and multimodal authoring, in which the participant was free to use their voice—the LLM mapper (local Qwen3 backend, the on-device configuration available for this study at the time) for authoring and navigation, and the deterministic keyword tier for the execution verbs—alongside hand interaction. The two backends share the identical schema, dispatcher, and safety gate, so the pilot speaks to the multimodal interaction loop rather than to the specific mapper that the benchmark selects (
Section 5). The five participants were a mix of graduate students and non-experts (three male, two female; all aged 25–35); three had prior AR/HoloLens experience and one had prior robot-programming experience. The condition order was counterbalanced across participants. We measured completion time, corrective actions, task success, and, for the multimodal condition, the number of times the LLM asked the user to rephrase a command. Usability was rated after each condition with the System Usability Scale (SUS) and the Single Ease Question (SEQ), the same instruments used in the conference study [
4]. This is a preliminary pilot; with five participants, we report means, standard deviations, and paired differences, and treat the results as indicative rather than statistically conclusive. The larger 20-participant study of the base (hand) interface against a teach pendant is reported in the conference paper [
4].
Table 6 reports the results. The two modalities perform comparably overall, and the informative pattern is how their difference changes with task size. On the shortest task (three waypoints), multimodal authoring essentially broke-even with hand authoring (
s,
): for a task this short, the fixed overhead of speaking a command, waiting for the mapper, and occasionally rephrasing a misunderstood command roughly cancels out the time saved over manual placement. As the task grows, the balance shifts in favor of voice commands: multimodal authoring is
faster at five waypoints and
faster at seven waypoints (
s), where manual placement and editing of many waypoints becomes the bottleneck and spoken offset edits are comparatively cheap. Task success at seven waypoints improved from three of five trials under hand authoring to four of five under multimodal, and corrective actions were slightly lower for multimodal at five and seven waypoints. The number of LLM rephrasings rose with task size (0.2, 0.8, and 1.6 on average at three, five, and seven waypoints), consistent with more commands being issued, but remained low enough that voice still reduced the total time on the larger tasks.
On the subjective measures, the two modalities were comparable, with a small edge for multimodal: SUS rose from
to
and SEQ from
to
(paired differences
and
). Both differences are small relative to their spread at this sample size, so we read them as evidence that adding the voice modality did not reduce the perceived ease of use, rather than as a demonstrated usability gain. Taken together, the pilot indicates that adding the voice modality did not reduce the perceived ease of use in this small sample and became advantageous for completion time as the authoring task grew, while never requiring the user to memorize commands. No paired difference reported here is statistically reliable at
;
Table 6 reports descriptive small-sample data only. These observations are preliminary given the sample size, and a larger comparative study of hand, voice, and hybrid authoring remains for future work.
6. Discussion
6.1. Summary of Findings
The results answer the three questions posed in
Section 4 directly. On reliability, a constrained LLM mapper translates free-form commands into AURaPath’s command set with high per-field accuracy: the deployed backend reaches 95.3% exact-match and near-perfect scores on reference, offset, and axis (all above 99%), which is sufficient to replace the memorized commands in the authoring loop. On deployment, the free-tier cloud backend is the choice under the zero-cost constraint, ahead of the local models on accuracy, false-accept behavior, and latency at once. On value over the keyword baseline, the LLM tier improves exact-match from 21.9% to 91.7% zero-shot, and the gap is categorical rather than incremental: the keyword grammar scores zero on every authoring field because it cannot represent parameters at all.
Two secondary findings are worth stating. First, the reject bucket discriminates between backends more sharply than accuracy. A model can be fluent and still unusable: Llama 3.1 8B accepts an average of 28.5 of 29 reject-worthy utterances and in one case maps an assent phrase to delete_all. The ability to decline, not raw accuracy, is the property that determines whether a model is acceptable for a pipeline in which a deterministic tier and the preview-before-execute confirmation gate remain responsible for execution safety. This result also supports the two-tier design: because the safety-critical verbs are handled by the deterministic keyword layer and never by the LLM, an over-accepting model cannot trigger execution on its own.
Mapping errors are also not uniform in consequence. A wrong reference or axis produces a visible, correctable action; a sign or magnitude error moves the waypoint the wrong way but within the same class of action; a spurious delete or an over-accepted execution intent is the most consequential failure mode, since it removes authored work or could arm execution on an unintended command. The false-accept results above concentrate in this last, most severe category, which is the practical reason the two-tier design keeps execution verbs off the LLM entirely, regardless of its accuracy.
Second, the latency result reverses a common assumption. The local models are slower than the cloud backend, not faster: ∼2.6 s on the RTX 5060 against 1.0 s for Gemini Flash-Lite, and 0.3 s for the same Llama 3.1 8B weights served on Groq’s LPU hardware. This same-weights comparison remains confounded by quantization—the local copy is 4-bit and the Groq copy is full-precision (
Section 7)—so the gap reflects both the serving stack and quantization rather than serving location alone; within that caveat, serving infrastructure, not physical proximity, is what makes the cloud and Groq backends faster for an interactive authoring loop.
6.2. Design Implications
Few-shot prompting helps every model, and helps the weaker ones most (the few-shot ablation in
Section 5), which suggests that a small demonstration set is a cheap way to raise a local model toward deployment quality when a cloud API is not an option. Schema-constrained decoding, by contrast, has little effect on accuracy for the models tested and acts mainly as a safeguard against occasional malformed output. A deployment can therefore treat constrained decoding as a reliability measure rather than a source of accuracy, and can rely on few-shot demonstrations when accuracy is the constraint.
7. Limitations and Threats to Validity
7.1. Scope Boundaries
The evaluation isolates the language-to-command mapper by holding speech recognition and robot execution constant. This is a deliberate boundary, not an oversight: it produces an objective, single-metric measurement of the one learned component the pipeline adds. Errors introduced by speech recognition in the field, and the interaction between recognition errors and mapping errors, are outside the present scope and left to future work.
Reachability checking, which the deployed pipeline performs before any trajectory runs (
Section 3), confirms that the UR10 can attain the commanded pose; it is not collision-checking or global motion planning, and no such checks are performed here. Mapping errors that pass reachability, and collisions with objects in the workspace, are outside the present scope.
The offset directions are defined in the UR10 base frame under a fixed authoring convention that the operator is expected to follow (
Section 3.1.3). Resolving directions relative to the operator’s viewpoint, which changes as the operator moves around the workspace, is out of scope. The research questions concern the accuracy and latency of language-to-command mapping, not egocentric frame resolution.
7.2. Evaluation Scope and Confounds
The backends are a convenience sample of what is available at zero cost, and the set is small: the results characterize free-tier and consumer-hardware options at the time of evaluation rather than LLMs in general, and paid frontier models are excluded by the cost constraint rather than being tested and found wanting. The specific rankings should be read as a snapshot, since the open-model landscape moves quickly, as the gap between the 2024 and 2026 local models in our own results shows.
Two variables are also not held constant across all backends. The constraint mode differs for the Groq backend, which offers only JSON-validity mode for Llama 3.1 8B rather than schema-constrained decoding; we address this with a controlled run of the local Llama under the same mode (
Table 4). The local-versus-Groq comparison of the same model additionally mixes quantization with the serving stack, since the local copy is 4-bit-quantized and the served copy is not, and the two effects are not separated here.
Accuracy, per-field, and latency figures are means over 10 independent runs of each backend at temperature 0 (
Section 4.4), with the run-to-run standard deviation reported alongside every mean in
Table 1,
Table 2 and
Table 3 rather than being left unquantified. That variation is small (at most ∼1.5 points of exact-match SD for any backend) and present for local and cloud backends alike, so it is not a local-versus-cloud asymmetry. The small between-backend gaps, such as the 95.3% (Gemini) versus 88.2% (Qwen3 8B, few-shot) exact-match difference, are backed by these repeated-run estimates rather than by single-run point estimates.
Finally, the prompt was revised once during development, to define the execution verbs by meaning rather than by enum value; the revision was checked to contain no dataset phrasings, and the few-shot example pool is disjoint from the test set by construction. The benchmark itself is a single test set of 169 utterances with no train/test split, which is appropriate for evaluating off-the-shelf mappers but limits the statistical precision of small per-category differences, and it reflects the authors’ own construction of natural phrasings rather than a corpus of operator language collected in the field.
8. Conclusions
This work extended AURaPath, a mixed-reality system for authoring 6-DoF waypoints on a UR10 arm, with a voice modality built on a two-tier design: an on-device keyword layer that owns the safety-critical verbs, and an LLM mapper that translates free-form language into the system’s structured command set. The contribution is one of integration and completeness rather than a new component in isolation. Mapping language to robot commands is an established capability; the value here is packaging a constrained, safety-gated instance of it into a complete AR authoring pipeline for untrained users, and characterizing the one new component well enough to make a deployment decision.
To that end, we evaluated the mapper against an objective, single-metric benchmark that isolates language-to-command mapping from speech recognition and robot execution. Across a keyword baseline, four local models spanning two generations, and free-tier cloud backends, the evaluation shows that a constrained LLM mapper replaces memorized commands reliably (up to 95.3% exact-match), that a fixed keyword grammar cannot express parameterized authoring at all (0% on every authoring field), and that the choice between local and cloud turns on the operating constraint rather than on accuracy alone. The reject analysis further shows that the ability to decline out-of-scope input, not raw accuracy, is what determines whether a model is suitable for the language-mapping tier behind the deterministic safety layer and preview-before-execute gate, and it justifies keeping the safety-critical verbs on the deterministic layer.
The benchmark selects the backend the deployed system uses, and the preliminary modality study (
Section 5.6) exercised the multimodal authoring loop end-to-end on the physical UR10 using the local Qwen3 backend, finding it usable in a small exploratory sample; the specific deployed mapper (Gemini) was selected by the benchmark rather than validated by the pilot, a point we note as a limitation. This closes the loop from the isolated measurement to the integrated system; the project repository provides the base-system demonstration together with the evaluation harness and dataset.
Future work follows the scope boundaries drawn above. The largest open direction is an end-to-end evaluation that reintroduces speech recognition and studies how recognition errors interact with mapping errors in the field. A preliminary two-condition pilot comparing hand-only and multimodal authoring is reported in
Section 5.6; a larger, fully powered comparative study of hand, voice, and hybrid authoring, reusing the interaction protocol of the conference study, remains future work and would measure the authoring experience rather than the mapper in isolation. Resolving offset directions relative to the operator’s viewpoint, rather than a fixed base-frame convention, would relax one of the present scope boundaries and broaden the range of natural commands the system accepts.