Next Article in Journal
An Evolving AI-Driven Ensemble Learning Framework for Sickle Cell Crisis Prediction Using MIMIC-III Data
Previous Article in Journal
PPG-FusionNet: A Dual-Branch Neural Architecture for Cuffless Blood Pressure Estimation from Photoplethysmography
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

An AI-Driven Framework for Automating SME Commercial Workflows with Robotics and Immersive Technologies

1
Department of Applied and Computer Sciences, Faculty of Applied Sciences and Creative Industries, Barleti University, 1001 Tirana, Albania
2
Department of Statistics and Applied Informatics, Faculty of Economics, University of Tirana, 1010 Tirana, Albania
*
Author to whom correspondence should be addressed.
Computers 2026, 15(9), 568; https://doi.org/10.3390/computers15090568 (registering DOI)
Submission received: 23 July 2026 / Revised: 14 August 2026 / Accepted: 17 August 2026 / Published: 29 August 2026

Abstract

Commercial operations across trading, import/export, logistics, and technology distribution are being reshaped by the convergence of artificial intelligence (AI), machine learning, multi-agent robotics, and Extended Reality (XR). Small and Medium-sized Enterprises (SMEs) feel this shift acutely: they face the same pressures as their larger competitors; labor shortages, high-SKU inventories that resist tidy categorization, narrow margins, and customer expectations set by Amazon-grade fulfilment, but rarely command the capital or the structured warehouse environments that make industrial automation straightforward. Existing frameworks for AI-driven automation and digital twins have been developed primarily for large-scale industrial settings and do not account for the capital, infrastructure, and organizational constraints specific to SMEs, leaving a gap in SME-scoped integration models. This research addresses that gap by asking how AI-driven robotics and immersive technologies can be integrated to optimize commercial workflows in SMEs operating in dynamic logistics and trading environments. The proposed framework is grounded in Sociotechnical Systems Theory, which treats technology and organizational workflows as jointly designed and mutually adapting, and follows a Design Science orientation in which the architecture itself is constructed as an evaluable artifact rather than a purely descriptive model. Methodologically, the study conducts a narrative synthesis of literature on embodied AI, computer vision, digital twins, VR training, and AR-assisted operations, combined with workflow analysis to identify where SMEs lose the most time and money. These are translated into a four-layer system architecture (perception, cognition, execution, integration) deployed through a four-phase implementation model: needs assessment, digital-twin and VR pre-training, controlled hardware pilot, and AR-supported scaling. The contribution of the study is twofold: conceptually, it brings together several technologies that are often discussed separately in the literature, while focusing specifically on the needs and constraints of SMEs while practically, it proposes a phased roadmap that can help SMEs adopt these technologies gradually, reducing both financial and operational risks. The approach also emphasizes human–robot collaboration rather than replacing human workers. The study does not include experimental validation, it presents a conceptual architecture and implementation roadmap consistent with a Design Science artifact-construction stage that can serve as a basis for empirical testing in real commercial environments.

1. Introduction

Artificial intelligence, robotic automation, and immersive technologies are quietly rewriting how modern commerce gets done. In logistics, warehousing, import/export, and technology distribution, companies of every size are layering intelligent systems on top of operations that were, until recently, run almost entirely by hand. The arrival of Industry 4.0 has only accelerated this shift, pulling AI-powered robotics, machine learning, computer vision, and Extended Reality (XR) into the daily mechanics of the supply chain. The familiar showcase examples; Amazon’s Kiva fleets, Walmart’s automated distribution centers, JD.com’s lights-out facilities, demonstrate what is technically possible when capital and structured environments are not constraints.
Most businesses, however, are not Amazon. Small and Medium-sized Enterprises (SMEs) operate under a fundamentally different set of conditions: tighter capital, fragmented digital infrastructure, warehouses that change layout from one quarter to the next, and inventories that mix hardware, software licenses, machinery components, and consumer goods on the same shelf. Traditional industrial automation, designed around fixed conveyors, uniform pallets, and predictable lighting, does not transfer cleanly into these settings. Yet SMEs face the same downstream pressures as their larger competitors: rising customer expectations for next-day fulfilment, persistent labor shortages in physical roles, growing SKU complexity, and the need to defend already-thin margins.
Recent advances in embodied AI, computer vision, and multi-agent coordination have begun to close this gap. Six-axis robotic arms paired with local vision models can now handle product recognition, inventory tracking, adaptive sorting, and packing with a flexibility that earlier generations of factory robotics simply did not have. Running on-device, rather than depending on a constant connection to a cloud service, also changes the economics: latency drops, sensitive operational data stays inside the building, and the system continues to function when the internet does not. In parallel, XR technologies, Virtual Reality (VR) for immersive workforce training and safety simulation, and Augmented Reality (AR) for real-time on-the-job guidance, make it realistic for non-specialist staff to work alongside these systems without first becoming roboticists. Recent work on this convergence in commercial settings argues that AI, AR, and VR are no longer separate strands of digital transformation but are increasingly bundled together in customer-facing and operational contexts, and that the principal barriers to wider SME adoption are not technical but organizational, perceived complexity, internal capacity, and financial constraint [1].
The academic literature on each of these strands is rich, but it is also fragmented. Warehouse robotics, XR-based training, digital twins, and enterprise AI platforms tend to be studied in isolation, and the small body of integration-focused work is dominated by large-scale industrial or manufacturing case studies. Comparatively little attention has been paid to how these technologies can be combined, sequenced, and de-risked for SMEs operating in commercial trading and logistics, settings characterized by high variability, modest budgets, and small teams that cannot absorb a failed deployment.
This research addresses that gap. It proposes a conceptual and implementation-oriented framework for integrating AI-driven robotics and immersive technologies into SME commercial workflows. The work is guided by a single research question: how can AI-driven robotics and immersive technologies be integrated to optimize commercial workflows in SMEs operating in dynamic logistics and trading environments? The study has three objectives: first, to examine the role of embodied AI and computer vision in automating warehouse and inventory management processes; second, to explore how XR technologies can support workforce training and human–robot collaboration; and third, to develop a phased implementation methodology that reduces the financial and operational risk of SME digital transformation.
The study makes three main contributions. First, it proposes an SME-oriented integrated architecture that brings together AI, robotics, digital twins, and XR. Second, it introduces a human-in-the-loop approach that separates AI-based decision-making from optimization, safety, and robot control. Third, it presents a phased implementation roadmap that considers the financial, technological, and organizational challenges SMEs face when adopting advanced automation.
The paper is conceptual and implementation-oriented, not experimental. It does not provide empirical results on detection accuracy, robot performance, latency, or ROI but it outlines the architecture, implementation steps, and evaluation criteria for future simulation, prototyping, and SME pilot testing.

2. Literature Review

Recent academic and industry research shows that artificial intelligence, robotics, and immersive technologies are no longer isolated fields. They are converging, sometimes deliberately, sometimes by accident of adjacent adoption, to reshape enterprise automation and commercial logistics. The benefits documented across these studies are familiar: better operational efficiency, more accurate inventories, sharper predictive analytics, and stronger human–machine collaboration. The open questions are less so. Researchers continue to debate how well these systems scale outside the demonstration environments they were built in, how much implementation effort they realistically demand, what they do to the workforce that operates alongside them, and how robust they remain when warehouse conditions shift week to week. Large enterprises have largely worked through these questions in public; the picture inside SMEs, with their leaner budgets, fragmented infrastructure, and less standardized workflows, is considerably less clear.

2.1. The Architecture of Multimodal Frameworks

Modern enterprise automation increasingly depends on multimodal frameworks, systems that can take in heterogeneous business inputs (visual data from cameras, structured records from ERP and inventory systems, natural-language workflow descriptions, and live sensor feedback) and act on them coherently. A recurring architectural choice in recent work is the use of multi-agent orchestration in place of a single monolithic model. Specialized agents are assigned distinct responsibilities: supervisor agents handle high-level coordination, orchestrator agents plan workflows, and execution agents carry out individual tasks. The motivation is partly engineering, small specialized models are easier to update and debug than a single large one, and partly conceptual: business workflows are themselves hierarchical, and the agent topology can be made to mirror them. Recent architectural work formalizes this pattern explicitly, proposing that agents function as specialized executors with clearly defined input/output contracts, decoupled from a central orchestrator, with human involvement treated as a regularized risk-management step; such designs are reported to shorten process cycle times and reduce errors at role interfaces [2].
A second architectural pillar is knowledge integration. Enterprise data is typically scattered across CRM, WMS, accounting, and email systems, each with its own schema and access controls. Recent frameworks pull these sources into activity-centric knowledge graphs that support real-time, human-in-the-loop feedback, so that an operator can intervene mid-workflow without having to query four separate dashboards. The localized Large Language Model (LLM), running on-premises rather than through a public API, is becoming a common component of this layer because it allows knowledge-graph reasoning to happen close to the data it depends on. “Recent work on hybrid neural architectures illustrates this integration challenge concretely: Chechkin et al. [3] propose a hybrid neural network transformer that combines heterogeneous representations for content detection and classification, offering a template for how attention-based and interpretable decision mechanisms might be incorporated into the perception and cognition layers of commercial automation systems an integration strategy largely absent from the multimodal frameworks reviewed above.”
There is, however, a live debate over how far this architectural flexibility actually scales. Proponents argue that multi-agent systems adapt to changing operational conditions in ways that rule-based automation simply cannot, when a SKU is renamed or a shelf is moved, the system can be told once rather than reprogrammed. Critics counter that highly decentralized AI architectures introduce their own costs: agents need to interoperate, message protocols drift, and the cumulative maintenance burden can outweigh the benefit, particularly for SMEs that do not have a dedicated ML platform team.
A related debate concerns the choice between cloud-hosted and locally processed AI. Cloud architectures offer effectively unlimited compute and convenient model updates, but they introduce latency that can be fatal for closed-loop robotic control, depend on a stable internet connection that smaller warehouses cannot always guarantee, and concentrate sensitive operational data in third-party infrastructure. Local processing reverses each of these trade-offs: lower latency, better data sovereignty, graceful degradation when the network fails, at the cost of capital expenditure on edge hardware and a steeper internal skills requirement. For warehouse robotics in particular, where object recognition and trajectory planning need to complete in tens of milliseconds, the latency argument tends to win.
Most existing studies, however, examine these architectural choices in industrial manufacturing, environments where workflows are predictable, products are standardized, and the operating envelope is narrow. The dynamic commercial trading setting that this research takes as its subject, diverse inventories, layouts that shift with each season, irregular product geometries, has received comparatively little attention. The framework proposed here addresses that gap by emphasizing modular integration, locally hosted computer vision, and adaptive robotic automation scoped specifically to SME logistics and warehouse operations.
Figure 1 illustrates this kind of layered pattern in practice, showing the three-tier TwinXR architecture proposed by Tu et al. [4], which links a Smart Factory, its digital-twin (DT) description, and an XR application layer.

2.2. AI-Driven Robotics in Unstructured Commerce

The intelligence layer of modern robotics depends on AI for a stack of related jobs: predictive analytics, demand forecasting, object recognition, adaptive manipulation, and trajectory planning. These challenges become especially apparent in unstructured commercial environments, where products differ significantly in shape, size, weight, and packaging across orders. Traditional rigid automation, built around the assumption that the next item to arrive looks like the previous one, tends to fail in this setting. Research into embodied AI offers a way forward: by tightly coupling local computer vision networks with robotic control loops, robots can perform real-time object detection, adaptive grasping, and dynamic path planning, and continue to function when shelf layouts, lighting, or product mixes change without warning. Mobile robots and 6-axis arms equipped with sensor fusion can navigate autonomously, run continuously, and absorb material handling and other repetitive operations around the clock. Reported impact at scale is substantial: deployments in high-volume settings have cut operational costs by up to 20% and lifted inventory accuracy by more than 30%. At the manipulation level, memory-efficient deep-learning methods for grasping-point detection now let a single vision network evaluate feasible grips on randomly arranged, non-uniform objects and select the optimal grasp for different end-effectors such as vacuum cups or parallel grippers, a capability directly relevant to the mixed, high-SKU inventories that SME warehouses handle [5].
One such implementation, the HoloRobot mixed-reality application, is captured in Figure 2 from within the Unity development environment.
In retail environments, AI-powered robotic systems like OSHbot use voice recognition, image recognition, and machine learning technologies to actively welcome customers, answer questions, and provide personalized shopping guidance [6].
Customer-facing robots of this kind mark a shift toward continuously available retail assistance with reasonably consistent service quality. The most significant impact of AI-driven robotics appears in supply chain and warehouse operations, where mobile robots equipped with advanced sensors and AI-driven decision-making capabilities navigate autonomously through large spaces, performing material handling, transportation, and repetitive operations with remarkable precision. These systems can work 24/7, increasing the speed and accuracy of order processing while reducing operational costs [7].
The software architecture behind these implementations, the four core TwinXRscripts (GlobalInstance.cs, JsonReader.cs, JsonWriter.cs, and SimpleJSON.cs) that manage digital-twin data exchange between the Smart Factory and the XR application, is outlined in Figure 3.
Notable implementations have demonstrated substantial results, with Amazon’s AI-driven warehouse robots cutting operational costs by 20% while improving inventory management accuracy by 30%. The intelligence layer of these robotic systems relies heavily on AI and machine learning for predictive analytics and demand forecasting, which optimize inventory management and enhance production scheduling while minimizing inefficiencies. These AI capabilities have enabled retailers to achieve remarkable efficiency gains, including an 18% reduction in inventory carrying costs through improved demand forecasting and a 25% increase in production output with 10% quality control improvements in automated assembly lines [8].
A second implementation is shown in Figure 4, the HoloCrane application, demonstrating how the same TwinXR framework extends to a different piece of warehouse equipment.
Modern smart shopping systems represent the evolution of AI-driven robotics into personalized customer experiences, with autonomous robotic assistants using real-time tracking algorithms and reinforcement learning to follow customers throughout stores while providing personalized recommendations and navigation assistance. The integration of these technologies has been accelerated by recent global events and labor shortages, leading major retailers like Amazon, Ocado, and Walmart to rapidly adopt robotic systems for packaging, sorting, and distribution tasks [9].
The six-axis robotic manipulator configuration referenced throughout this paper, along with its rotational degrees of freedom, appears in Figure 5.

2.3. Extended Reality (XR) and Human–Robot Interaction (HRI)

Extended Reality (XR) technologies, Virtual Reality (VR), Augmented Reality (AR), and Mixed Reality (MR), bridge the gap between the physical warehouse and its digital representation. VR has proven particularly effective at producing usable digital twins of warehouse environments: operators can practise complex tasks in a safe simulation before they touch a live system, which materially improves procedural retention and shortens onboarding. AR plays a complementary role on the active floor, providing guided picking through smart glasses, surfacing inventory that is otherwise hidden behind a shelf or buried in a system, and walking technicians through maintenance protocols step by step. When AI sits behind the XR layer, the relationship between operator and robot starts to look less like supervision and more like collaboration; AR can render a robot’s intended trajectory in the operator’s field of view, so that a human can intervene before, rather than after, an unsafe motion.
This workflow, VR-based staff training ahead of deployment, AI-driven robotic picking during normal operations, and AR-guided human intervention when exceptions arise, is summarized in Figure 6.
Empirical work on XR in industrial settings is consistent on the broad direction even where it disagrees on magnitudes. VR-based simulations improve employee onboarding, procedural learning, and safety training when workers can interact with digital warehouse environments before real deployment; immersive 3D scenarios let staff rehearse real-world tasks in a safe, repeatable setting and emerge measurably better prepared for the live floor [10]. AR overlays delivered through wearable devices likewise reduce error rates on order-picking tasks compared with traditional paper or screen-based instructions. The mechanism is not particularly mysterious: contextualized, in situ information lowers the cognitive cost of switching between an instruction and the object it refers to.
That said, the literature is far from uncritical. A recurring counterweight to the enthusiasm is the question of long-term wear: prolonged headset use is associated with motion sickness, hardware discomfort, cognitive overload, and eye fatigue, and the magnitude of these effects depends on the device, the task, and the individual. Several studies argue that the productivity gains seen in short pilots tend to attenuate when the same workers run an XR-assisted shift day after day, which is a meaningful caveat for any deployment plan that assumes the pilot numbers will hold.
Scalability and integration are the other open questions. XR systems perform predictably in laboratory and pilot conditions; rolling them out into a live warehouse exposes a less forgiving set of constraints, hardware capex, integration with existing WMS and ERP systems, infrastructure for tracking and rendering, and the maintenance burden of a heterogeneous device fleet. Concerns over workforce experience cut in two directions at once: some scholars treat XR-assisted automation as an opportunity for upskilling, while others warn that the same data layer that enables AR guidance can enable fine-grained worker surveillance, and that the line between the two is mostly a matter of policy rather than technology.
The framework proposed in this research responds to these concerns by adopting an explicitly human-centered posture. VR is scoped to workforce training and simulation-based onboarding, where the wear-out and ergonomic risks are bounded by session length. AR is reserved for real-time human intervention during robotic workflows, surfacing information when an operator is already in the loop, rather than imposing a continuous overlay across the working day. By embedding XR inside a phased implementation strategy rather than a single-shot rollout, the framework leaves room to recalibrate as the workforce, the hardware, and the tasks evolve together.

2.4. Enterprise Workflow Automation Applications

The practical face of multimodal automation in enterprise settings is more varied than headline case studies suggest, spanning business functions and levels of technical depth. Cognitive Business Process Management represents one of the most comprehensive application areas, where systems can automatically discover client business applications, understand dependencies on IT components, and interpret business processes from unstructured information sources. These cognitive BPM systems use natural language processing to digest textual descriptions of processes in documents and deploy intelligent assistants that can identify process steps, ordering relationships, ownership structures, and connections to enterprise systems and applications. Comparable applications have been documented in public-sector and integration-oriented settings, where AI is used both to automate routine administrative workflows and to identify procedural gaps that human reviewers would otherwise miss [11]; see also [12].
Democratized AI Automation through no-code platforms addresses the practical constraints and entry barriers that prevent widespread AI adoption in enterprises. These multimodal LLM-based multi-agent systems enable users without programming knowledge to easily build and manage AI systems for automating and improving business processes [13].
More advanced implementations allow non-technical users to describe workflows in structured natural language, with AI agents interpreting business intent and interacting with enterprise systems through APIs. Enterprise integration workflows can particularly benefit from agentic architectures that automate complex workflows spanning multiple business systems [14]. A related strand pairs conventional robotic process automation with LLM agents that supervise bot behaviour in real time, flagging exceptions and adjusting execution rules without a human having to rewrite the underlying automation, which points toward a lighter-weight route for SMEs already running rule-based RPA to add adaptive oversight incrementally [15].
These applications often incorporate formal verification modules that analyze generated programs to prove properties like “A contract is never sent unless the user is created in Okta,” ensuring automation is both functional and reliable. The systems also support adaptive process composition that can handle addition, change, or removal of services during execution, including integration with crowdsourced workforce platforms [12].

2.5. Theoretical Foundation: A Sociotechnical Systems Perspective

The technologies reviewed above, computer vision, embodied AI, digital twins, and extended reality, are frequently discussed in the literature as discrete technical modules, evaluated primarily on performance metrics such as accuracy, latency, or throughput. This treatment is insufficient for SME contexts, where adoption outcomes depend as much on organizational readiness, workforce acceptance, and workflow fit as on the technical performance of any individual component. To account for this, the present study grounds its conceptual framework in Sociotechnical Systems Theory, the view that technical and social subsystems within an organization are interdependent and must be jointly optimized rather than designed in isolation, as established in the foundational literature reviewed systematically by Sony and Naik, [16]. This perspective has recently been extended to AI-driven industrial contexts by Freire et al. [17], who frame human–robot collaboration in industrial settings as a sociotechnical system in which agency is distributed between human operators and non-human agents such as robots and sensors, rather than assumed to reside with human operators alone. Applied here, the technical subsystem comprises the perception and cognition layers described in Section 3, the sensing, detection, and decision-making components that determine what the system can perceive and infer. The social subsystem comprises the execution and integration layers, where human operators, existing workflows, and organizational systems (WMS/ERP) absorb and adapt to automated decisions, and where human–robot collaboration is negotiated in practice rather than assumed by design.
This theoretical lens carries two implications for the framework developed in this paper. First, it justifies the four-layer architecture not merely as a convenient engineering decomposition, but as a structure that separates while explicitly linking technical and organizational subsystems, consistent with sociotechnical design principles. Second, it motivates the phased implementation model presented in Section 4, in which technology is introduced incrementally (needs assessment, simulated pre-training, controlled pilot, scaled deployment) so that organizational and workforce adaptation can proceed alongside technical deployment rather than being treated as an afterthought. This positions the contribution of the paper as a Design Science artifact: the framework itself is the evaluable construct, developed through synthesis of existing technical and organizational literature, and offered as a basis for future empirical evaluation rather than as a validated solution.

2.6. Industrial Reference Architectures

Reviewer feedback rightly noted that the novelty of a four-layer architecture cannot be established without situating it against established industrial reference architectures. This section compares the proposed framework with three widely used architectures, RAMI 4.0, the Industrial Internet Reference Architecture (IIRA), and ISA-95, and with digital twin reference architectures, to clarify which elements of the proposed framework are genuinely novel and which build on existing architectural principles.

2.6.1. RAMI 4.0

The Reference Architecture Model Industrie 4.0 (RAMI 4.0) provides a three-dimensional reference architecture for Industry 4.0, organizing industrial systems along hierarchy levels, life-cycle/value-stream considerations, and functional layers [18]. Its purpose is to provide a common architectural language and to support interoperability and structured Industry 4.0 implementation. RAMI 4.0 is highly relevant to the proposed framework because both approaches emphasize layered organization and the connection between physical and digital assets. However, RAMI 4.0 is a general reference architecture rather than an SME-specific implementation framework for AI-enabled commercial warehouses: it does not by itself specify how local LLMs, multi-agent reasoning, computer vision, robotic manipulation, XR-based intervention, or SME deployment gates should be organized. The proposed framework should therefore not be presented as a replacement for RAMI 4.0. Instead, RAMI 4.0 can provide a reference architectural foundation, while the proposed framework specializes the application logic for dynamic SME commercial environments.

2.6.2. Industrial Internet Reference Architecture (IIRA)

The Industrial Internet Reference Architecture (IIRA) provides a common architectural framework for interoperable Industrial Internet of Things systems across diverse industrial domains [19]. Its relevance to the present study lies in its emphasis on system viewpoints, interoperability, functional decomposition, and the integration of heterogeneous industrial components. However, IIRA is intentionally broad and domain-independent, and consequently does not prescribe the detailed configuration of a commercial warehouse architecture involving multimodal perception, agentic reasoning, robotic execution, XR, and SME-specific implementation stages. The proposed framework builds on the same principle of functional decomposition but introduces a more operationally specific structure for SME commercial workflows.

2.6.3. ISA-95

ISA-95 focuses on enterprise-control system integration and defines models and interfaces for information exchange between enterprise functions and manufacturing/control functions [20]. The standard is based on a hierarchical model and is designed to reduce integration risk, cost, and errors. This is particularly relevant to the proposed framework because the Integration Layer connects robotic and AI systems with WMS and ERP platforms. ISA-95 should not be portrayed as an inadequate or obsolete architecture: its strength is precisely its definition of enterprise-control interfaces and information exchange. The proposed framework instead extends this perspective by considering the AI reasoning, multimodal perception, robotic execution, digital twin, and XR components that operate around these enterprise-control interfaces. The ISA-95 Part 1 standard was substantially updated in 2025 [20], and the present discussion references this current version rather than relying exclusively on the earlier 2010 edition.

2.7. Comparative Analysis of Existing Architectures

Table 1 summarizes this comparison across architectural dimensions relevant to SME commercial automation. The comparison indicates that while RAMI 4.0, IIRA, ISA-95, and digital twin reference architectures [21] each address important aspects of layered, interoperable industrial system design, none of them jointly specifies multimodal perception, LLM-based reasoning, multi-agent orchestration, XR-based training and intervention, and SME-specific resource constraints within a single, phased implementation roadmap. This is the specific gap the proposed framework addresses.
The pattern in Table 1 is not a deficiency of the existing architectures but a reflection of their intended scope: RAMI 4.0, IIRA, and ISA-95 were developed prior to the maturation of large language models and embodied AI, and were designed for large-scale, capital-intensive industrial deployments rather than resource-constrained SME environments. The proposed framework does not claim architectural superiority; rather, it specializes and operationalizes principles already established in these architectures for a context SME commercial automation under variable, high-SKU conditions that none of them was designed to address directly.

3. System Architecture and Conceptual Framework

Consistent with the sociotechnical perspective outlined in Section 2.5, the architecture proposed here separates technical and organizational concerns into distinct but interdependent layers. Translating the goals of the previous section into something that can actually be built in an SME warehouse requires a system architecture that is modular, scalable, and capable of operating both on-premises and, where appropriate, through the cloud. The architecture proposed here is organized into four layers, perception, cognition, execution, and integration, each of which can be developed, tested, and replaced independently. The layering is deliberate: it limits the blast radius of an upgrade in any one layer, and it lets an SME begin with off-the-shelf components in the layers it cares least about while investing more carefully in the layers that touch its core workflows.
Perception Layer: High-definition optical cameras and depth sensors (RGB-D or LiDAR) feed local vision models; YOLO-family detectors for general object localization, and lightweight CNNs trained on the enterprise’s own SKUs, that identify products in real time by shape, color, barcode, or QR code. Running on-device keeps the round-trip from camera to control loop within the tens of milliseconds the gripper needs to act on it.
Cognitive Layer: A locally hosted LLM, or a small multi-agent ensemble, ingests live warehouse state alongside incoming orders and decides what should happen next. It sequences sorting logic, optimizes the physical picking route to limit travel time and energy, and arbitrates between agents when two tasks compete for the same hardware. Route optimization of this kind is an active research problem in its own right: recent work models the warehouse layout itself as a graph and combines graph neural networks with deep reinforcement learning to learn picking-path policies that generalize across warehouse sizes and layouts, which is the kind of learned routing component the cognitive layer is expected to host as it matures beyond hand-coded heuristics [22].
Execution Layer: A modular 6-axis robotic arm carries out the physical work. Because commercial inventories are heterogeneous by definition, the arm must accept interchangeable or adaptive end-effectors, soft pneumatic grippers for fragile electronics, vacuum suction cups for standard cardboard, finger grippers for irregular machinery parts, and the cognitive layer is responsible for choosing the right one before the motion begins.
Integration Layer: Robotic hardware and AI logic interface with the enterprise’s existing Warehouse Management System (WMS) and Enterprise Resource Planning (ERP) software through well-defined APIs. The data flow is bi-directional: the WMS tells the robots what to pick; the robots tell the WMS what they actually picked, in real time, so that inventory does not drift between the floor and the system of record.
These four layers come together in the flowchart in Figure 7, which traces the hierarchical data flow from sensory input in the Perception Layer through decision-making in the Cognitive Layer to execution and, finally, integration with the WMS/ERP systems described above.
Compared with architectures proposed for large-scale industrial automation, the framework described here is deliberately less ambitious in raw capability and more demanding on adaptability. Where industrial reference architectures (e.g., the RAMI 4.0 stack, or the integration patterns) assume stable workflows and dedicated platform teams, the architecture here assumes the opposite: workflows that change with the season, and IT capacity measured in two or three people rather than two or three departments. The trade-off is explicit. Each layer is sized so that an SME can swap a component, replace a vision model, add a second arm, migrate from a local LLM to a hosted one, without rebuilding the whole stack.

4. Methodology for SME Integration

This study follows a qualitative, conceptual methodology aimed at producing a workable framework for integrating AI-driven robotics and immersive technologies into the daily operations of SME commercial enterprises. Rather than measuring an intervention experimentally, the work proceeds by analyzing recent literature, mapping the workflows where commercial SMEs typically lose time and money, and assembling a system architecture and implementation plan that responds to what the analysis reveals. The intention is to give practitioners a defensible starting point and to give researchers something concrete to test.
The literature underpinning this framework (Section 2) was assembled through targeted, purposive searches in Google Scholar and Scopus, focusing primarily on the literature published between 2015 and mid-2026, using search terms organized around the framework’s four thematic clusters: embodied AI and computer vision in commercial robotics, extended reality and human–robot interaction, digital twin and simulation-based training, and enterprise workflow automation (e.g., combinations of “embodied AI,” “warehouse robotics,” “digital twin SME,” “AR-assisted picking,” “multi-agent workflow automation”). Sources were prioritized for inclusion where they addressed unstructured or variable commercial environments, SME-relevant resource constraints, or empirical performance data transferable to warehouse or logistics contexts; sources addressing only large-scale, capital-intensive industrial deployments were retained selectively, where relevant for architectural comparison (Section 2.6), but not treated as representative of the SME context this paper addresses. This synthesis follows a narrative rather than systematic review protocol: it does not report a PRISMA-style screening count, and does not claim exhaustive coverage of the field. This choice is deliberate and stated explicitly, consistent with the paper’s conceptual and Design Science orientation (Section 2.5), where the object of contribution is the resulting framework rather than the literature review itself.
The methodological backbone is a four-phase implementation model, qualitative needs assessment, digital-twin and VR pre-training, controlled hardware pilot, and AR-supported scaling, sequenced so that each phase de-risks the next. Risk-mitigation is treated as a first-class design constraint rather than an afterthought: capital exposure is held back until pilot results justify it, technical assumptions are validated in virtual environments before they touch production, and workforce changes are introduced gradually with structured retraining. The framework is intentionally modular, so that an SME can stop after Phase 2, scale Phase 3 horizontally across multiple sites, or skip ahead to Phase 4 if a prior automation initiative has already cleared the early ground.
This approach is consistent with case-by-case strategies observed in adjacent work on digital transformation in smart, sensor-equipped environments, where IoT, AI, and robotic systems are introduced incrementally and validated within bounded test contexts before any wider rollout [23]. Crucially, the methodology does not assume the resource profile of a large enterprise. It is calibrated to the conditions SMEs actually face: legacy WMS or ERP systems that resist clean integration, small IT teams that cannot maintain a complex on-premise AI stack, warehouses whose layout changes more often than their software documents, and a workforce that has not previously worked alongside robotic systems. Every design choice in the phases that follow reflects this profile.

4.1. Phase 1: Qualitative Needs Assessment and Gap Analysis

The first phase establishes the operational and strategic foundation for the rest of the project. Its purpose is to understand the SME on its own terms, its existing workflows, its constraints, its tolerances for change, before any technology decision is made. Skipping this phase is the most common reason commercial automation projects fail: a system that is technically excellent but mismatched to the actual operating environment will not be used, no matter how impressive the demo looked.
The phase begins with a structured stakeholder engagement exercise. Interviews and facilitated workshops are held with warehouse managers, operations and logistics staff, IT personnel, and finance leadership. The objective is twofold: to surface the day-to-day pain points that the people closest to the work already know about, and to align expectations across functions before commitments are made. Sales and customer-service teams are deliberately included, because fulfilment problems often originate in upstream order capture rather than in the warehouse itself.
A targeted strategic analysis is then carried out, drawing on a SWOT and a lightweight PESTLE assessment scoped to the automation question rather than the business as a whole. The strengths and weaknesses surface internal readiness, quality of existing infrastructure, technical competence on the team, financial headroom, while opportunities and threats locate the project against the wider market context, including competitor moves, regulatory shifts, and labor-market pressure. The output is a clear-eyed view of what the organization can plausibly absorb in the next twelve to eighteen months.
The phase closes with detailed workflow analysis. Every step in the inventory and order-fulfilment lifecycle, receiving, put away, picking, packing, shipping, returns, is mapped end-to-end and time-stamped. Bottlenecks, manual checkpoints, and high-error stages are flagged as candidates for automation; processes that are already efficient, or that depend heavily on human judgment, are explicitly excluded. The deliverable is a prioritized list of automation targets, each with a baseline performance measurement against which later phases can be evaluated. Without this baseline, claimed improvements in Phase 3 cannot be defended.

4.2. Phase 2: Digital Twin Modeling and XR Training

Once the operational baseline is in place, the project moves into a virtual phase before any hardware is purchased. The targeted warehouse zone is digitized into a 3D model, a digital twin, that mirrors aisle geometry, shelving heights, fixed obstacles, and the routes that staff and forklifts actually use. The twin is built to a level of fidelity sufficient for the robotic arm’s reach envelope and path-planning algorithms to be tested against realistic spatial constraints, including the cramped corners and irregular clearances that idealized CAD layouts tend to omit. Path conflicts and reach failures that would have been expensive to discover on the warehouse floor are caught here at the cost of an afternoon’s simulation time. This staged, low-commitment use of the twin matters specifically because SMEs adopt digital twin technology on different terms than large manufacturers: a recent study of maintenance-focused digital twin adoption in small and medium-sized enterprises finds that limited capital, thin in-house modelling expertise, and uncertainty about integration with legacy systems are the binding constraints, rather than any technical ceiling on what the twin itself can do, which is consistent with confining the twin here to a bounded pre-hardware validation step rather than a continuously maintained shadow of the whole operation [24].
The same virtual environment then doubles as a training ground for the workforce. Staff put on VR headsets and rehearse interactions with the simulated robotic systems in scenarios that progress from passive observation through guided collaboration to anomaly handling, a damaged package on the belt, a barcode that will not scan, an emergency stop. The point is not novelty: it is to build muscle memory for the cadence of human–robot collaboration, to normalize the safety protocols, and to surface the resistance and concerns that staff would otherwise raise only after the equipment arrives. Trainees who have already worked through a dozen virtual shifts approach the live system with a measurable reduction in hesitation, and the project gains a feedback channel that informs the next phase.

4.3. Phase 3: Controlled Hardware Pilot Deployment

Hardware enters the picture in Phase 3, but on deliberately constrained terms. A localized, low-risk pilot is initiated to validate the solution without disrupting core business functions: a single 6-axis robotic arm is installed at a fixed sorting station, fed by products tagged with QR codes routed from a portion of the regular inbound flow. The station operates in parallel with existing manual processes rather than replacing them, which means the business continues to function on day one even if the pilot does not, a precondition for getting executive sign-off in a resource-constrained SME.
Vision calibration is the most underestimated activity in the phase. The computer vision algorithms must be tuned under the warehouse’s actual lighting conditions, which typically vary by time of day, season, and which lights happen to be working. Thresholds for blob detection and color recognition are optimized iteratively against a representative SKU sample, and edge cases, reflective shrink-wrap, transparent clamshells, partially occluded labels, are deliberately stress-tested rather than avoided. A vision system that performs at 99% on a curated test set and at 78% on the actual inventory will not survive contact with operations; the pilot is the right place to find that out.
Throughout the pilot the system is evaluated against an explicit set of KPIs rather than against impressions. Cycle time captures speed per pick; pick-up success rate captures grip reliability across SKU variation; detection accuracy captures vision precision; and system uptime captures whether the platform stays available across shifts and not just during demos. Each KPI is measured against the baseline established in Phase 1, so claimed improvements are defensible rather than rhetorical. A go/no-go decision for Phase 4 is gated on this data.

4.4. Phase 4: Scaling and AR Integration

Phase 4 begins only after the Phase 3 KPI gate has been cleared, and it scales the system on two axes at once: more of the workflow, and tighter integration with the people who work alongside it. Additional sorting stations are brought online progressively, and the perception and cognitive stack is hardened against a broader range of inventory than the pilot encountered. Each scale step is itself a small pilot, capacity is not doubled in a single weekend, so that an unexpected failure mode never threatens more than one station’s throughput.
Live human support is introduced through Augmented Reality. Workers are equipped with AR tools, typically tablets and smart glasses depending on task ergonomics, and the cognitive layer routes anomalies, an unreadable barcode, a damaged package, a SKU that the model has never seen, directly to the nearest human via the AR overlay. Instructions are pinned to the physical object in question rather than buried in a separate screen, which shortens the time from anomaly to resolution and keeps the human in a supervisory rather than reactive posture. The same channel carries on-the-job microlearning for newer staff, building expertise without removing people from the line for formal training.
Closing the loop, every failed pick, every misidentification, and every human correction is fed back into the localized AI model. Over weeks and months, the system becomes progressively more accurate against the specific inventory and the specific environment it operates in, a form of continuous learning that does not depend on sharing operational data with an external cloud service. The retraining pipeline is itself kept lightweight, both to fit the resource profile of the SME and to keep model behavior traceable when something does go wrong.
This four-phase roadmap, from the initial Needs Assessment, through Digital-Twin/VR pre-training and the Hardware Pilot, to the AR-supported Scaling phase described here, is set out in Figure 8.

4.5. Technical Implementation of the Perception Layer: Real-Time Vision Algorithms

The Perception Layer described above is more than a conceptual block: it is the place where the framework meets the warehouse floor. The code listing below (Listing 1) is the working implementation that runs on the on-arm vision module during the Phase 3 pilot. A continuous capture loop pulls snapshots from the camera and routes them through local machine vision routines that identify products by color and shape; nothing leaves the device. The find_blobs primitive performs the heavy lifting of object segmentation, while the draw_cross and draw_rectangle calls visualize detected coordinates back onto the live image, both for debugging and, more importantly, for operator transparency. The output is a set of actionable spatial coordinates that the cognitive and execution layers consume to plan and carry out a pick. The point of including the code is to show that the framework can be specified concretely, not just in the abstract: a small SME team can read it, modify it for their own SKU palette, and run it on commodity hardware.
Four behaviors of the system are worth pointing out specifically. First, real-time perception: the loop runs continuously, so incoming products are detected as they arrive rather than on a scheduled scan. Second, product identification by color code (here, ‘R’ for red and ‘B’ for blue, extended in deployment to a richer category palette), which keeps adaptive sorting decisions local. Third, coordinate mapping: the crosses and rectangles drawn at the center of detected blobs provide the spatial input the 6-axis arm needs for path planning. Fourth, human-in-the-loop transparency: writing the processed image to an LCD lets the operator see what the robot is seeing, which is the single most effective safeguard against the silent misclassification failures that erode trust in autonomous systems.
Listing 1. Pseudocode for the color-based object detection and tracking algorithm.
1. while True: img = sensor.snapshot ()
2. img.draw_cross(bx, by, color=(0, 0, 0), size=10, thickness=1)
3. img-draw_rectangle(bx -50, by -50, 100, 100, color=(0, 0, 0), thickness=1)
4. for blob in img.find_blobs (thresholds, roi=color_roi, x_stride=15, y_stride=15, pixels_threshold 150):
5. color= ‘’
6. code = blob.code()
7. if code == 1:
8. color = ‘R’
9. img.draw_string(blob.x(), blob.y() + 10, ‘R’)
10. elif code == 2:
11. color = ‘B’
12. img.draw_string(blob.x(), blob.y() + 10, ‘B’)
13. img.draw_cross(blob.cx(), blob.cy())
14. img.draw_rectangle(blob.rect())
15. lcd.write(img, roi=lcd_roi)
The methodological contribution of this research is the integration itself. Most existing work treats organizational analysis, digital-twin simulation, embodied AI, XR-supported training, and phased deployment as separate disciplines, with their own studies, conferences, and assumptions about who the user is. Bringing them into a single strategy that an SME can actually execute requires more than listing them in a paragraph, it requires deciding the order in which they appear, the de-risking dependencies between them, and the points at which each one earns or loses its place in the project. That is the work the four-phase model is doing.
Equally important is what the framework chooses not to optimize for. It does not pursue maximum theoretical throughput; it does not assume cloud connectivity; it does not treat the workforce as a residual category. Interoperability with existing legacy systems, gradual capital exposure, structured workforce adaptation, and explicit risk reduction at each phase are treated as design constraints rather than nice-to-haves. The result is a roadmap shaped for the SMEs who would otherwise be priced or complexity-priced out of the conversation, and one that researchers can pick up and stress-test in real settings.

4.6. Illustrative Case Study: Applying the Framework to a Hypothetical SME

To make the four-phase implementation model concrete, this section walks through a hypothetical application of the proposed framework. The scenario is illustrative rather than empirical: it does not report data from an actual deployment, and no claims of validated performance are made. Its purpose is to demonstrate that the framework’s layers and phases translate into specific, actionable steps for a realistic SME context, rather than remaining at a purely abstract level.
Scenario. Consider a hypothetical SME operating as an importer and distributor of consumer electronics, managing approximately 600 SKUs ranging from small accessories to boxed appliances, with three warehouse staff and a single manually operated forklift. The company faces the pressures described in Section 1: inconsistent SKU packaging, seasonal order spikes, and customer expectations for next-day fulfilment, but operates with a warehouse management system (WMS) that lacks any automation beyond barcode scanning.
Phase 1—Needs Assessment and Gap Analysis. A qualitative assessment identifies that the greatest time loss occurs during picking of small, visually similar SKUs (e.g., cable accessories) and during peak-season put-away, where staff spend a disproportionate share of shift time locating shelf space. This maps directly onto the Perception and Cognition layers: the enterprise needs reliable object identification for visually similar items and route optimization for put-away, rather than, for example, heavy investment in the Execution layer (robotic arms), which would address a comparatively smaller share of lost time at this stage.
Phase 2—Digital Twin Modeling and VR Pre-Training. Before any hardware investment, a digital twin of the existing warehouse layout is constructed from the WMS’s shelf-location data, and staff use VR walkthroughs to pre-train on a proposed reorganization of high-turnover SKUs. This phase allows the SME to test whether a reorganized layout reduces travel distance in simulation before committing floor space or budget—directly addressing the need for adoption costs and organizational readiness to be considered before hardware deployment.
Phase 3—Controlled Hardware Pilot. A single fixed camera with a lightweight CNN, trained on a small sample of the enterprise’s own SKU images, is piloted at one picking station to assist staff in distinguishing visually similar accessories, without any robotic manipulation. The Cognition layer here is deliberately minimal: a rule-based decision layer flags likely SKU matches for human confirmation rather than acting autonomously, consistent with the human-in-the-loop principle established in Section 2.5. KPI gates from Section 3 (e.g., minimum detection accuracy before any expansion) determine whether the pilot proceeds to Phase 4.
Phase 4—Scaling and AR-Supported Operation. If Phase 3 KPI thresholds are met, AR-assisted picking (e.g., a handheld or wearable display highlighting bin locations) is introduced at additional stations, and the Integration layer connects the vision system’s output to the existing WMS so that picking confirmations update inventory in real time. Robotic execution (the Execution layer) is treated as a later-stage option contingent on sustained volume growth, rather than a default component—reflecting the framework’s resource-conscious sequencing for SMEs, in contrast to reference architectures such as RAMI 4.0 or ISA-95, which do not prescribe this kind of staged, cost-gated progression (Section 2.6).
Interpretation. This scenario illustrates two features of the framework that are difficult to convey in the abstract. First, not every layer needs to be activated at once: a small SME can meaningfully adopt the Perception and Cognition layers alone, well before Execution-layer investment becomes justified. Second, the phased model functions as a risk-limiting sequence rather than a fixed deployment pattern, allowing an SME to exit or pause after any phase if KPI thresholds are not met. The scenario does not substitute for empirical validation; it is offered to demonstrate operational plausibility, and the need for empirical testing with a real SME partner is addressed explicitly in Section 5.2.

4.7. Evaluation Criteria

Because the framework is conceptual, the question of how to evaluate it matters as much as the question of how to specify it. Evaluation in this research is therefore framed around two layers, distinguished by what they can plausibly tell us. The first layer is internal validation: whether the framework hangs together, whether its phases are well-sequenced, its layers are well-bounded, and its assumptions about SME constraints are consistent with the literature on which it is built. Internal validation is what the present research offers, and it is judged by the coherence of the argument and the fit between the proposed architecture and the operational realities reported across the sources reviewed.
The second layer is empirical validation, which the research deliberately does not claim to provide. When an SME adopts the framework, the relevant evaluation criteria are operational and measurable: cycle time per pick against the Phase 1 baseline, pick-up success rate across a representative SKU mix, vision detection accuracy under realistic lighting variance, system uptime across full shifts, mean time between human interventions, and the time required to onboard a new staff member from first VR session to live floor work. Complementing these are economic criteria, payback period, the ratio of cumulative savings to cumulative investment, and the marginal cost of adding a station in Phase 4, and human criteria including reported workforce confidence, voluntary turnover during the transition, and the rate of accepted system suggestions in AR-mediated tasks.
These criteria are not aspirational; they are the same metrics on which Phase 3 already gates its go/no-go decision, extended over the longer time horizons that distinguish a successful deployment from a successful pilot. Future empirical work is therefore not free to pick its own success measures; the framework specifies them in advance, which is itself a contribution.

4.8. Threats to Validity

Any conceptual framework should be honest about the conditions under which its conclusions would not hold. The most significant threat is selection bias in the literature on which the framework is built: published reports of AI-driven robotics in commercial settings skew towards successful or at least funded deployments, while abandoned projects and silently rolled-back pilots tend not to surface in peer-reviewed venues. The framework’s design choices are robust against this where possible, the insistence on baseline measurement, KPI-gated phases, and parallel rather than replacement operation in Phase 3, but the bias cannot be eliminated.
A second threat concerns external validity across SME sub-segments. The research scopes itself to commercial trading, import/export, logistics, and technology distribution; the framework is not claimed to transfer cleanly to cold-chain logistics, pharmaceutical distribution, or sectors with heavy regulatory inspection regimes, where the cognitive and integration layers would face additional constraints not modelled here. Adaptation, not adoption, is required for those settings.
A third threat is the assumption of a baseline level of digital readiness. The framework presumes the SME has, or is willing to acquire, a working WMS or at least a structured inventory database, a network capable of supporting on-premise compute, and a leadership team that will protect the project across the eight-to-twelve months a full four-phase rollout typically takes. SMEs that lack any of these, and they are not rare, should expect Phase 1 to extend in scope before Phase 2 becomes feasible.
Finally, the framework’s emphasis on locally processed AI and human-centered deployment makes it relatively robust to the failure modes that have plagued cloud-centric automation projects (latency, vendor lock-in, opaque model behavior). It is correspondingly more exposed to threats specific to on-premise systems: hardware obsolescence, the recruitment difficulty of small but specialized in-house teams, and the long-term drift of locally retrained models in the absence of external validation. Naming these risks is part of the framework’s honesty about its own boundaries.

5. Discussion and Ethical Considerations

5.1. Operational Transformation and Enterprise Efficiency

If the proposed framework does what it claims to do, the most visible change for an SME is not a single dramatic moment of automation but a steady, observable shift in how the warehouse runs day to day. Embodied AI handles product recognition, sorting, and packing; computer vision turns the shelf into a continuously readable surface; and XR-supported workflows give human staff a clearer picture of what the system is doing and why. The result is a measurable improvement in inventory accuracy, a reduction in repetitive manual labor, and a smoother order-fulfilment process. Just as importantly, it frees staff to spend more of their time on quality assurance, exception handling, customer relations, and the kind of judgement calls that intelligent systems still struggle with.
The literature is broadly consistent on this point: where robotic automation is applied to repetitive warehouse tasks, consistency improves and human error drops. The size of the effect, however, is not constant. It depends heavily on environmental conditions, on the maturity of the technology stack, and on whether the organization is actually ready to receive it. Studies on large enterprises tend to report dramatic gains because the underlying environments, uniform pallets, stable lighting, fixed conveyors, predictable SKUs, were essentially designed for the robots. SME warehouses are not. They carry diverse, high-SKU inventories, the floor plan changes more often than the WMS knows about, packaging is irregular, and inventory movement patterns shift with the season. The same robot performs very differently in those two settings.
The phased implementation model takes this seriously. Rather than committing capital up front for a full-facility rollout, an SME using this framework calibrates the AI models, refines the workflow, and confirms operational feasibility on a deliberately small footprint before expanding. The financial logic follows the same shape: a 6-axis arm plus sensor stack is a non-trivial capital outlay, but in the deployments most comparable to the SME case the return on investment is typically realized within 18 to 24 months, driven by fewer mis-picks, denser use of floor space, and the kind of consistent uptime that manual operations rarely achieve on their own.
Digital twins and XR-based simulation environments support this calibration by letting enterprises test robotic workflows virtually before anything physical moves. The cost of a layout error caught in simulation is a few hours of an engineer’s time; the cost of catching it after the racks are bolted to the floor is considerably higher. In that sense, intelligent automation succeeds or fails less on the sophistication of any individual technology than on the discipline of how the technologies are sequenced, and on whether the organization around them is flexible enough to adapt as the system learns.

5.2. Technological Limitations and Systemic Challenges

There is a familiar gap between what AI-driven automation can do in a paper and what it can do in a warehouse, and it shows up almost immediately when these systems are deployed in commercial environments. The existing literature tends to highlight the capabilities and downplay the instability of the conditions those capabilities have to work in. In practice, computer vision systems remain vulnerable to a fairly mundane list of real-world factors: inconsistent and seasonal lighting, reflective or transparent packaging that confuses depth cues, partial occlusions when stock is densely shelved, narrow aisles that limit camera angles and force tight kinematic envelopes, and inventory configurations that change faster than any retraining cycle can absorb. Each of these on its own is manageable; together they compound, and a robot that performs cleanly under test conditions can degrade visibly within a week of live use.
Their immediate effect is on the three things that matter most operationally: object detection accuracy, grasping reliability, and navigation stability. Fully autonomous operation in this kind of environment is therefore difficult to sustain without ongoing human supervision and periodic recalibration of the underlying models. This is not a flaw of the proposed framework; it is a feature of the world the framework is designed for, and the framework’s job is to make that supervision and recalibration efficient rather than to pretend it can be eliminated.
A second class of challenge is interoperability. SMEs do not arrive at automation from a blank slate. They typically already run a Warehouse Management System that is partly outdated, a set of databases that do not talk cleanly to each other, and inventory processes that are still partly manual by design rather than by accident. Inserting AI-driven robotics into this stack usually means more than buying hardware: it implies API development, data standardization, careful sequencing of legacy-system retirement, and an honest accounting for software integration, workforce retraining, infrastructure adaptation, and long-term maintenance. The cost line that vendors quote rarely includes any of this; the cost line that finance signs off on must.
A third concern is the data dependency of the AI models themselves. Machine-learning systems need continuous, high-quality operational feedback to keep their accuracy stable. Where that feedback is thin, and for many SMEs it will be, because the historical datasets simply do not exist at scale, performance drifts, misclassification rates creep up, and the system becomes progressively less adaptable to new inventory categories. The framework’s continuous-learning loops and AR-assisted human corrections are direct responses to this, but they require a deliberate organizational commitment to label, review, and feedback what the system gets wrong.
Finally, cybersecurity and data governance deserve more attention than they typically receive in this literature. Once robotic systems, cloud services, enterprise databases, and head-mounted XR devices are connected together, the attack surface expands quickly. Comparable concerns have been raised in the literature on AI-driven public-administration workflows, where data-quality assurance, encryption, and explicit privacy governance are treated as preconditions for deployment rather than optional add-ons [11]. Wireless sensor networks and other IIoT links compound this exposure at the perception layer specifically, and recent work on securing WSN-driven smart manufacturing under Industry 5.0 argues that layered authentication, encrypted sensor-to-controller channels, and anomaly-detection models tuned to plant-floor traffic need to be designed in from the start rather than bolted onto a working pilot after the fact [25]. Operational data, inventory records, and pricing-sensitive workflows become exposed if the network and identity layers are not designed with the same care as the perception layer. The privileging of localized, on-premise processing in this framework is partly motivated by exactly this concern: keeping the model and the data inside the building reduces the surface that needs to be defended and shortens the path of any breach.
Taken together, these limitations argue not against intelligent automation but against the assumption that it can be deployed without organizational adaptation. The framework attempts to embed that adaptation through localized AI processing, phased pilot deployment, structured human-assisted intervention, and continuous-learning mechanisms, and through an explicit acknowledgement that operational uncertainty is a permanent feature of the environment, not a transitional one.

5.3. Human–Robot Collaboration and Workforce Transition

Few questions in this literature attract more disagreement than the relationship between intelligent systems and the human workforce. One body of work treats AI-driven automation as fundamentally a labor-reduction mechanism; another treats it as augmentation and upskilling. This study takes the second view as a design principle, in line with broader observations that AI is most usefully understood as a reordering of work between humans and machines, taking over the repetitive and the mechanical so that human capacity is redirected toward judgment, innovation, and exception handling [11]. The argument is partly ethical and partly practical: sustainable automation, in environments as variable as commercial SME warehouses, depends on the human operators staying in the loop, because the system will keep encountering cases the model was not trained for.
Within the framework, human workers are not residual components of an otherwise automated process. They are the supervisory and decision-making layer of the operational ecosystem. The robotic systems handle what is physically repetitive and computationally heavy; the people handle exception cases, judgement calls, quality control, and the small adjustments that keep the rest of the pipeline aligned with what the business actually needs. This division of labor is not philosophically tidy, there will always be a grey zone between what the robot decides and what the operator overrides, but it is operationally honest.
VR and AR matter to this division for very practical reasons. VR-based training environments give employees a safe place to develop fluency with the robotic systems before the systems go live, which lowers technological anxiety and softens the change-management curve that automation projects so frequently founder on. AR-assisted operational guidance then keeps that fluency current during normal operations by surfacing anomalies, instructions, and contextual data exactly where they need to be acted on, in the operator’s field of view, beside the object in question, at the moment the decision matters. Trust between the operator and the robot is not a fixed quantity that AR training simply installs once; recent work modelling operators’ visual attention during human–robot collaboration finds that where a person looks in real time tracks their moment-to-moment trust in the robot’s next action, which suggests that AR interfaces could eventually be tuned to the operator’s live attentional state rather than delivering the same fixed prompt regardless of how much the operator already trusts the system in that instant [26].
That said, the transition is not simply a technical exercise. Workforce change of this kind carries real social weight. Employees are reasonable to worry about job security, about the rate at which their skill requirements are shifting, and about whether the organization will invest in helping them keep up. Without transparent communication and a structured reskilling plan, automation initiatives generate resistance well before they generate efficiency, and the resistance is usually the rate-limiting step.
There is also a quieter ethical concern that runs through the literature without quite being addressed: the increase in workplace surveillance that comes naturally with AI-powered monitoring and wearable XR devices. The same data that lets a system improve its model also lets it track individual workers in detail that would have been inconceivable a decade ago. Enterprises deploying intelligent automation need explicit governance frameworks, what is collected, who can see it, what it may and may not be used for, that balance operational optimization against employee privacy and organizational trust. Treating this as an afterthought is, in practice, a way of guaranteeing that it becomes an incident.
What the discussion converges on is that successful automation is not only a technological accomplishment. It requires ethical workforce management, inclusive organizational planning, and a sustained commitment to professional development. The framework’s role is to make those commitments easier to honor, not to substitute for them.

5.4. Strategic and Economic Implications for SMEs

Seen from the finance side, intelligent automation looks like both a substantial opportunity and a substantial barrier, depending on which row of the spreadsheet is in focus. The upside is the part most often cited: better operational scalability, lower long-term labor costs, faster supply-chain responsiveness, and the kind of real-time inventory visibility that gives an SME genuine competitive ground in increasingly digital markets. The downside is the part most often understated: the implementation costs are front-loaded, the technology cycle is short, and the internal expertise needed to keep the system productive is non-trivial.
SMEs face structural constraints that the literature on large-enterprise automation tends to skim past. Recent SME-focused reviews of AI and immersive technology in commerce report that, even when the technical case is clear, adoption is slowed by perceived complexity, limited internal capacity, and uncertainty about the actual return, concerns that scale poorly when an organization does not have a dedicated digital-transformation function [1]. A broader systematic review of AI adoption across SMEs generally, not limited to commerce or logistics, reaches the same conclusion from a different angle: financial limitations, a shortage of in-house AI skills, and inadequate digital infrastructure are the barriers that recur across sectors and regions, while top-management support and staff capability are the factors most consistently associated with adoption succeeding once it is attempted [27]. Initial capital outlay is heavier in relative terms because the fixed costs of integration do not scale down linearly. ROI horizons are uncertain because the operational variability is higher. Technical expertise is harder to retain because the labor market for it is competitive. And the technology itself moves quickly enough that a stack chosen today may need to be partially rebuilt within a few years. Models drawn from multinational case studies, with their dedicated automation teams, multi-million-dollar pilot budgets, and tolerant capital structures, simply do not translate.
The phased implementation methodology proposed here is a direct response to that mismatch. It promotes incremental deployment so capital exposure tracks demonstrated value rather than projected value; modular infrastructure design so individual components can be replaced without redoing the whole stack; localized pilot testing so unsuccessful approaches fail cheaply; and scalable integration so a successful pilot can grow horizontally rather than requiring a separate rollout effort. Taken together, the structure lets an SME evaluate operational effectiveness in stages, with most of the financial exposure deferred to the point where the operational case has already been made.
The strategic contribution is the framing as much as the mechanics. Treating intelligent automation as a long-term organizational transformation, rather than as a one-off hardware purchase, changes how the business plans for it, budgets for it, and absorbs the resulting changes. Sustainable digital transformation rests on continuous adaptation, workforce integration, and organizational learning. The hardware is necessary but not, on its own, sufficient.

5.5. Theoretical and Practical Implications

On the theoretical side, this study contributes to the Industry 4.0 literature by integrating embodied AI, XR technologies, enterprise workflow automation, and human–robot collaboration into a single conceptual framework rather than treating them as four loosely related research areas. This distinguishes the framework from existing reference architectures such as RAMI 4.0, IIRA, and ISA-95, which were developed for large-scale, capital-intensive industrial deployments and do not specify how LLM-based reasoning, multimodal perception, or SME-specific resource constraints should be organized (Section 2.6); the novelty claimed here lies specifically in operationalizing these elements for SME-scale commercial automation, rather than in the four component technologies themselves, each of which already has an established literature of its own. The interdependencies between them are where most of the difficulty actually lives, not within any one of them on its own, and the framework attempts to make that visible. It also offers a vocabulary in which the cognitive layer, the perception layer, and the workforce that sits alongside both can be discussed in the same terms.
On the practical side, the framework gives SMEs a concrete starting point. It is not a finished product, and it is not a substitute for the specific judgement that any individual deployment will require, but it is a structured roadmap, phased, modular, calibrated to resource-constrained settings, and built around human-centered automation. For organizations whose previous exposure to AI has been limited to procurement decks and vendor demos, a roadmap that names its phases, identifies its risks, and accepts its own limits is more useful than a more ambitious framework that pretends to none of those things.
The framework’s value will ultimately be measured by what happens to it once it is taken into a live setting. As a conceptual model, it stops short of empirical validation. That validation, through case studies, simulation experiments, and real-world pilot implementations, is the next step, and the design of those studies is in itself a research question worth taking seriously.

5.6. Future Work and Empirical Validation

Three lines of empirical work would do most to test the framework. The first is a longitudinal case study in a single SME warehouse, instrumented to capture the four headline KPIs defined in §4.3, cycle time, pick-up success rate, detection accuracy, and uptime, across the full Phase 1 to Phase 4 sequence. The point of running it longitudinally is that the most interesting effects, especially the ones tied to workforce adaptation, take months rather than weeks to surface. A natural extension of this same study is cross-country replication, since regulatory environments, labor-market conditions, and SME digital-readiness baselines vary substantially across regions and would affect how directly these findings generalize beyond the initial deployment context.
The second is a comparative simulation study using digital twins of contrasting warehouse archetypes: a high-SKU electronics distributor, a mixed consumer-goods importer, and a B2B industrial-component reseller. Holding the framework constant and varying the environment would help separate the effects that are intrinsic to the approach from those that are an artefact of any one operational context. A simulation-first design also lowers the cost of running the comparison at the scale needed for it to mean something statistically.
The third is a controlled study of the workforce-transition mechanisms themselves, the VR pre-training programme, the AR-assisted live operations, and the changes to role design that the framework recommends. Self-reported anxiety, retention through the transition, time-to-competence on the new tasks, and productivity at the new tasks are all measurable, and the existing literature gives almost no guidance on what effect sizes to expect.
Beyond the empirical agenda, three technical extensions are worth pursuing. First, the cognitive layer would benefit from a deeper investigation of small, locally deployable language models, particularly in the 7B-parameter range that recent open-weights releases have made viable, as orchestration components for warehouse-floor decision-making. Second, the perception layer should be revisited as commodity depth sensors and event-based cameras mature; both have the potential to address the lighting and occlusion failure modes that currently constrain Phase 3. Third, sustainability metrics, energy consumption per pick, embodied carbon of the hardware over its useful life, end-of-life recyclability, belong in the evaluation set, and were left out of this version of the framework only because the literature does not yet provide consistent baselines for SME contexts.

6. Conclusions

This research has set out a conceptual framework for automating the commercial workflows of small and medium-sized enterprises through the convergence of AI-driven robotics, computer vision, and immersive technologies. The starting observation is straightforward: SMEs face the same competitive pressure as their larger rivals, labor shortages, complex inventories, narrow margins, demanding customers, without commanding the capital, the structured environments, or the in-house technical depth that make industrial automation straightforward. The framework offered here is intended to make automation a realistic option for those enterprises rather than an aspirational one.
Three threads tie the work together. The first is the integration of technologies that the literature usually keeps apart: embodied AI and on-device computer vision for perception, local LLM and multi-agent systems for cognition, six-axis robotic arms for execution, and digital-twin, VR, and AR layers for training, validation, and live operational support. The second is a four-phase implementation model, qualitative needs assessment, digital-twin and VR pre-training, controlled hardware pilot, AR-supported scaling, that treats risk mitigation as a design constraint rather than an afterthought, and that gives an SME explicit stopping points and exit ramps along the way. The third is a deliberate orientation toward human augmentation rather than labor replacement: AR overlays elevate operators into supervisors of intelligent systems, VR pre-training lowers the social cost of organizational change, and exception handling, ethical judgment, and quality control remain firmly in human hands.
The framework is conceptual, and the research does not claim to have proven it in the field. What it offers instead is a defensible starting point. The system architecture and phased model give practitioners a way to begin without betting the business on the first deployment, and they give researchers a structured artefact against which empirical work can be designed. The accompanying evaluation criteria and threats-to-validity discussion are intended to make that follow-up easier: future studies have clear KPIs to instrument and clear assumptions to interrogate.
Taken together, the contribution is both academic and practical. Academically, the research consolidates a fragmented literature around a single SME-scoped object of study and articulates the gaps that remain. Practically, it gives SMEs in logistics, trading, and technology distribution a roadmap that respects the operational realities they actually live with, fragmented WMS or ERP systems, small IT teams, warehouses whose layout changes more often than their documentation. Intelligent automation, on these terms, becomes a long-running organizational transformation rather than a one-off hardware purchase. The next step, and the natural continuation of this work, is empirical validation in real SME settings, and the framework is built to receive it.

Author Contributions

S.S. contributed to the conceptualization, methodology, software, validation and formal analysis. E.Z. contributed to the investigation, resources, data curation and writing—original draft preparation. L.B. contributed to the visualization, supervision and project administration. All authors have read and agreed to the published version of the manuscript.

Funding

This publication has been made possible with the financial support of NASRI. Authors are solely responsible for its content. Views expressed do not necessarily reflect those of NASRI.

Data Availability Statement

All data and information analyzed are available from the sources cited in the manuscript.

Acknowledgments

During the preparation of this work, the authors used Claude, Sonnet 5 (Anthropic) to improve the clarity, grammar, and style of the text, as well as to support the design and visualisation of conceptual figures. After using this tool, the authors reviewed and edited the content as needed and take full responsibility for the content of the published article.

Conflicts of Interest

The authors declare no conflict of interest.

References

  1. Shurdhi, S.; Bekteshi, L.; Rexha, G. AI and immersive technologies in e-commerce: A literature review. WSEAS Trans. Syst. 2026, 25, 243–266. [Google Scholar] [CrossRef] [Scilit]
  2. Garg, N. Agentic AI for enterprise workflow automation: Impact and architectural principles for multi-agent orchestration. Int. J. Res. Commer. Manag. Stud. 2026, 8, 122–136. [Google Scholar] [CrossRef] [Scilit]
  3. Chechkin, A.; Pleshakova, E.; Gataullin, S. A hybrid neural network transformer for detecting and classifying destructive content in digital space. Algorithms 2025, 18, 735. [Google Scholar] [CrossRef] [Scilit]
  4. Tu, X.; Autiosalo, J.; Ala-Laurinaho, R.; Yang, C.; Salminen, P.; Tammi, K. TwinXR: Method for using digital twin descriptions in industrial eXtended reality applications. Front. Virtual Real. 2023, 4, 1019080. [Google Scholar] [CrossRef] [Scilit]
  5. Dolezel, P.; Stursa, D.; Kopecky, D. Memory efficient deep learning-based grasping point detection of nontrivial objects for robotic bin picking. J. Intell. Robot. Syst. 2024, 110, 110. [Google Scholar] [CrossRef] [Scilit]
  6. Yunzhu, Y. Research on the Application of Computer and Information Technology in Brand Management. J. Phys. Conf. Ser. 2020, 1648, 042006. [Google Scholar] [CrossRef] [Scilit]
  7. Mobile robots meet augmented reality technologies: Transforming human-robot interaction in Industry 4.0 scenarios. In Proceedings of the 2024 ACM/IEEE International Conference on Human-Robot Interaction (HRI ’24), Boulder, CO, USA, 11–15 March 2024. [CrossRef] [Scilit]
  8. Patro, S. Emerging technologies for precision supply chain management. Int. J. Intell. Syst. Appl. Eng. 2024, 12, 564–574. [Google Scholar]
  9. Kumar, A.; Simangunsong, S.; Carreno-Medrano, P.; Cosgun, A. Mixed reality outperforms virtual reality for remote error resolution in pick-and-place tasks. In Proceedings of the 2025 ACM/IEEE International Conference on Human-Robot Interaction (HRI ’25), Melbourne, Australia, 4–6 March 2025. [Google Scholar] [CrossRef] [Scilit]
  10. Shabbir, S.; Banerjee, J.; Singh, D. Application of virtual reality for process based training in warehouse logistics to evaluate effectiveness. In Proceedings of 2025 International Conference on Intelligent Computing and Virtual and Augmented Reality Simulations (ICVARS), Birmingham, UK, 25–27 July 2025. [Google Scholar] [CrossRef] [Scilit]
  11. Shurdhi, S.; Bekteshi, L. Artificial intelligence in the process of European integration. Rom. J. Econ. 2025, 60, 132–142. Available online: https://revecon.ro/sites/default/files/2025-1-10.pdf (accessed on 23 June 2026).
  12. Motahari-Nezhad, H.R.; Shwartz, L. Towards open smart services platform. In Proceedings of the 50th Hawaii International Conference on System Sciences (HICSS), Hawaii, HI, USA, 4–7 January 2017. [Google Scholar] [CrossRef] [Scilit]
  13. Jeong, C. Beyond Text: Implementing Multimodal Large Language Model-Powered Multi-Agent Systems Using a No-Code Platform. J. Intell. Inf. Syst. 2025, 31, 191–231. [Google Scholar] [CrossRef] [Scilit]
  14. Kumar, P. Agentic AI-driven enterprise architecture: A foundational framework for scalable, secure, and resilient systems. Int. J. Comput. Exp. Sci. Eng. 2025, 11, 8262–8278. [Google Scholar] [CrossRef] [Scilit]
  15. Kakkar, S. Prompt-robotics: Leveraging large language model (LLM) agents to enhance robotic process automation (RPA) supervision and efficiency. Int. J. Financ. Data Sci. 2025, 3, 7–29. [Google Scholar] [CrossRef] [Scilit]
  16. Sony, M.; Naik, S. Industry 4.0 integration with socio-technical systems theory: A systematic review and proposed theoretical model. Technol. Soc. 2020, 61, 101248. [Google Scholar] [CrossRef] [Scilit]
  17. Freire, I.T.; Guerrero-Rosado, O.; Amil, A.F.; Verschure, P.F.M.J. Socially adaptive cognitive architecture for human-robot collaboration in industrial settings. Front. Robot. AI 2024, 11, 1248646. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  18. ZVEI (German Electrical and Electronic Manufacturers’ Association). Reference Architecture Model Industrie 4.0 (RAMI4.0); ZVEI/Plattform Industrie 4.0; ZVEI (German Electrical and Electronic Manufacturers’ Association): Frankfurt am Main, Germany, 2015; Available online: https://www.zvei.org/en/press-media/publications/the-reference-architectural-model-industrie-40-rami-40 (accessed on 23 June 2026).
  19. Industrial Internet Consortium (IIC). The Industrial Internet of Things, Volume G1: Reference Architecture (Version 1.9). 2019. Available online: https://www.scribd.com/document/472610245/IIRA-v1-9 (accessed on 23 June 2026).
  20. International Society of Automation. ANSI/ISA-95.00.01-2025 (IEC 62264-1 Mod): Enterprise-control system integration-Part 1: Models and terminology. ISA. 2025. Available online: https://www.isa.org/products/ansi-isa-95-00-01-2025-iec-62264-1-mod-enterprise (accessed on 23 June 2026).
  21. Tao, F.; Zhang, H.; Liu, A.; Nee, A.Y.C. Digital twin in industry: State-of-the-art. IEEE Trans. Ind. Inform. 2019, 15, 2405–2415. [Google Scholar] [CrossRef] [Scilit]
  22. Čelik, N.; Škraba, A. Graph neural networks and deep reinforcement learning for warehouse order picking and representation learning. Systems 2026, 14, 659. [Google Scholar] [CrossRef] [Scilit]
  23. Shurdhi, S.; Shtylla, S. Digitalization and smart lab concept. In Greening Our Cities: Sustainable Urbanism for a Greener Future; Springer: Berlin/Heidelberg, Germany, 2024; pp. 373–380. [Google Scholar] [CrossRef] [Scilit]
  24. Nasirinejad, M.; Afshari, H.; Sampalli, S. The adoption of digital twin technologies for maintenance in small and medium-sized enterprises: Challenges and benefits. Adv. Eng. Inform. Adv. Eng. Inform. 2026, 74, 104645. [Google Scholar] [CrossRef] [Scilit]
  25. Babbar, H.; Rani, S.; Boulila, W. Fortifying the connection: Cybersecurity tactics for WSN-driven smart manufacturing in the era of Industry 5.0. IEEE Open J. Commun. Soc. 2025, 6, 3417–3428. [Google Scholar] [CrossRef] [Scilit]
  26. Goubard, C.; Demiris, Y. Cognitive modelling of visual attention captures trust dynamics in human–robot collaboration. ACM Trans. Hum.-Robot Interact. 2025, 14, 67. [Google Scholar] [CrossRef] [Scilit]
  27. Hamid, E.S.; Artha, B. Artificial intelligence adoption, implementation barriers, and business impact in small and medium-sized enterprises: A systematic literature review. Arch. Bus. Res. 2026, 14, 86–110. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Overall architecture of the TwinXR method setup, with three layers from Smart Factory, DT document, to XR application. (Source: [4]).
Figure 1. Overall architecture of the TwinXR method setup, with three layers from Smart Factory, DT document, to XR application. (Source: [4]).
Computers 15 00568 g001
Figure 2. Screenshot of the implemented MR application “HoloRobot” (Unity developer view on desktop). (Source: [4]).
Figure 2. Screenshot of the implemented MR application “HoloRobot” (Unity developer view on desktop). (Source: [4]).
Computers 15 00568 g002
Figure 3. Technical description of the TwinXR package consisting of four scripts, namely, GlobalInstance.cs, JsonReader.cs, JsonWriter.cs, and SimpleJSON.cs. (Source: [4]).
Figure 3. Technical description of the TwinXR package consisting of four scripts, namely, GlobalInstance.cs, JsonReader.cs, JsonWriter.cs, and SimpleJSON.cs. (Source: [4]).
Computers 15 00568 g003
Figure 4. Screenshot of the implemented MR application “HoloCrane” (Unity developer view on desktop). (Source: [4]).
Figure 4. Screenshot of the implemented MR application “HoloCrane” (Unity developer view on desktop). (Source: [4]).
Computers 15 00568 g004
Figure 5. Six-axis industrial robotic manipulator showing the rotational degrees of freedom. (Source: Authors).
Figure 5. Six-axis industrial robotic manipulator showing the rotational degrees of freedom. (Source: Authors).
Computers 15 00568 g005
Figure 6. Workflow of a mixed human–robot warehouse system, showing VR-based staff training before deployment, AI-driven robotic picking during active operations, and AR-guided human intervention for exception handling. (Source: Authors).
Figure 6. Workflow of a mixed human–robot warehouse system, showing VR-based staff training before deployment, AI-driven robotic picking during active operations, and AR-guided human intervention for exception handling. (Source: Authors).
Computers 15 00568 g006
Figure 7. Architectural layers of an agentic robotic system. The flowchart illustrates the hierarchical data flow from sensory input (Perception Layer) to decision-making (Cognitive Layer). (Source: Authors).
Figure 7. Architectural layers of an agentic robotic system. The flowchart illustrates the hierarchical data flow from sensory input (Perception Layer) to decision-making (Cognitive Layer). (Source: Authors).
Computers 15 00568 g007
Figure 8. Implementation roadmap for the integrated AI-robotic system. The four-phase approach details the progression from initial strategy (Needs Assessment) through virtual simulation (Digital Twin) and physical testing (Hardware Pilot). (Source: Authors).
Figure 8. Implementation roadmap for the integrated AI-robotic system. The four-phase approach details the progression from initial strategy (Needs Assessment) through virtual simulation (Digital Twin) and physical testing (Hardware Pilot). (Source: Authors).
Computers 15 00568 g008
Table 1. Comparison of established industrial architectures and the proposed SME framework.
Table 1. Comparison of established industrial architectures and the proposed SME framework.
Architectural DimensionRAMI 4.0IIRAISA-95Digital Twin ArchitecturesProposed SME Framework
Primary purposeIndustry 4.0 reference architectureIIoT reference architectureEnterprise-control integrationPhysical–virtual representation and simulationSME commercial automation
Layered architectureYesYesYesVariableYes
Enterprise integrationYesYesCore focusPartialYes
Physical asset integrationYesYesYesCore focusYes
Multimodal perceptionNot prescribedNot prescribedNoPossibleYes
Computer visionNot prescribedNot prescribedNoPossibleYes
LLM-based reasoningNot prescribedNot prescribedNoNot coreYes
Multi-agent orchestrationNot prescribedPossibleNoPossibleYes
Deterministic optimizationNot prescribedPossiblePossiblePossibleYes
Robotic executionSupported conceptuallySupportedIndirectSupportedYes
Digital twinCompatibleCompatibleNot coreCore focusYes
XR training/interventionNot coreNot coreNoPossibleYes
Human-in-the-loopSupported conceptuallySupportedPartialPossibleExplicit
Safety/validation gateNot specifically definedCross-cutting concernControl-orientedPossibleExplicit
SME-specific resource constraintsNoNoNoNot necessarilyCore focus
Phased implementation roadmapNoNoNoNot necessarilyYes
KPI-gated scale-upNoNoNoPossibleYes
WMS/ERP integrationCompatibleYesYesPartialYes
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Shurdhi, S.; Zyka, E.; Bekteshi, L. An AI-Driven Framework for Automating SME Commercial Workflows with Robotics and Immersive Technologies. Computers 2026, 15, 568. https://doi.org/10.3390/computers15090568

AMA Style

Shurdhi S, Zyka E, Bekteshi L. An AI-Driven Framework for Automating SME Commercial Workflows with Robotics and Immersive Technologies. Computers. 2026; 15(9):568. https://doi.org/10.3390/computers15090568

Chicago/Turabian Style

Shurdhi, Sokol, Eglantina Zyka, and Luan Bekteshi. 2026. "An AI-Driven Framework for Automating SME Commercial Workflows with Robotics and Immersive Technologies" Computers 15, no. 9: 568. https://doi.org/10.3390/computers15090568

APA Style

Shurdhi, S., Zyka, E., & Bekteshi, L. (2026). An AI-Driven Framework for Automating SME Commercial Workflows with Robotics and Immersive Technologies. Computers, 15(9), 568. https://doi.org/10.3390/computers15090568

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop