Next Article in Journal
EKA—Enterprise Knowledge Assistant: Collaborative Multi-Agent AI for Large Claims Handling
Next Article in Special Issue
Machine Learning Operations on ZYNQ FPGA Board for Real-Time Face Recognition
Previous Article in Journal
Design of the Electric Power Control System for a Hydrogen-Fed AEMFC Polymeric Fuel Cell Generator to Power a 0.75 KW DC Motor
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

INTELLECTUM: A Hybrid AR-VR Metaverse Framework for Smart Cities

The Artificial Intelligence Research Center, Novosibirsk State University, Novosibirsk 630090, Russia
*
Author to whom correspondence should be addressed.
These authors contributed equally to this work.
Appl. Syst. Innov. 2026, 9(3), 61; https://doi.org/10.3390/asi9030061
Submission received: 6 January 2026 / Revised: 7 March 2026 / Accepted: 13 March 2026 / Published: 17 March 2026
(This article belongs to the Special Issue Information Industry and Intelligence Innovation)

Abstract

This work presents INTELLECTUM as a reference architecture and design-time evaluation framework for multi-entity XR–AI–digital twin systems. Rather than optimizing a specific implementation, the paper formalizes architectural invariants, event semantics, and coordination mechanisms that precede and inform system realization. INTELLECTUM provides a conceptual framework for structuring interactions across physical and virtual environments, emphasizing human-centered design, immersive digital twins, and collaborative extended-reality workspaces. The technical specification defines core architectural components, human integration modalities via WebXR and heterogeneous sensor networks, and representative usage scenarios within smart city ecosystems. By enabling AI-assisted urban planning, interactive simulation, and multi-actor coordination, INTELLECTUM positions itself as an XR-based architectural foundation for next-generation smart city platforms.

1. Introduction

Rapid urbanization and increasing digitalization of metropolitan environments have intensified the need for advanced human–machine interfaces that bridge physical and virtual systems in a coherent manner. Smart city (SC) initiatives increasingly rely on digital twins, immersive visualization, and artificial intelligence (AI) assisting decision support to manage complex urban infrastructures. However, existing platforms often fragment these capabilities across disconnected tools, lack formal coordination mechanisms between heterogeneous actors, and provide limited support for continuous interaction across reality–virtuality boundaries.
INTELLECTUM is introduced in this work as a conceptual architecture that addresses these challenges at the design level. Rather than proposing a specific implementation, the framework formalizes architectural principles for integrating extended reality (XR), artificial intelligence agents, the Internet of Things (IoT), and digital twins (DTs) into a unified interaction environment. The focus is on defining event semantics, coordination logic, and interaction invariants that enable consistent behavior across physical and virtual contexts, independent of underlying hardware or deployment choices.
A continuous reality–virtuality spectrum allows human participants to smoothly transition between augmented, mixed, and fully virtual representations of the same urban environment. Unlike conventional XR systems that treat AR and VR as separate modes, the architecture conceptualizes immersion as a controllable parameter within a unified interaction model. This design supports continuity of perception, spatial orientation, and task context, which is essential for urban planning, infrastructure monitoring, and collaborative decision-making scenarios. Shared environments also support transparent scenario exploration and participatory decision-making among municipal stakeholders in urban digital twins (UDTs).
Beyond human interaction, the framework formalizes multi-entity coordination across humans, AI agents, and embodied systems such as robots or sensor-equipped infrastructure components. By abstracting interaction through a common event-driven framework, the architecture enables heterogeneous actors to operate within a shared spatiotemporal model of the city. This approach supports collaborative workflows in which human intent, automated reasoning, and environmental data jointly influence system behavior.
The objective of this work is architectural soundness rather than empirical optimization. INTELLECTUM integrates heterogeneous subsystems (XR, AI, UDT, and IoT) whose full operational integration across distributed physical deployments represents a multi-year engineering effort beyond the scope of a single research prototype. The architecture spans multiple coordination boundaries: real-time perceptual rendering in XR clients, event-driven synchronization through WebRTC/WebSocket channels, deterministic state evolution in the Event Store, asynchronous agent inference loops, and physical sensor feedback from robotic inspectors (Figure 1). Validating these interactions empirically would require simultaneous deployment of all subsystems under production-scale conditions, which is infeasible at the architectural design stage.
Consequently, evaluation is conducted at design time through formal validation of architectural properties rather than runtime benchmarking. Design-time evaluation assesses whether the proposed architecture satisfies its stated requirements under formal interaction models, structural constraints, and complexity bounds. This approach is familiar in distributed systems and reference architecture research, where feasibility, consistency, and scalability are validated prior to implementation [1,2].

2. Related Works

2.1. Public Urban Digital Twins

DTs represent next-generation SC architectures by integrating planning, management, and service delivery within a unified computational framework. Their application spans urban development, infrastructure coordination, and public service optimization [3]. UDTs are not neutral technical artifacts—they reflect underlying governance objectives, planning paradigms, and institutional priorities. Without explicitly aligning social, political, and technical dimensions, deployment of DT may risk internal inconsistency and limited policy relevance in urban management contexts [4].
SC concepts can be applied within municipal infrastructure management by embedding digital monitoring, analytics, and coordination tools directly into existing engineering workflows. Such integration supports policy-relevant decision-making by linking technical system states with asset management and long-term investment planning [5].
Recent research in UDTs has increasingly incorporated immersive technologies to move beyond static visualization toward interactive, human-centered decision-making environments. Luo et al. [6] present a perception-powered UDT that integrates photo-realistic rendering with immersive virtual reality to capture and predict subjective human perceptions such as safety and liveliness in simulated urban scenarios. By combining semantic segmentation of urban imagery with VR-based perception experiments, their work demonstrates how human perceptual feedback can be systematically embedded into a DT pipeline. While the focus is primarily on urban planning and sustainability, the study illustrates a shift toward treating the DT as an interactive spatiotemporal model in which human perception becomes an operational input rather than an external evaluation layer.
DTs constitute a parallel research stream relevant to engineering-oriented metaverse systems. Mazzetto [7] reviews UDT integration and highlights their importance for SC management, sustainability, and predictive urban analytics. Similarly, another bibliometric study of over 4200 DT-related papers identified core themes in IoT-enabled simulation, real-time data fusion, and urban infrastructure modeling [8], while these DT studies emphasize large-scale environmental and infrastructural modeling, they generally lack mechanisms for deterministic multi-agent collaboration or decentralized and trusted contributions.
DT cities represent a shift toward data-driven models of urban governance, where real-time data streams, simulation, and automation inform policy coordination. The convergence of AI, IoT, and blockchain technologies enables new governance mechanisms but also introduces complexity in accountability and system oversight [9]. Integration of IoT and agents under UDT is particularly important for adaptive urban systems, requiring solutions to ensure data fusion, interoperatibility, and real-time consistency for agent–robot–human coordination across city-scale environments [10]. Recent research demonstrates that coordinated DT ecosystems improve interoperability, hotspot prediction, and eco-driving assistance in metropolitan environments [11].
XR technologies can broaden public participation in urban planning by enabling immersive visualization and remote engagement. Such tools improve transparency and facilitate feedback collection, but their effectiveness depends on usability, accessibility, and integration with existing participatory processes [12]. Web-based XR solutions demonstrate potential for professional urban planning applications by improving accessibility and reducing hardware barriers. Adoption is strongly influenced by interface familiarity and alignment with established planning practices [13].

2.2. XR/VR and Shared Spatiotemporal Interaction

Integration of AR experience with robotics expands possibilities for industrial and urban tasks. Qian et al. [14] provided a comprehensive review of AR integration in robotic-assisted surgery. Their survey analyzed hardware architectures, application paradigms, and clinical findings showing how AR can enhance surgical precision, reduce visual distraction, and improve situational awareness in minimally invasive procedures. The authors also discuss challenges such as visual clutter and propose mechanisms to ensure safe and efficient AR-based interfaces. This demonstrates how AR can already enhance precision in various contexts, being adaptable also to urban environments.
Diehl et al. [15] introduced an AR interface that enables humans to verify robot learning outcomes prior to execution. The system embedded a virtual simulation directly into the physical environment and coupled it with semantic action information, allowing users to assess whether the robot had acquired a task correctly. The authors compared three visualization designs, including AR-enhanced and traditional RViz-based interfaces, and conducted a user study with 18 participants. Results displayed that AR enhances realistic feedback, improves error detection performance, and reduces the need for environment modeling, highlighting the advantages of AR for human–robot collaboration.
Beyond planning-oriented applications, XR-enhanced DTs have been proposed as coordination environments where humans and autonomous agents jointly operate. Wang et al. [16] introduce XR-DT, an extended reality–enhanced DT framework for agentic mobile robots that tightly couples real-time sensor data, simulated environments, and human feedback through AR, VR, and MR interfaces. Their hierarchical architecture enables bidirectional understanding between humans and robots, allowing both to reason over a shared spatiotemporal representation of the environment. Importantly, the framework treats human interactions, robot actions, and environmental dynamics as co-evolving elements within the same DT, thereby supporting interpretable and trustworthy human–robot interaction in dynamic, shared spaces.
Complementing system- and agent-oriented perspectives, recent work on embodied interaction in DTs highlights the importance of precise spatial and temporal alignment for social and collaborative experiences. Zhang et al. [17] extend the DT paradigm to location-based social VR by integrating full-body avatars, hand tracking, and synchronized physical sensations. Their results show that accurate alignment between physical and virtual environments enables natural social gestures and interactions, fostering trust and presence among participants. Although focused on social VR, this work reinforces the idea that DTs can serve as shared spatiotemporal spaces where humans and AI can coordinate actions and perceptions in real time.

2.3. Human Integration Technical Specification

Immersion and transparency remain central challenges in AR-enhanced robotic applications. Although AR can significantly improve visual feedback and situational awareness, the perception is still constrained by the field of view (FoV), whether using a monitor, head-mounted display, or projection system. A limited FoV restricts the amount of contextual information available, potentially reducing immersion and impairing the operator’s ability to interpret robot behavior or environmental cues. Therefore, achieving effective AR–robot integration requires careful design of visual overlays, adaptive rendering strategies, and interface techniques that mitigate FoV limitations while preserving clarity and operational transparency [18].
The role of XR/VR in integrating humans into adaptive decision loops is further explored by Yigitbas et al. [19], who investigate human-in-the-loop adaptive systems using DTs coupled with immersive VR interfaces. Through procedural and mixed-initiative control strategies, humans can either directly intervene in system adaptations or specify goal states that are resolved by automated planners. This approach demonstrates how XR-based DTs can function as shared operational models where human intent and machine reasoning interact within a unified temporal and spatial context.

2.4. Metaverse Modalities

Recent research on metaverse systems, DTs, and cyber–physical integration reveals a rapidly evolving ecosystem that spans XR interfaces, real-time data infrastructures, AI agents, and decentralized computation layers. A recent large-scale survey by Rawat et al. [20] provides an extensive overview of metaverse fundamentals, highlighting its architectural requirements, emerging standards, enabling technologies (XR, AI, and ubiquitous computing), and cross-domain applications. Their work underscores the need for architectures capable of supporting multimodal inputs, distributed computation, and complex cyber–physical capabilities for interactions.
From the perspective of real-time metaverse environments, Hatami et al. [21] analyze the technological foundations required for low-latency, high-fidelity virtual worlds. Identified challenges involved synchronization delays, sensor–network bottlenecks, and interoperability limitations between digital and physical systems.
The vision in conventional social metaverse (CSM) focuses primarily on social presence, embodiment, and avatar-centric interaction, coupled with infrastructure that relies on centralized cloud rendering, actor replication, and proprietary XR pipelines [22]. This results in strong performance for social spaces but limited guarantees for interaction determinism, verifiable state evolution, or persistent multi-entity coordination. In contrast, INTELLECTUM is designed as a coordination-centric XR environment, emphasizing deterministic event synchronization, multi-entity interaction governance, and adaptive reality–virtuality blending.

2.5. Privacy in Real-Time Applications

Recent literature highlights that security and privacy remain critical barriers to the deployment of large-scale XR-enabled SC platforms. Braun et al. [23] demonstrated that SCs introduce uniquely high-dimensional data streams, which significantly amplify the privacy attack surface. Their analysis identifies five fundamental vulnerabilities: privacy leakage from sensor-rich environments, large-scale network exposure, untrustful data exchange, misuse of AI analytics, and cascading failures in interconnected systems.
On security and system integrity, Wang et al. [24] provide a detailed survey of metaverse security and privacy challenges, noting issues such as identity spoofing, cross-domain data tampering, and the absence of verifiable synchronization between virtual and physical worlds.
Client-side local re-projection and time-warp techniques aim to capture previously received update frames in accordance with the most recent tracking or state information (display refresh), thus bridging the temporal gap when remote updates arrive late. The time warp represents an extended VR experience in which motion-to-photon (MTP) delays are reduced by corrections for the optical aberration of VR lenses and transformation of the stereoscopic rendering from real-time head tracking information to the correct view at applicable display [25]. The time warp has been applied to increase the perceived frame rate and their consistency in asynchronous stereoscopic rendering, allowing optimized hardware usage as rendering with the display refresh cycle does not require direct synchronization [25].
Real-time communication (RTC) systems use an ABR algorithm to achieve minimal delays and maximum quality. RTC adjusts video bitrate dynamically based on available network bandwidth. Since issues typically arise from the conflict between the network’s bitrate targets and codec-based bitrate control, Li et al. [26] proposed the utilization of multi-agent reinforcement learning with multi-dimensional ABR algorithm to provide enhanced quality of experience (QoE). Their approach seeks adaptive behavior with collectively adjusted encoding factors, such as resolution, state-based frame rates, and quantization parameters. These pre-configurations established low latencies without compromising quality in peer-to-peer contexts by dynamically adjusting encoded parameters based on available bandwidth and video feedback [26].
Dead reckoning allows robotic entities, AI agents, and virtual avatars to maintain plausible continuity of motion even during packet loss or delayed updates. It is applied in distributed applications that require controlled bandwidth usage and latency [27,28]. Simulations in distributed virtual environments (DVEs) and online gaming have particularly drawn interest. Dead reckoning methods typically apply kinematic models directly for movement estimates, but further optimizations have been obtained by extrapolating environmental variables and human interactions [28]. Solutions for dead reckoning modeling can touch different contexts. Lin et al. [29] introduced multinomial logit models (MLM) with predict function sets for image smoothness, establishing discrete decision boundaries based on smoothing steps. MLM was applied to aerospace applications but also sets one set of tools for DVEs in highly diverse contexts.
A 2025 taxonomy-based review by Antuley et al. [30] emphasizes that future SC ecosystems will require hybrid authentication and authorization models capable of spanning AI-driven services, blockchain-secured infrastructures, and quantum-resilient communication layers. Their work highlights the growing demand for frameworks that guarantee confidentiality, integrity, availability, and trust in distributed urban infrastructures. From a governance perspective, Barcik et al. [31] show that the digitalization of cities requires robust mechanisms to protect critical urban services from cyberattacks, misinformation streams, and systemic disruptions. Findings underscore that DTs expand the security footprint by mirroring real infrastructure, making the virtual model itself a target.
Latency-related security vulnerabilities have also been highlighted in recent XR research. Monaco et al. [32] demonstrate that predictive latency models can substantially reduce user-perceived delay in distributed interactive systems such as cloud gaming. Their machine-learning-based approach forecasts network conditions to mitigate spikes and concept drift. Peter and Yazdanian [33] apply pose-prediction methods to compensate for network delays in offloaded XR computations, ensuring that user movements remain synchronized with remotely rendered content. INTELLECTUM integrates these latency-compensation techniques (dead reckoning, adaptive bitrate control, and predictive rendering) to maintain temporal stability and reduce attack vectors that exploit timing inconsistencies.

3. INTELLECTUM Interaction Pipelines

The protocols introduced in the following sections describe logical interaction flows rather than low-level implementations. The interaction protocols are closer to state-transition systems than numerical algorithms. Data structures, serialization formats, and failure recovery are implementation-dependent and outside the scope of this architectural specification.

3.1. Cognitive Environment

The cognitive environment necessitates context-aware data between entities with sufficient precision. High precision is required, especially in non-VR states where actuated AI entities or robots are integrated into real urban environments. However, network capacity may become a bottleneck if all raw sensor data are processed.
One approach for optimization is to apply reinforcement learning to establish a dynamic mode switch between lightweight updates and high-fidelity sensory sharing [34]. This is particularly applicable for AI agents, which could recognize and prioritize messages based on urgency. For example, concise status and event messaging over the global or regional network provides a persistent record that can be handled with lightweight updates. In high-urgency situations (e.g., emergency maneuvers), agents switch to high-fidelity mode to provide immediate response with rich sensory data. To ensure system integrity and trust, each agent has a reputation score that defines its access to high-fidelity communications. Since the agent interacts in a 3D workspace, its positions and event locations should be mapped in order to estimate its state at a specific time. Following previous research [34], the state for the agent at time t can be defined as (Equation (1)):
s t = { p ( t ) , T ( t ) , E ( t ) , L ( t ) }
where p ( t ) is the position and proximity to other agents or entities, T ( t ) expresses the active task, E ( t ) is an event indicator, and L ( t ) specifies the computational load in the current network.
Algorithm 1 lays the high-level interaction pipeline for interactive cross-communications for cognitive systems. Implementation-specific concerns such as internal data structures, message serialization, synchronization primitives, error handling, and failure recovery are delegated to the underlying runtime and network middleware and are therefore omitted for clarity.
The cognitive environment seeks to integrate both symbolic and probabilistic components in hybrid AI reasoning, which influences the formalization of the trust system. Intelligent agents operating within this layer rely on a dual inference pipeline where a symbolic rule base enforces deterministic constraints and an LLM-driven probabilistic reasoner provides flexible contextual inference. To ensure that agent decisions remain both adaptive and verifiable, a hybrid trust model is employed to evaluate the correctness of LLM-generated outputs against symbolic constraints. Proposition 1 does not assert logical completeness or probabilistic optimality. It provides a practical evaluation model under the assumption that symbolic constraints are sound and that LLM confidence estimates are calibrated.
Proposition 1
(Hybrid AI Trust Evaluation Model). Assume the following conditions hold:
1. 
Let a hybrid agent combine a symbolic rule base R with a probabilistic LLM reasoner L .
2. 
Given an inference result x proposed by L , its correctness probability under the hybrid model is:
Pr [ x is correct ] = λ · Pr [ x R ] + ( 1 λ ) · C L ,
where Pr [ x R ] { 0 , 1 } indicates whether the symbolic rule base verifies the inference, C L is the LLM confidence score, and λ [ 0 , 1 ] is the symbolic grounding coefficient.
3. 
Let trustworthy condition be
Pr [ x is correct ] τ ,
for a trust threshold τ, the system accepts the inference as trustworthy, providing a design-time trust evaluation criterion under explicit assumptions for hybrid symbolic–LLM reasoning.
Algorithm 1 INTELLECTUM Cognitive Environment Interaction Protocol
  1:
procedure XRInteractionLoop
  2:
     e v t CaptureXRInput ( )     ▹ Head pose, controllers, gesture, camera frame
  3:
     SDK . SendEvent ( e v t )             ▹ Forward to SDK-JS layer
  4:
end procedure
  5:
procedure SDK.SendEvent( e v t )
  6:
    if  e v t . t y p e = SceneUpdate  then
  7:
         SyncLayer . Broadcast ( e v t )           ▹ WebRTC SFU → all peers
  8:
    else if  e v t . t y p e = ModelOp  then
  9:
         Store ( e v t . a s s e t )                    ▹ Asset persistence
10:
         Runtime . Notify ( e v t )               ▹ Inform agent runtime
11:
    else if  e v t . t y p e = UserAction  then
12:
         Runtime . QueryAgent ( e v t )          ▹ Interactive cognition request
13:
    end if
14:
end procedure
15:
procedure SyncLayer.Broadcast( e v t )
16:
     s e s s i o n SessionManager . Resolve ( e v t . u s e r )
17:
     WebRTC . Send ( s e s s i o n . c h a n n e l s , e v t )      ▹ Propagate via data channels
18:
end procedure
19:
procedure Runtime.QueryAgent( e v t )
20:
     c o n t e x t ContextBuilder . Collect ( e v t )       ▹ Build multimodal context
21:
     m AI . Retrieve ( c o n t e x t )                  ▹ Associative recall
22:
     a AI . Reason ( m , c o n t e x t )               ▹ Continual inference
23:
     Runtime . Emit ( a )
24:
end procedure
25:
procedure Runtime.Emit(a)
26:
    if  a . t y p e = SceneChange  then
27:
         SDK . EmitToClient ( a )           ▹ Visual updates to XR client
28:
    else if  a . t y p e = AgentAction  then
29:
         SyncLayer . Broadcast ( a )           ▹ Multi-entity propagation
30:
    else if  a . t y p e = Store  then
31:
         Store ( a . p a y l o a d )                 ▹ Persist model output
32:
    end if
33:
end procedure
34:
procedure SDK.EmitToClient( m s g )
35:
     XRClient . Apply ( m s g )          ▹ Modify scene, overlay, avatar, UI
36:
end procedure

3.2. Reality–Virtuality Continuum

While Algorithm 1 describes the cognitive event-processing pipeline that governs communication between the XR Client, software development kit (SDK), synchronization layer, and agent runtime, it abstracts away the details of how the visual environment is actually composed. INTELLECTUM separates these concerns by using the cognitive environment that manages what events occur and which entities respond, whereas the reality–virtuality subsystem determines how the world is rendered for the user. The transition coefficient γ therefore operates orthogonally to the cognitive loop—events propagate identically regardless of immersion level, but the perceptual outcome is blended according to the user’s position along the reality–virtuality continuum. Algorithm 2 formalizes this rendering-specific process and shows how physical and virtual layers are composited in real time. The blending mechanism in Algorithm 2 activates Equation (4) by directly integrating the reality coefficient γ into the rendering pipeline.
Algorithm 2 Reality–Virtuality Blending Protocol in INTELLECTUM
  1:
γ InitRealityCoefficient ( )    ▹ 0 = VR (fully virtual), 1 = physical passthrough
  2:
procedure HandleXRInput
  3:
     e v t CaptureXRInput ( )        ▹ User gestures, headset passthrough, sliders
  4:
    if  e v t . t y p e = AdjustGamma  then
  5:
         γ e v t . v a l u e                ▹ User/device modifies AR–VR blend
  6:
    end if
  7:
     UpdateRealityBlend ( γ )           ▹ Recompute rendered environment
  8:
end procedure
  9:
procedure UpdateRealityBlend( γ )
10:
     P CapturePhysicalFeed ( )        ▹ Camera or passthrough feed as texture
11:
     V RenderVirtualScene ( )            ▹ 3D assets, UI layers, virtual entities
12:
     R γ · P + ( 1 γ ) · V        ▹ Reality–virtuality blending (Equation (4))
13:
     Display ( R )                       ▹ Final composited frame
14:
end procedure
15:
procedure AgentInfluence
16:
     a AgentOutput ( )        ▹ AI retrieval and reasoning may adjust immersion
17:
    if  a . t y p e = AdjustGamma  then
18:
         γ a . v a l u e                  ▹ Agent suggests immersion level change
19:
         UpdateRealityBlend ( γ )
20:
    end if
21:
end procedure
Equation (2) is designed to provide a seamless transition mechanism between reality states through a central toggle controller that manages rendering pipelines, interaction modalities, and data flow based on user positioning along the reality–virtuality spectrum. The continuum is mathematically represented as (Equation (4)):
R = ( γ 1 ) P + γ V
where R is the rendered composite environment, representing the blending ratio between physical and virtual content, with 1 presenting virtual dominance. γ is the virtuality coefficient (a scalar in [ 0 , 1 ] ) ranging from 0 (fully physical) to 1 (fully virtual). V and P are the virtual and physical render buffers, respectively. Physical components include video textures, passthrough XR, depth occlusion, and real-time motion tracking. Virtual components contain 3D models, overlays, holograms, effects, digital avatars, and scene lighting. Four operational regions are recognizable for the rendered environment:
  • Enhanced Reality (ER): Predominantly physical environment with minimal virtual overlays ( γ 0.3 ). ER with a low virtual coefficient is applicable for navigation and overlay displays.
  • Mixed Reality (MR): Balanced physical–virtual composition ( 0.3 < γ < 0.7 ). MR has an evenly blended virtual coefficient level, providing deeper interactions for tasks such as training and design.
  • Augmented Reality (AR): Predominantly virtual environment with residual physical grounding ( 0.7 γ < 1.0 ). AR is dominantly virtual, ideal for fast prototyping and simulations.
  • Virtual Reality (VR): Fully virtual environment ( γ = 1.0 ). VR is a completely virtual state for broadcasting events and DTs.
While γ is a continuous variable, categorical regions ER, MR, AR, and VR are considered as semantic abstractions between different ranges for policy selection, rendering pipelines, and interaction constraints. Device capability influences the maximum achievable reality coefficient γ within the reality–virtuality continuum. Low-end devices maintain higher reliance on physical passthrough ( γ 1 ), while professional XR systems can reliably sustain near-virtual scenes ( γ 0 ) without compromising frame rate or spatial coherence. This adaptive control ensures uniform experience quality across various user hardware.
Seamless reality switching during adaptations to different multi-modal tasks is assigned to the reality switch (Equation (5)), which is defined as:
D T u r b a n = i = 1 n S i · Φ ( I i ) + λ · RealitySwitch ( A R , V R )
where D T u r b a n represents the current rendered state of the UDT, I i denotes raw input from the i-th sensing modality from input streaming devices, S i is a weighting or activation coefficient reflecting the relevance or reliability of that sensor stream, and Φ ( · ) is a modality-specific transformation function that maps raw sensor inputs into a unified DT representation involving actions such as feature extraction, spatial alignment, or semantic encoding. The term RealitySwitch ( A R , V R ) represents a control function governing transitions between augmented and virtual rendering modes, while the scalar coefficient λ regulates the influence of perceptual mode switching relative to sensor-driven updates.
Depending on device capabilities (such as camera passthrough availability, WebXR subsystem support, or desktop camera feeds), the reality coefficient γ (float between 0 and 1) is dynamically adjusted to reflect the user’s perceptual state. A value of γ = 0 corresponds to fully virtual VR environments, whereas γ > 0 introduces increasing contributions of the physical feed, enabling ER and MR experiences. Because this protocol is decoupled from cognitive routing, INTELLECTUM supports seamless transitions between AR and VR modes, continuous reality blending during agent interaction, and adaptive immersion levels that can be modulated by user input, device sensors, or higher-level agent reasoning. This establishes a unified and extensible perceptual layer that integrates physical and digital worlds.

3.3. Multi-Entity Integration Framework

Multi-entity integration framework is organized using the synchronization layer, the core runtime responsible for real-time message propagation, state harmonization, and cross-entity coordination. It serves as the canonical interface through which all agents, robots, and humans exchange state updates.
The synchronization layer provides three fundamental services: (i) temporal synchronization of distributed entities, (ii) session-level consistency through WebRTC channels and verifiable state diffs, and (iii) event propagation toward persistence (data infrastructure layer). Through this mechanism, the synchronization layer acts as the arbitration point that guarantees causality, ordering, and reproducibility of multi-entity interactions across diverse environments.
Algorithm 3 for interactions is executed within the synchronization layer, which functions as the authoritative coordination point for multi-entity communication. Upon permission validation, the synchronization layer establishes a session channel between the entities, ensuring that interaction events follow the system’s global ordering constraints. The synchronization layer logs the interaction, updates both local and distributed entity states, and disseminates the resulting changes to all subscribed peers.
A unified synchronization core guarantees secure, deterministic, and extensible multi-entity coordination across agents, robots, and human users. The synchronization layer functions as the authoritative gateway for all cross-entity communication, enforcing global ordering, validating permissions, and propagating authenticated state transitions. Each interaction request is validated through lightweight cryptographic checks, ensuring that entity identities, session keys, and capabilities are verified before any state mutation occurs. This establishes a trust mechanism across distributed XR, robotic, and AI-driven environments.
Algorithm 3 Multi-Entity Interaction Protocol
1:
procedure EntityInteraction( e n t i t y 1 , e n t i t y 2 , i n t e r a c t i o n _ t y p e )
2:
     a u t h VerifyPermissions ( e n t i t y 1 , e n t i t y 2 )
3:
    if  a u t h = T r u e  then
4:
         s e s s i o n CreateInteractionSession ( e n t i t y 1 , e n t i t y 2 )
5:
         LogInteraction ( s e s s i o n , i n t e r a c t i o n _ t y p e )
6:
         UpdateEntityStates ( e n t i t y 1 , e n t i t y 2 )
7:
         BroadcastStateChange ( e n t i t y 1 , e n t i t y 2 )
8:
    end if
9:
end procedure
Synchronization is achieved through a hybrid real-time pipeline where WebRTC data channels are combined with session-level consistency guarantees. The synchronization layer broadcasts entity updates using timestamped state diffs, enabling precise temporal alignment between participants while preventing race conditions and divergent world states. To maintain deterministic replay and conflict-free state reconstruction, the synchronization layer archives all events to the data infrastructure and selectively anchors critical state commitments to the persistent and event-based storage. This upward and downward propagation ensures coherence across volatile, semi-persistent, and immutable storage domains. Consistency of multi-entity synchronization is ensured when (Equation  (6)):
τ i < τ j S i S j
The cross-entity trust score is computed as:
T i = α · R i + ( 1 α ) · H i
where R i represents the entity’s reputation and H i provides historical consistency and compliance. Trust score blending parameter α balances the influence of recent reputation versus historical consistency when computing an entity’s overall trust score T i .
Extensibility is embedded into the framework through modular interfaces that expose standardized protocols for introducing new entity types, external systems, or behavioral modules. Interoperability with robotics ecosystems, IoT networks, and simulation pipelines is achieved through the interoperability layer, which consumes synchronized state events and translates them into external semantics. Similarly, intelligent agents and human XR clients can be extended with new sensing modalities, behavioral policies, or learning components without modifying the synchronization core.
To formalize cross-entity coordination, we define each entity e i by a state vector:
s i ( t ) = { p i ( t ) , v i ( t ) , c i ( t ) , θ i ( t ) }
where p i is position, v i velocity, c i capability vector, and θ i cognitive (or intentional) state. The synchronization layer computes a global scene state:
S ( t ) = i = 1 N s i ( t )
All interaction events form a time-ordered log:
L = ( E 1 , τ 1 , a 1 ) , ( E 2 , τ 2 , a 2 ) , , ( E k , τ k , a k ) ,
where E k { e 1 , , e N } is the set of entities participating in the k-th event (allowing unary, pairwise, or multi-entity interactions), τ k is the Lamport timestamp, and a k is the interaction action.
The causality condition in Equation (6) ensures that all clients and agents interpret events in the same logical order, even if network latency causes different physical arrival sequences. When the synchronization layer receives updates, it applies a deterministic merge operator that reconstructs the world state based on Equation (9) precedence, guaranteeing that every participant converges toward the same global scene regardless of timing variations or packet loss.
The integration of trust scoring (Equation (7)) further refines state interpretation by weighting updates based on the reputation and historical reliability of each entity. This becomes essential in environments where physical robots, IoT devices, human operators, and autonomous agents contribute simultaneously to the evolving world state. Entities with low consistency or repeated contradictory updates can be automatically down-weighted, while reliable entities gain stronger influence on the synchronized global state S ( t ) . Proposition 2 states that coordination remains polynomial assuming bounded entity state size and fixed merge operators, not as a universal complexity claim.
Proposition 2
(Bounded Coordination Complexity). Assume the following:
1. 
There are n active entities and m events in the Lamport log L .
2. 
State merging is performed using algorithm MergeState with complexity O ( n · f ( n ) ) , where f ( n ) is the cost of merging individual entity states.
3. 
Lamport timestamp evaluation is performed for all events in L in O ( m ) time.
4. 
Conflict-Free Replicated Data Type (CRDT) reconciliation for conflict-free updates has complexity O ( n · g ( n ) ) , where g ( n ) is the cost of merging individual CRDT objects.
Then, the total decision-making time for any multi-entity interaction in INTELLECTUM satisfies:
T decision = O ( n · f ( n ) + m + n · g ( n ) ) .
If f ( n ) and g ( n ) are polynomial, the overall complexity is polynomial in n and m.
Proposition 2 represents polynomial-time bounded coordination for all synchronized multi-entity interactions. Because all interaction events are recorded in the log (Equation (10)), the presented architecture can perform deterministic replay, anomaly detection, and post hoc verification. The log serves as a causal trace of the entire system’s evolution, enabling external systems to reconstruct past states using the same deterministic merge rules as the live system.
Overall, this design supports coordinated reasoning, conflict-free state propagation, and extensible behavior across multi-entities. Whether an update originates from a robot, an XR client, an AI agent, or an IoT sensor, the same mathematical structure defines how it enters and influences the shared environment.
An actor is defined as an entity that can emit and conduct event-triggered synchronization. Actor abstraction involves human XR clients, AI agents, and robotics/IoT devices. Equation (11) formalizes the actor abstractions:
a i : = i d i , s i ( t ) , C i , P i
where i d i is a globally unique identifier, s i ( t ) is the local state vector for coordination (Equation (8)), C i is the capability set, and P i is the trust profile. Additionally, the decoupling of actors can be formally defined as Equation (12):
a i , a j A , i j : e i e j | sync ( S ( t ) )
where a i , a j A are distinct actors in the system, e i , e j are the event streams emitted by actors a i and a j , respectively, S ( t ) is the global synchronized scene state (Equation (9)), and sync ( S ( t ) ) denotes the synchronization layer that merges events. ⊥ indicates that the event streams are independent given the synchronization process.

3.4. Latency Management

To compensate for delays that arise from network jitter, sensor processing, WebRTC transport, and rendering overhead, INTELLECTUM uses predictive motion modeling and time-warp rendering strategies, similar to latency-compensation methods described in recent XR research [32,35,36,37].
The predictive rendering function for maintaining synchronization demands is (Equation (13)):
P r e n d e r ( t + Δ t ) = S c u r r e n t + V · Δ t + 1 2 A · ( Δ t ) 2
where P r e n d e r is the predicted render position of an entity, S c u r r e n t is the current state, V is the velocity vector, and A is the estimated acceleration extracted from inertial measurement unit (IMU) and positional stream data.
To mitigate network variability, adaptive bitrate control (ABR), client-side local re-projections and time warp techniques, dead-reckoning motion models, and frame-level interpolation can be applied.

4. System Architecture

Several decisions affected the architectural solutions. These were formalized as requirements: event-driven design, formal event taxonomy, architectural properties, and public domain considerations.

4.1. Requirements

The architecture is derived from a set of design-time requirements that constrain how interactions across humans, agents, robots, and DTs can be coordinated between reality states. These requirements act as preconditions that rule out monolithic, synchronous, or tightly coupled designs and motivate the event-driven architecture introduced in Section 4.2. Validation of interactions between intelligent and automatic entities across reality states, architecture of INTELLECTUM was established based on the following requirements:
  • Integration of adaptive cognitive systems
    The architecture must support heterogeneous reasoning processes (human intent, AI inference, and automated control) without assuming a shared execution model or synchronized internal state. This precludes centralized control logic and necessitates event-based coordination.
  • Controllable reality–virtuality continuum mechanism for dynamic state transitions in interface
    Reality transitions (ER-AR–MR–VR) must be dynamically adjustable at runtime without invalidating system state or coordination logic. This requires separation between authoritative state changes and view-dependent rendering updates.
  • Multi-entity collaboration across heterogeneous actors
    Humans, AI agents, robots, and infrastructure components must participate as first-class entities in shared interactions, with no hard-coded role assumptions. The architecture must therefore abstract interaction through a uniform event model rather than direct method coupling.
  • Asset tracking and state synchronization
    Distributed entities must observe a consistent view of shared assets and environments under partial failures and asynchronous communication, requiring explicit event ordering, replayability, and conflict-tolerant state reconciliation.

4.2. Event-Driven Design

Figure 1 introduces a hybrid event-driven architecture following Command–Query Responsibility Segregation (CQRS), where state mutation events and read-model synchronization are decoupled to ensure scalability and deterministic replay.
By construction, the event propagation graph induced by Figure 1 admits a strict partial order over components. All durable events are appended to an append-only log L with Lamport timestamps. Role separation uses CQRS, meaning that projections never emit commands, nor do renders mutate authoritative state. Ingress, domain, projection, and real-time signal events are stacked into non-overlapping classes with monotonic timestamp assignment, preventing any causal feedback loop. Consequently, the event flow graph is directed and acyclic. The system is designed to scale to any number of actors without requiring coupling between them. This is constrained by the event-driven architecture (Figure 1) and the actor abstraction formalism (Equation (11)).

4.3. Formal Event Taxonomy

A disjoint and complete event taxonomy is defined as E , where each class encodes distinct semantic roles and routing constraints. Perceptual events affect only representation, state mutation events modify authoritative DT state, and governance events provide verifiable accountability without influencing perception. Formal event taxonomy is presented in Table 1.

4.4. Architectural Properties

Architectural properties define structural and behavioral guarantees, including clear boundaries between components, enforcement of event ordering, and consistency of the global synchronized scene state S ( t ) under multi-entity interactions. Assume the following formal conditions hold:
  • Let the global synchronized scene state be defined as S ( t ) (Equation (9)).
  • Let all interaction events be organized into the Lamport-ordered log L (Equation (10)).
  • Assume that the synchronization layer performs:
    (i)
    state-vector merging on S ( t ) ;
    (ii)
    causality and ordering checks on L ;
    (iii)
    trust-score and permission validation, each using polynomial algorithms.
  • Then the total decision-making time for any multi-entity interaction in INTELLECTUM satisfies:
    T decision = O ( n c ) ,
    where n is the number of active entities and c is a constant determined by the combined complexity of state merging, Lamport timestamp evaluation L , and CRDT-based reconciliation.
These bounds hold under standard assumptions of reliable session connectivity and bounded message delays, consistent with distributed interactive systems literature. INTELLECTUM does not model intelligence nor its own cognition—its primary goal is to model coherent events between entities.

4.5. Public Domain Considerations

Architectural adaptations for the public domain and UDTs were constrained to three core considerations. Generalization of conditions was defined for collaborative workspaces, adaptive public services, and privacy and surveillance management.

4.5.1. Collaborative Workspaces

Collaborative environments allow distributed teams to co-manipulate DT layers, annotate real-time sensor data, and coordinate urban interventions. Recent SC research indicates that coordinated multi-user DTs improve decision-making efficiency and reduce operational friction in complex urban tasks [11]. Collaboration efficiency is expressed (Equation (15)):
E c o l l a b = i = 1 n T s u c c e s s ( i ) i = 1 n T t o t a l ( i ) · C s y n c · Q d a t a
where E c o l l a b represents multi-user performance, T s u c c e s s denotes successful interactions or tasks, C s y n c is the synchronization quality between distributed users, and Q d a t a is the fidelity of shared data streams.

4.5.2. Adaptive Public Services

Context-aware public services dynamically adapt using both real-world sensor readings and AR/VR state overlays. SC administrators may adjust traffic management, environmental controls, or emergency response based on combined physical and digital insights. DT research identifies these adaptive service loops as key drivers of next-generation municipal management systems [38]. Context-aware public services adapt based on reality states (Equation (16)):
S a d a p t e d = β · U p r e f + ( 1 β ) · C e n v · γ ( R s t a t e )
where S a d a p t e d is the computed service output, U p r e f represents user or citizen preferences, C e n v is the environmental context extracted from sensors, and γ ( R s t a t e ) represents the reality-state from INTELLECTUM’s AR/VR continuum.

4.5.3. Privacy and Surveillance Management

To protect user identity and sensor-derived data, INTELLECTUM uses federated learning together with differential privacy. The effective privacy-preserving transformation is defined as (Equation (17)):
M ( x ) = f ( x ) + Lap Δ f ϵ
where M ( x ) is the privacy-preserving output, f ( x ) is the local function (model updates), Δ f is its sensitivity, and ϵ is the privacy budget. Only the privatized updates are transmitted to the coordination layer, while raw sensor data remains on the user’s device.

5. Technical Evaluation

Technical evaluation assesses the feasibility, correctness, and performance characteristics of the architecture by combining formal design-time analysis (Section 5.1) with synthetic evaluation scenarios (Section 5.2).

5.1. Design-Time Evaluation

Design-time evaluation assesses architectural validity without executing a full system implementation. Rather than relying on empirical benchmarks, this evaluation examines whether the proposed architecture satisfies its stated requirements under formal interaction models, structural constraints, and complexity bounds. This approach is common in distributed systems and reference architecture research, where feasibility, consistency, and scalability are established prior to implementation.
Design-time evaluation is supported by the reference constructs summarized in Table 2. Each construct corresponds to a formal model (algorithm, equation, or invariant) defined elsewhere in the paper, providing a structured basis for reasoning about correctness, separation of concerns, and scalability.
To provide design-time evidence that the architecture satisfies the stated requirements, a synthetic evaluation scenario (Section 5.2) symbolically executes representative multi-entity interactions. This scenario verifies that coordination, event semantics, privacy preservation, and scalability constraints are upheld under formal interaction rules.
  • Coordination correctness: all events are routed, merged, and reflected in the global state S ( t ) .
  • Privacy compliance: raw data remains local, and federated updates enforce differential privacy.
  • Scalability: decision-making and synchronization time remain within polynomial bounds, supporting increased agent counts.
  • Event semantics and invariant enforcement: disjoint event types and command-handling invariants are preserved.
By formalizing global state S ( t ) and Lamport-ordered events L , all Table 2 criteria can be reasoned about at design time. Together, the reference constructs and synthetic scenario provide design-time assurance that the architecture meets its intended requirements without requiring a full system implementation. Full empirical validation and latency–cost measurements are deferred to future work.

5.2. Synthetic Evaluation Scenario

To illustrate design-time feasibility, the minimal architecture is applied in a representative collaborative workspace comprising three entities: a human planner, an agent, and a robotic inspector. This configuration captures all core interaction classes supported by the XR input, autonomous reasoning, and sensor-driven actuation. Each entity produces interaction events, state updates, or sensor readings according to its role, formalized by Equations (15) and (16). Algorithm 4 symbolically executes this scenario by generating events per entity, enforcing causal ordering through a Lamport log, merging updates into the global synchronized state S ( t ) , and propagating results through projections.
Rather than simulating concrete values, the algorithm evaluates whether architectural constraints—coordination correctness, privacy preservation, invariant enforcement, and polynomial decision-time bounds—hold under a representative multi-entity workload, thereby validating the architecture at design time.
Federated learning with differential privacy (Equation (17)) is applied to all sensor-derived data. Each agent locally computes updates to the shared DT, applies the Laplace mechanism to privatize sensitive information, and transmits only the perturbed updates to the synchronization layer. Within this synthetic scenario, the architecture can be validated for:
  • Coordination correctness: All events are correctly routed, merged, and reflected in the global state S ( t ) .
  • Privacy compliance: Raw data never leaves local nodes; federated updates maintain differential privacy guarantees.
  • Scalability and timing: Decision-making and synchronization time (Equation (14)) remain within polynomial bounds, demonstrating feasibility for larger teams.
  • Event semantics and invariant enforcement: Disjoint event types (Table 1) and command-handling invariants are preserved.
Algorithm 4 Synthetic Multi-Entity Interaction Evaluation
  1:
procedure SyntheticEvaluation( E n t i t i e s ,   E v e n t R a t e s ,   P r o j e c t i o n s ,   ϵ )
  2:
     n | E n t i t i e s |                   ▹ Number of active entities
  3:
     S ( t ) InitializeGlobalState ( )
  4:
     L InitializeLamportLog ( )
  5:
    for all  e n t i t y E n t i t i e s  do
  6:
         r E v e n t R a t e s [ e n t i t y ]
  7:
         T decision event 0            ▹ Initialize per-entity decision time accumulator
  8:
        for  i = 1 to r do
  9:
            e GenerateEvent ( e n t i t y )
10:
            L . Append ( e )
11:
                                 ▹ Decision-making computation
12:
            C event O ( n c )
13:
            T decision timestep e n t i t y T decision event
14:
                                      ▹ State update and merging
15:
            S ( t ) MergeState ( S ( t ) , e )
16:
            VerifyCausality ( L , e )
17:
            VerifyTrust ( e n t i t y , e )
18:
                                     ▹ Federated learning/privacy
19:
            u ComputeLocalUpdate ( e n t i t y )
20:
            u ˜ u + Lap ( Δ f / ϵ )
21:
            TransmitUpdate ( u ˜ )
22:
        end for
23:
    end for
24:
                         ▹ Propagation through projections/XR output
25:
    for all  p P r o j e c t i o n s  do
26:
         UpdateProjection ( p , S ( t ) )
27:
    end for
28:
     RenderXRScene ( S ( t ) )
29:
                                                 ▹ Scalability check
30:
     T decision timestep e T decision event
31:
    if  T decision timestep > T max  then
32:
        ReportScalingLimit(n)
33:
    end if
34:
end procedure

6. Discussion

The layered approach separating cognition, synchronization, interoperability, and decentralization provides clear modular boundaries. This separation reduces cross-layer coupling while enabling flexible scaling, particularly when integrating AI agents or robotics. The synchronization layer acts as the critical mediator, ensuring determinism, causal message ordering, and re-playable event history.
The reality–virtuality continuum mechanism introduces a dynamic perceptual layer that can adapt to user context, device capability, or agent-driven decisions. However, enhanced perceptual flexibility raises challenges such as visual coherence across various devices, secure mixing of real and virtual visual cues, and potential manipulation risks if attackers alter rendered environments. Mitigating these requires strict integrity checks and perceptual safety audits.
A key open challenge concerns balancing user privacy with high-fidelity multi-modal sensing. Even with federated learning and differential privacy, sensor-rich XR devices generate fine-grained behavioral signatures that remain difficult to keep fully anonymous. Additional research into privacy-preserving XR compression, encrypted sensor fusion, and zero-knowledge interaction proofs is necessary.
An interesting future research direction would involve an integration of memory-centricity as an internal agent property to add persistent and continual agent memory. One such option could harness sparse representations and base-2 or base-3 systems.

6.1. Security and Privacy Challenges

Privacy-preserving mechanisms such as differential privacy are essential for protecting sensitive data in edge computing-based SC applications. Resource constraints and distributed architectures introduce additional challenges that must be considered in system design [39].
In the context of design-time evaluation, using federated learning together with the Laplace mechanism provides a synthetic yet representative validation framework: it allows us to model how updates from multiple distributed entities can be aggregated and privatized without transmitting raw sensor data, ensuring that the architecture correctly supports privacy-preserving coordination, while no real deployment is performed, this combination is sufficient to verify data flow, invariants, and event handling under privacy constraints, demonstrating feasibility within the architectural model.
Extended reality platforms supported by advanced network infrastructures introduce significant security and privacy challenges, particularly due to the handling of biometric data, immersive sensing, and distributed computation. Vulnerabilities related to virtualization, application interfaces, and data exposure must therefore be addressed as first-class design constraints in XR-enabled urban systems [40].
Differential privacy (Equation (17)) enforces a formal privacy guarantee. This approach aligns with modern recommendations for privacy-preserving sensor fusion in smart-city infrastructure [23] and for protecting biometric data collected from XR devices [32,33].

6.2. Cross-System Consistency and Scalability

Achieving deterministic state propagation across XR clients, robots, and AI agents presents complex consistency challenges. SC systems often experience contradictory data sources, asynchronous event timing, and heterogeneous device resolutions. Recent work in DT organization highlights the necessity of federated consistency layers to unify multi-platform updates [11].
INTELLECTUM resolves these challenges through: (i) Lamport-timestamped event ordering inside the synchronization layer, (ii) CRDTs for user annotations and scene edits, and (iii) chain-anchored state hashes to guarantee global consensus. This ensures replayable timelines and verifiable world-state convergence across all entities.
As the user count or sensor density increases, maintaining real-time performance becomes challenging. Large cities may produce millions of data points per minute, requiring massive parallel processing. Recent scalable digital-twin studies emphasize cloud-edge distribution as essential for maintaining responsiveness in high-load SC environments [41].

6.3. Ethical Governance

UDTs are not neutral technical artifacts but reflect underlying governance objectives, planning paradigms, and institutional priorities. Without explicitly aligning social, political, and technical dimensions, DT deployments risk internal inconsistency and limited policy relevance in urban management contexts [4]. Large-scale socio-technical systems introduce ethical risks that extend beyond technical performance, particularly when decision-making is delegated to data-driven and automated infrastructures. Responsible system design therefore requires explicit consideration of stakeholder impacts, governance structures, and long-term social impact, especially in contexts where technology mediates public life and institutional authority [42].
Algorithmic governance systems deployed in SC contexts can reinforce existing social and spatial inequalities when data asymmetries and surveillance mechanisms are insufficiently regulated. These risks highlight the need for accountability and ethical oversight in automated urban decision-making [43].
DTs increasingly function as governance instruments used by multiple urban stakeholders, including policymakers, planners, and infrastructure operators. Challenges arise when representational scope, accountability mechanisms, and institutional responsibilities are insufficiently defined [44]. Social acceptance of XR and AI interfaces in public urban environments depends on perceived usefulness, transparency, and opportunities for feedback. Participatory mechanisms that integrate citizen input into planning workflows enhance legitimacy and trust in SC deployments [45].

7. Conclusions

INTELLECTUM represents a novel metaverse architecture grounded in a reality–virtuality continuum, multi-entity coordination, and event-driven design. Beyond conceptual integration, the architecture is validated at design time through formal interaction models, explicit event semantics, and structural constraints that demonstrate feasibility prior to full system implementation.
The system advances mixed-reality platforms in three key ways. First, it enables seamless perceptual transitions across physical and virtual environments, allowing users and agents to navigate smoothly between AR, MR, and VR states through a formally defined reality-blending mechanism. Second, it supports coordinated interaction among humans, AI agents, and robotic entities via a unified actor abstraction and deterministic event routing, ensuring consistency across shared DT environments.
Design-time evaluation criteria demonstrate that these capabilities are not merely aspirational. Formal reference constructs complete the pipeline to establish architectural coherence, event semantic correctness, and scalability under increasing numbers of entities and sessions. In particular, the separation of real-time and authoritative processing paths, the absence of cyclic dependencies in the event propagation graph, and polynomial coordination bounds establish that the architecture remains tractable and deterministic without reliance on empirical tuning.
Future research directions include optimized cross-chain communication, compression techniques for reality-state transitions, and privacy-preserving mechanisms for multimodal urban data. Further work will also explore agent-based cognitive modeling and phased deployment in real-world urban environments. As a reference architecture, INTELLECTUM provides a formally grounded foundation for next-generation SCs, enabling trustworthy human–AI–robot collaboration while remaining extensible to emerging technologies and deployment scales.

Author Contributions

Conceptualization, A.N. and J.R.; methodology, J.R.; software, J.R.; validation, A.N.; formal analysis, A.N.; investigation, A.N. and J.R.; resources, J.R.; data curation, J.R.; writing—original draft preparation, A.N. and J.R.; writing—review and editing, A.N. and J.R.; visualization, J.R.; supervision, A.N.; project administration, A.N.; funding acquisition, A.N. All authors have read and agreed to the published version of the manuscript.

Funding

This work was supported by a grant for research centers, provided by the Ministry of Economic Development of the Russian Federation in accordance with the subsidy agreement with the Novosibirsk State University dated 17 April 2025 No. 139-15-2025-006: IGK 000000C313925P3S0002.

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

Benchmark implementation and example scripts will be made available at https://github.com/intellectumvr (accessed on 10 March 2026), upon publication.

Acknowledgments

Authors would like to express gratitude to the Artificial Intelligence Research Center of Novosibirsk State University for their support of this research.

Conflicts of Interest

The authors declare no conflicts of interest.

Abbreviations

The following abbreviations are used in this manuscript:
ABRAdaptive Bitrate Control
AIArtificial Intelligence
ARAugmented Reality
CSMConventional Social Metaverse
CQRSCommand–Query Responsibility Segregation
CRDTConflict-Free Replicated Data Type
DTDigital Twin
DVEDistributed Virtual Environment
EREnhanced Reality
FoVField of View
IMUInertial Measurement Unit
IoTInternet of Things
MLMMultinomial Logit Models
MRMixed Reality
MTPMotion-to-Photon Latency
QoEQuality of Experience
RTCReal-Time Communication protocol
SDKSoftware Development Kit
SFUSelective Forwarding Unit
UDTUrban Digital Twin
VCVirtual City
VRVirtual Reality
XRExtended Reality

References

  1. Bass, L.; Clements, P.; Kazman, R. Software Architecture in Practice; Pearson Education: London, UK, 2013. [Google Scholar]
  2. Michelson, B. Event-Driven Architecture Overview; Number 681; Patricia Seybold Group: Boston, MA, USA, 2006; p. 681. [Google Scholar] [CrossRef]
  3. Deren, L.; Wenbo, Y.; Zhenfeng, S. Smart city based on digital twins. Comput. Urban Sci. 2021, 1, 4. [Google Scholar] [CrossRef]
  4. Al-Sehrawy, R.; Kumar, B.; Watson, R. The pluralism of digital twins for urban management: Bridging theory and practice. J. Urban Manag. 2023, 12, 16–32. [Google Scholar] [CrossRef]
  5. Alabi, B.N.; Du, R.; Guo, Y.; Alinizzi, M.; Labi, S. A framework for incorporating smart city concepts into the management of municipal infrastructure. Front. Eng. Built Environ. 2025, 5, 109–124. [Google Scholar] [CrossRef]
  6. Luo, J.; Liu, P.; Xu, W.; Zhao, T.; Biljecki, F. A perception-powered urban digital twin to support human-centered urban planning and sustainable city development. Cities 2025, 156, 105473. [Google Scholar] [CrossRef]
  7. Mazzetto, S. A Review of Urban Digital Twins Integration, Challenges, and Future Directions in Smart City Development. Sustainability 2024, 16, 8337. [Google Scholar] [CrossRef]
  8. El-Agamy, R.F.; Sayed, H.A.; AL Akhatatneh, A.M.; Aljohani, M.; Elhosseini, M. Comprehensive analysis of digital twins in smart cities: A 4200-paper bibliometric study. Artif. Intell. Rev. 2024, 57, 154. [Google Scholar] [CrossRef]
  9. Deng, T.; Zhang, K.; Shen, Z.J.M. A systematic review of a digital twin city: A new pattern of urban governance toward smart cities. J. Manag. Sci. Eng. 2021, 6, 125–134. [Google Scholar] [CrossRef]
  10. Sacoto-Cabrera, E.J.; Perez-Torres, A.; Tello-Oquendo, L.; Cerrada, M. IoT, AI, and Digital Twins in Smart Cities: A Systematic Review for a Thematic Mapping and Research Agenda. Smart Cities 2025, 8, 175. [Google Scholar] [CrossRef]
  11. Nguyen, D.V.; Dao, M.S.; Zettsu, K. Digital Twin Orchestration: Framework and Smart City Applications. In Proceedings of the AI4DT&CP Workshop, in Conjunction with IJCAI 2024, Jeju, Republic of Korea, 3–9 August 2024. [Google Scholar]
  12. Höftberger, K.; Konrath, A.; Berger, A.; Allerstorfer, D.; Krebs, R. XR-Supported Communication in Green Urban Projects. Participating in Urban Change through Virtual and Augmented Reality. In LET IT GROW, LET US PLAN, LET IT GROW. Nature-Based Solutions for Sustainable Resilient Smart Green and Blue Cities. Proceedings of REAL CORP 2023, 28th International Conference on Urban Development, Regional Planning and Information Society; CORP—Competence Center of Urban and Regional Planning: Vienna, Austria, 2023; pp. 1071–1076. [Google Scholar]
  13. Rzeszewski, M.; Orylski, M. Usability of WebXR Visualizations in Urban Planning. ISPRS Int. J. Geo-Inf. 2021, 10, 721. [Google Scholar] [CrossRef]
  14. Qian, L.; Wu, J.Y.; DiMaio, S.P.; Navab, N.; Kazanzides, P. A Review of Augmented Reality in Robotic-Assisted Surgery. IEEE Trans. Med. Robot. Bionics 2020, 2, 1–16. [Google Scholar] [CrossRef]
  15. Diehl, M.; Plopski, A.; Kato, H.; Ramirez-Amaro, K. Augmented Reality interface to verify Robot Learning. In Proceedings of the 2020 29th IEEE International Conference on Robot and Human Interactive Communication (RO-MAN), Naples, Italy, 31 August–4 September 2020; pp. 378–383. [Google Scholar] [CrossRef]
  16. Wang, T.; Byeon, J.; Yehia, A.; Wang, H.; Xu, Y.; Zeng, T.; Wang, Z.; Jiao, J.; Claudel, C. XR-DT: Extended Reality-Enhanced Digital Twin for Agentic Mobile Robots. arXiv 2025, arXiv:2512.05270. [Google Scholar] [CrossRef]
  17. Zhang, J.; Ma, F.; Pi, Y.; Du, H.; Pan, X. Digital twin embodied interactions design: Synchronized and aligned physical sensation in location-based social VR. Front. Virtual Real. 2025, 6, 1499845. [Google Scholar] [CrossRef]
  18. Fu, J.; Rota, A.; Li, S.; Zhao, J.; Liu, Q.; Iovene, E.; Ferrigno, G.; Momi, E.D. Recent Advancements in Augmented Reality for Robotic Applications: A Survey. Actuators 2023, 12, 323. [Google Scholar] [CrossRef]
  19. Yigitbas, E.; Karakaya, K.; Jovanovikj, I.; Engels, G. Enhancing Human-in-the-Loop Adaptive Systems through Digital Twins and VR Interfaces. arXiv 2021, arXiv:2103.10804. [Google Scholar] [CrossRef]
  20. Rawat, D.B.; alami, H.E.; Hagos, D.H. Metaverse Survey & Tutorial: Exploring Key Requirements, Technologies, Standards, Applications, Challenges, and Perspectives. arXiv 2024, arXiv:2405.04718. [Google Scholar] [CrossRef]
  21. Hatami, M.; Qu, Q.; Chen, Y.; Kholidy, H.; Blasch, E.; Ardiles-Cruz, E. A Survey of the Real-Time Metaverse: Challenges and Opportunities. Future Internet 2024, 16, 379. [Google Scholar] [CrossRef]
  22. Mystakidis, S. Metaverse. Encyclopedia 2022, 2, 486–497. [Google Scholar] [CrossRef]
  23. Braun, T.; Fung, B.C.M.; Iqbal, F.; Shah, B. Security and privacy challenges in smart cities. Sustain. Cities Soc. 2018, 39, 499–507. [Google Scholar] [CrossRef]
  24. Wang, Y.; Su, Z.; Zhang, N.; Xing, R.; Liu, D.; Luan, T.H.; Shen, X. A Survey on Metaverse: Fundamentals, Security, and Privacy. IEEE Commun. Surv. Tutor. 2023, 25, 319–352. [Google Scholar] [CrossRef]
  25. van Waveren, J.M.P. The asynchronous time warp for virtual reality on consumer hardware. In Proceedings of the 22nd ACM Conference on Virtual Reality Software and Technology, Munich, Germany, 2–4 November 2016; pp. 37–46. [Google Scholar] [CrossRef]
  26. Li, Y.; Zhang, Z.; Chen, H.; Ma, Z. Mamba: Bringing Multi-Dimensional ABR to WebRTC. arXiv 2023, arXiv:2308.03643. [Google Scholar] [CrossRef]
  27. Chen, Y.; Liu, E.S. A Path-Assisted Dead Reckoning Algorithm for Distributed Virtual Environments. In Proceedings of the 2015 IEEE/ACM 19th International Symposium on Distributed Simulation and Real Time Applications (DS-RT), Chengdu, China, 14–16 October 2015; pp. 108–111. [Google Scholar] [CrossRef]
  28. Chen, Y.; Liu, E.S. Comparing Dead Reckoning Algorithms for Distributed Car Simulations. In Proceedings of the 2018 ACM SIGSIM Conference on Principles of Advanced Discrete Simulation, New York, NY, USA, 23–25 May 2018; SIGSIM-PADS ’18. pp. 105–111. [Google Scholar] [CrossRef]
  29. Lin, K.C.; Wang, M.; Wang, J.; Schab, D.E. Smoothing a dead reckoning image in distributed interactive simulation. J. Aircr. 1996, 33, 450–452. [Google Scholar] [CrossRef]
  30. Antuley, U.; Hameed, S.; Siddiqui, S.; Shah, S.A. Securing Smart City Ecosystems: A Taxonomy-Based Review of Emerging Technologies and Frameworks for Scalable Collaborative Services. IET Smart Cities 2025, 7, 8–25. [Google Scholar] [CrossRef]
  31. Barcik, P.; Coufalikova, A.; Frantis, P.; Vavra, J. The Future Possibilities and Security Challenges of City Digitalization. Smart Cities 2022, 6, 137–155. [Google Scholar] [CrossRef]
  32. Monaco, D.; Sacco, A.; Spina, D.; Strada, F.; Bottino, A.; Cerquitelli, T.; Marchetto, G. Real-time latency prediction for cloud gaming applications. Comput. Netw. 2025, 264, 111235. [Google Scholar] [CrossRef]
  33. Peter, B.; Yazdanian, Y. Compensation for Latency in XR Offloaded Tasks Using Pose Prediction. Master’s Thesis, Lund University, Lund, Sweden, 2024. [Google Scholar]
  34. Dorokhov, I.; Ruponen, J.; Shutsky, R.; Nechesov, A. Time-Exact Multi-Blockchain Architectures for Trustworthy Multi-Agent Systems. In Proceedings of the MathAI 2025, San Diego, CA, USA, 6 December 2025; Available online: https://openreview.net/forum?id=2PLPmN5QW3 (accessed on 28 February 2026).
  35. Richter, F.; Zhang, Y.; Zhi, Y.; Orosco, R.K.; Yip, M.C. Augmented Reality Predictive Displays to Help Mitigate the Effects of Delayed Telesurgery. In Proceedings of the 2019 International Conference on Robotics and Automation (ICRA), Montreal, QC, Canada, 20–24 May 2019; pp. 444–450. [Google Scholar] [CrossRef]
  36. Nabiyouni, M.; Scerbo, S.; Bowman, D.A.; Höllerer, T. Relative Effects of Real-world and Virtual-World Latency on an Augmented Reality Training Task: An AR Simulation Experiment. Front. ICT 2017, 3, 34. [Google Scholar] [CrossRef]
  37. Gard, N.; Hilsmann, A.; Eisert, P. Combining Local and Global Pose Estimation for Precise Tracking of Similar Objects. In Proceedings of the 17th International Joint Conference on Computer Vision, Imaging and Computer Graphics Theory and Applications, Virtual Event, 6–8 February 2022; pp. 745–756. [Google Scholar] [CrossRef]
  38. Turek, T.; Pawełoszek, I. Digital Twins Technology for Smart City Development. A Case Study of Poland’s Largest Cities. Procedia Comput. Sci. 2024, 246, 4863–4872. [Google Scholar] [CrossRef]
  39. Yao, A.; Li, G.; Li, X.; Jiang, F.; Xu, J.; Liu, X. Differential privacy in edge computing-based smart city Applications:Security issues, solutions and future directions. Array 2023, 19, 100293. [Google Scholar] [CrossRef]
  40. Christopoulou, M.; Koufos, I.; Xilouris, G.; Dimitriou, N. 5G/6G Architecture Evolution for XR and Metaverse: Feasibility Study, Security, and Privacy Challenges for Smart Culture Applications. IEEE Access 2025, 13, 103077–103094. [Google Scholar] [CrossRef]
  41. Crespo-Aguado, M.; Lozano, R.; Hernandez-Gobertti, F.; Molner, N.; Gomez-Barquero, D. Flexible Hyper-Distributed IoT–Edge–Cloud Platform for Real-Time Digital Twin Applications on 6G-Intended Testbeds for Logistics and Industry. Future Internet 2024, 16, 431. [Google Scholar] [CrossRef]
  42. Abbas, A.E. Next-Generation Ethics: Engineering a Better Society; Google-Books-ID: sYK0DwAAQBAJ; Cambridge University Press: Cambridge, UK, 2019. [Google Scholar]
  43. Humphry, J. Policing Homelessness: Smart Cities and Algorithmic Governance. In Homelessness and Mobile Communication: Precariously Connected; Humphry, J., Ed.; Springer Nature: Cham, Switzerland, 2022; pp. 151–181. [Google Scholar] [CrossRef]
  44. Kogut, P.; van der Heijden, R. Digital Twins for Urban Governance: General Desires, Expectations, Challenges. In Decide Better: Open and Interoperable Local Digital Twins; Raes, L., Ruston McAleer, S., Croket, I., Kogut, P., Brynskov, M., Lefever, S., Eds.; Springer Nature: Cham, Switzerland, 2025; pp. 9–31. [Google Scholar] [CrossRef]
  45. Veloso-Luis, A.; Silva, A.; Neves-Silva, R. Socially Acceptable XR Technology for Community Engagement on Cultural and Urban Development. Procedia Comput. Sci. 2025, 270, 1641–1648. [Google Scholar] [CrossRef]
Figure 1. Event -driven XR architecture with explicit context selection, command handling, event sourcing, and read-side projections. Real-time signals bypass persistence via the synchronization layer, while durable commands are validated by aggregates and recorded in an append-only event store. Multiple projections derive authoritative state for XR rendering, collaboration, and agent reasoning. Red components are event generators, orange are event channels, yellow are event processors, and green are downstream activities.
Figure 1. Event -driven XR architecture with explicit context selection, command handling, event sourcing, and read-side projections. Real-time signals bypass persistence via the synchronization layer, while durable commands are validated by aggregates and recorded in an append-only event store. Multiple projections derive authoritative state for XR rendering, collaboration, and agent reasoning. Red components are event generators, orange are event channels, yellow are event processors, and green are downstream activities.
Asi 09 00061 g001
Table 1. Formal event taxonomy and temporal semantics in INTELLECTUM.
Table 1. Formal event taxonomy and temporal semantics in INTELLECTUM.
Event ClassTemporal SemanticsScopeFormal Constraints
Perceptual Events E P SynchronousClient/Session E E P : | E | = 1 E E S E E G
Interaction Events E I SynchronousUser/Agent E E I : E { e 1 , , e N } E ( E C E S )
Cognitive Events E C AsynchronousAgent Runtime e E C : s E S e is advisory
State Mutation Events E S OrderedGlobal/Shared E E S : t i T E i E i + 1
Physical Feedback Events E F AsynchronousEnvironment e E F : e E C e E P
Governance Events E G AsynchronousAudit/Trust E E S : g E G s . t . E g
Table 2. Design-time evaluation criteria references and reference constructs.
Table 2. Design-time evaluation criteria references and reference constructs.
DimensionCriterionDesign-Time Validation Reference
Pipeline
Coverage
Interaction classesCanonical interaction scenarios described in Section 4 and illustrated in Figure 1. Perception treated and covered in Algorithm 2 and Equation (4).
Event type completenessFormal event categories defined in Section 4.1 and Algorithm 1
Actor participationUnified actor abstraction for humans, agents, and robots defined in Section 3.3 and in Equation (11)
Architectural
Coherence
Responsibility separationLayered decomposition of XR client, synchronization layer, and runtime components (Context Selector, Command Handler/Aggregate, Event Store, Projections; Figure 1). Rendering pipeline decoupled from cognitive event loop in Algorithm 2.
Absence of cyclic dependenciesStructural analysis of event propagation graph in Section 4.1 induced by Figure 1, enforcing monotonic timestamping, role separation, and non-authoritative real-time bypass for directed acyclic event flow constraints.
Ingress/processing/egress boundariesEvent ingress and egress points defined in Algorithm 1
Fast vs. authoritative path separationVerification that real-time scene updates bypass persistence while authoritative events flow through Command Handler and Event Store to Projections (Figure 1). Reality switching (Equation (5)) is affecting only render state.
Event Semantics
Correctness
Event type disjointnessFormal event taxonomy defined in Section 4.3 and Table 1.
Routing determinismDeterministic event dispatch logic in Algorithm 1 and Equation (6).
Sync vs. async semanticsTemporal ordering and asynchronous handling formalized via Lamport-ordered event log (10) and actor-level event decoupling (12) in the synchronization layer.
Domain and persistence integrityValidation of invariant enforcement in Command Handler, append-only semantics in Event Store, and deterministic Projections ensuring correct read/write separation (Figure 1). Global scene state (Equation (9)) and interaction events form a time-ordered log are formalized (Equation (10)).
Scalability &
Extensibility
Agent count independenceEvent-driven decoupling (12) and actor abstraction (11)
Session scaling behaviorCoordination time complexity bounded by multi-entity decision-making time (Equation (14)). Proposition 2 states that coordination remains polynomial with bounded entity state size and fixed merge operators.
Agent closed-loop feedbackVerification that agent outputs ( e a i , e r a ) are correctly ingested into runtime without violating invariants or creating cycles, ensuring deterministic behavior under increasing agent numbers.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Nechesov, A.; Ruponen, J. INTELLECTUM: A Hybrid AR-VR Metaverse Framework for Smart Cities. Appl. Syst. Innov. 2026, 9, 61. https://doi.org/10.3390/asi9030061

AMA Style

Nechesov A, Ruponen J. INTELLECTUM: A Hybrid AR-VR Metaverse Framework for Smart Cities. Applied System Innovation. 2026; 9(3):61. https://doi.org/10.3390/asi9030061

Chicago/Turabian Style

Nechesov, Andrey, and Janne Ruponen. 2026. "INTELLECTUM: A Hybrid AR-VR Metaverse Framework for Smart Cities" Applied System Innovation 9, no. 3: 61. https://doi.org/10.3390/asi9030061

APA Style

Nechesov, A., & Ruponen, J. (2026). INTELLECTUM: A Hybrid AR-VR Metaverse Framework for Smart Cities. Applied System Innovation, 9(3), 61. https://doi.org/10.3390/asi9030061

Article Metrics

Back to TopTop