Abstract
Repeated dyadic interactions are characterized not only by average multimodal signal levels but also by how cross-modal consistency evolves over time. This study investigated whether the joint temporal organization of within-person face–body cross-modal consistency and probabilistic cross-modal affective agreement, quantified by the Consistency Index for Affective Synchrony (CIAS) and the Probabilistic Multimodal Consistency Index (PMCI), yielded reproducible session-level dynamic representations associated with repeated dyadic interaction. The primary analysis included 24 improvised and 17 naturalistic sessions, with 12 dyads in each condition. Time-varying signals from both partners were summarized through dynamic-state occupancy, dwell structure, transition characteristics, and temporal variability and evaluated using relational permutation analysis with shared-participant and technical controls. In naturalistic interaction, repeated sessions of the same dyad were significantly more similar ( = −2.0180, FDR q = 0.0114), with significant dyad-specificity beyond shared-participant similarity ( = −2.4670, q = 0.0033). Representation ablations showed that PMCI provided the dominant state-separation structure, whereas CIAS contributed modest and condition-dependent complementary information. Most notably, frozen participant-disjoint evaluation on 73 naturalistic sessions from six dyads and 11 previously unseen participants retained significant dyad-associated structure and achieved session-weighted nearest-dyad accuracy of 0.8356 and macro-dyad accuracy of 0.8512. Additional robustness analyses showed that the principal pattern persisted across state-count specifications, temporal-window settings, partner-order reversal, and explicit technical-quality adjustment. Overall, the findings support the existence of reproducible session-level multimodal dynamic structure associated with repeated dyadic interaction and its transfer to previously unseen participants, while not implying direct interpersonal synchrony or psychologically defined interaction states.
1. Introduction
Human affective behavior is expressed through multiple, concurrently operating modalities. Facial expressions, body posture, gestures, vocal cues, and movement dynamics constitute interrelated components of a multimodal expressive system. In many current multimodal emotion-recognition systems, these modalities are mainly used to improve downstream classification accuracy [1,2]. Accordingly, the conventional computational objective is to predict an emotion label from facial, bodily, or acoustic features. Although useful for classification, this formulation does not address whether the joint temporal organization of these modalities remains reproducible across repeated interactions of the same dyad.
In dyadic interaction, multimodal behavior is not defined only by the average intensity or consistency of individual signals. Two interaction sessions may have similar mean levels of face–body cross-modal consistency and probabilistic agreement, while differing substantially in their temporal organization. One dyad may consistently occupy a restricted region of the multimodal state space, whereas another may exhibit more frequent transitions, longer dwell periods, or greater variability. These temporal characteristics may provide information about the organization of interaction that is not captured by an emotion label or a single session-level mean.
This study builds on a broader line of research on multimodal consistency. The first component is the Unified Visual Synchrony framework and the Consistency Index for Affective Synchrony (CIAS), which quantify within-person temporal consistency between facial and bodily or gestural representations [3]. The second component is probabilistic cross-modal agreement: the Probabilistic Multimodal Consistency Index (PMCI) evaluates the compatibility of facial and gestural affect-related representations by comparing modality-specific probability distributions over shared affective prototypes [4]. The contribution of the present study does not stem from introducing CIAS or PMCI as new measures. Rather, the analysis extends these measures to a higher level of temporal organization by integrating their time-varying trajectories across both interaction partners into a session-level dynamic representation and evaluating whether this representation contains reproducible dyad-associated structure across repeated interactions.
CIAS and PMCI are treated as complementary time-varying observables rather than as components of a new scalar score. CIAS represents within-person temporal face–body consistency, whereas PMCI represents probabilistic agreement between modality-specific affect-related representations. At each time point, the measures from both interaction partners define a position in a joint dynamic state space. The resulting state sequence can then be summarized at the session level through occupancy, dwell duration, switching behavior, entropy, and signal variability. These features constitute a session-level dynamic representation that can be compared across repeated sessions. Representation-level ablations and same-dimensional shuffled/noise controls are used to determine the relative contributions of CIAS and PMCI and whether any improvement of the joint representation reflects signal content rather than merely increased dimensionality. The analytical continuity among CIAS, PMCI, and the present study is summarized in Table 1.
Table 1.
Analytical continuity across CIAS, PMCI, and the present dyadic-dynamics study.
This distinction is important because session-level averages do not fully capture the temporal organization of multimodal behavior. Two sessions may exhibit similar average CIAS and PMCI values while differing in the time spent in particular joint states, the duration of continuous state runs, the frequency of switching, or the entropy of the state sequence. The present study therefore shifts the analytical focus from isolated consistency scores toward the reproducibility of their joint temporal structure across repeated dyadic interactions.
The primary contribution of this study consists of four components. First, previously introduced CIAS and PMCI trajectories are integrated into a higher-level temporal representation of repeated interaction rather than combined into an additional scalar consistency index. Second, a session-level dynamic representation is constructed from joint-state occupancy, dwell duration, switching behavior, entropy, and signal variability. Third, reproducible dyad-associated structure is evaluated using relational permutation procedures, nearest-dyad identification, and MRQAP models that distinguish same-dyad similarity from shared-participant effects. Fourth, robustness is evaluated after removing absolute CIAS and PMCI levels, after pair-cross-fitted residualization of measured technical characteristics, and in participant-disjoint transfer data, with additional technical-quality and context controls, state-count sensitivity, partner-order reversal, and temporal-window sensitivity analyses. As a secondary exploratory extension, bidirectional cross-attention was additionally used to test whether learned face–body consistency and discrepancy signals provide information beyond the interpretable CIAS–PMCI representation.
Accordingly, the central question is whether the joint temporal organization of CIAS and PMCI provides reproducible dyad-associated information beyond shared-participant and measured technical structure, rather than whether these measures directly identify psychological states.
The remainder of the paper is organized as follows. Section 2 reviews multimodal affective analysis, face–body integration, interpersonal coordination, and dynamic approaches to dyadic interaction. Section 3 describes the Seamless Interaction samples, the CIAS and PMCI signals, construction of the joint dynamic state space, session-level dynamic representations, and relational statistical analyses. Section 4 presents the same-dyad, dyad-specificity, technical-control, and robustness results. Section 5 discusses the interpretation and limitations of the observed dyad-associated structure. Section 6 summarizes the study.
2. Related Work
Contemporary research conceptualizes human affective expression as a multimodal phenomenon rather than a property of a single isolated channel in which several expressive systems work together. These include the face, body, gestures, voice, gaze, posture, and movement dynamics. Classical approaches to emotional expression already emphasized that facial and bodily behavior are visible manifestations of affective states [5,6]. Later, categorical and dimensional theories expanded this understanding. In the first case, emotions are described as structured groups of basic affective categories, while in the second case, they are treated as continuous states that can be located along dimensions such as valence, arousal, and intensity [7,8,9]. Despite the differences between these approaches, one common assumption is important for computational modeling: human affective behavior is usually expressed not through a single channel, but through several partially independent, yet mutually interpretable, modalities.
This multimodal perspective is particularly important for interactive systems. In real communication, facial expression is almost never perceived separately from posture, gesture, gaze, speech rhythm, and the general context of the situation. For example, a smile may convey not only joy but also politeness, irony, embarrassment, or concealed tension. Its interpretation may therefore depend on whether it is accompanied by an open posture, relaxed gestures, a calm voice, or, conversely, bodily stiffness and avoidance of gaze. Similarly, an open posture and a broad gesture may strengthen the interpretation of positive engagement, whereas a closed posture or tense movement may weaken or even contradict the same facial signal. From this perspective, multimodal affective analysis should not only combine different modalities for emotion recognition but also evaluate to what extent these modalities really support one consistent affective interpretation.
Studies in affective neuroscience and perception provide strong empirical support for this approach. Emotional body language is processed by humans as a meaningful affective signal and can influence how facial expression is interpreted [10]. Moreover, evidence indicates rapid integration of facial and bodily emotional information. This shows that face–body consistency is not only the result of late cognitive evaluation, but also part of early affective perception [11]. Under conditions of ambiguity, high emotional intensity, or conflict between channels, body information may even become more decisive than facial expression, especially when emotional valence is evaluated [12]. In addition, other studies show a bidirectional contextual influence between the face and the body: facial expression affects the perception of the body, and bodily expression, in turn, changes the interpretation of the face [13]. Consequently, affective meaning under such conditions is relational, that is, it arises not from one separate modality, but from the general configuration of several expressive channels.
Such a relational approach is also consistent with studies showing that external facial expression does not always reliably accompany the experienced emotion [14]. Emotional expression can be suppressed, masked, regulated by social norms, or manifested only partially. Therefore, the absence of full agreement between modalities should not automatically be considered an error, noise, or a data artifact. In several cases, such inconsistency may indicate ambiguity, a transitional state, self-regulation, uncertainty, or a mismatch between the internal state and external manifestation. For interactive systems, this is important because an externally confident emotional label may be unreliable for an adaptive response if the channels on which it is based are weakly coordinated with each other.
Computational studies of bodily affective expression also emphasize the role of posture, movement, and kinematic characteristics. Reviews of affective body expression indicate that body movement carries emotionally meaningful information and can be used both in studies of perception and in automatic recognition [15]. Kinematic and postural features capture not only body position at a given moment but also its temporal dynamics, that is, movement direction, amplitude, speed, and local intensity of the gesture [16]. These results justify the use of pose-based affective analysis. In such an approach, keypoints of the face and body are considered not simply as geometric coordinates, but as structured temporal signals that reflect the dynamics of nonverbal behavior.
Multimodal affective analysis commonly employs early, late, and hybrid fusion strategies. In late fusion, the decisions of separate modal models are combined. Hybrid approaches combine the strengths of both levels [17,18]. Such methods are effective when the main task is to increase recognition accuracy, especially if one of the modalities is noisy, incomplete, or less informative. However, this leaves a distinct question unresolved: do such methods quantify the extent to which the modalities agree? A fusion classifier can output a correct or confident emotional label even when the underlying facial and bodily signals support different interpretations.
Deep multimodal models have extended this line through cross-modal learning, tensor fusion, attention mechanisms, and transformer architectures. Tensor fusion makes it possible to explicitly model unimodal, bimodal, and trimodal interactions in sentiment analysis and emotion analysis tasks [19]. Multimodal transformers, in turn, model interactions between unaligned sequences through cross-modal attention [20]. Cross-channel and prototype-grounded approaches have also been proposed to achieve more stable expression recognition and multimodal modeling of emotion [21]. These models are important because they improve representation learning on heterogeneous signals. Nevertheless, most of them remain primarily oriented toward prediction. Their main goal is usually to improve the accuracy of the final task, not to construct a separate, explicit, and interpretable measure of cross-modal consistency and its temporal organization.
Recent reviews of transformer-based multimodal emotion recognition similarly identify cross-modal attention, modality interaction, and fusion architecture as central directions in contemporary multimodal affective computing [1]. Recent multi-corpus work further emphasizes generalization across heterogeneous datasets and modality conditions, using cross-modal gated attention to integrate audio, visual, and textual representations [2].
In the psychological literature, interpersonal synchrony generally refers to temporal coordination between interacting individuals. Ramseyer and Tschacher operationalized nonverbal synchrony as coordinated body movement between patient and therapist, whereas Mayo and Gordon described interpersonal synchrony more broadly as the temporal coordination of behavioral and physiological processes within a dyad [22,23]. This definition differs from the within-person face–body consistency quantified by CIAS. Recent work further shows that temporal coordination can also occur within an individual across different communicative signals, such as gaze and gesture, a phenomenon described as intrapersonal synchrony [24]. Accordingly, CIAS is treated in the present study as a within-person cross-modal consistency signal rather than as direct evidence of interpersonal synchrony.
The distinction between prediction and consistency is central to the present study. Prediction concerns the affective interpretation assigned to an observation, whereas consistency concerns the compatibility and temporal organization of evidence across modalities [3,4,17,18]. Importantly, two sessions may have similar emotion labels or similar mean consistency values while differing in their state occupancy, dwell structure, switching frequency, and temporal variability. The present study therefore examines multimodal consistency as a dynamic property of repeated dyadic interaction rather than solely as a confidence signal for classification.
The first methodological component of this research line is the Unified Visual Synchrony framework and the Consistency Index for Affective Synchrony. In its original formulation, UVS/CIAS describes facial and gestural behavior as a coordinated visual system and uses the term synchrony for a time-resolved measure of within-person face–body consistency [3]. In this formulation, CIAS evaluates the similarity and short-term temporal stability of the latent trajectories of the face and body/gestures. Rather than serving solely as an emotion classifier, CIAS quantitatively assesses whether visual consistency between facial and gestural behavior is preserved over time. Accordingly, in the present study, CIAS is treated as a within-person cross-modal consistency signal rather than as a direct measure of coupling between interaction partners.
The second methodological component is the Probabilistic Multimodal Consistency Index. PMCI formulates probabilistic cross-modal consistency as an agreement between probability distributions over shared affective prototypes [4]. Instead of directly forcing facial and bodily embedding representations to coincide, PMCI allows each modality to form its own probability distribution in a shared affective prototype-space. These distributions are then compared through Jensen–Shannon agreement. This distinction is relevant because direct embedding similarity may capture not only emotional content, but also actor identity, posture style, background, recording conditions, or other shortcut features. PMCI is therefore better regarded as a diagnostic consistency index rather than as an emotion classifier or a general-purpose pair-matching model.
Previous studies evaluated CIAS and PMCI using matched/mismatched comparisons, perturbation, shuffle, and robustness controls [3,4]. These analyses provide computational validation of the measures’ sensitivity to manipulated cross-modal consistency, but they do not establish psychological construct or criterion validity. Accordingly, CIAS and PMCI are treated here as operational computational measures of cross-modal consistency rather than as validated measures of psychological states.
The term “affective” is retained because the underlying modality-specific representations and the PMCI prototype space are affect-related; it describes the semantic provenance of the computational representation rather than an established psychological measurement property. External validation against independent behavioral annotations, self-report measures, or experimentally controlled affective criteria remains to be established.
Recent computational research increasingly treats dyadic interaction as a temporally evolving system rather than as two independent individual observations. Dynamic approaches distinguish synchrony from other forms of interpersonal coordination, including complementarity, and show that coordination patterns can depend on task constraints and interaction context [25]. Contemporary multimodal dyadic modeling also preserves information from both interactants: recent studies have examined dyadic affect using participant-specific multimodal signals and have developed multi-perspective representations that explicitly model the affective dynamics of both interaction partners [26,27].
The present paper extends CIAS and PMCI to a higher session-level representation of repeated dyadic interactions. CIAS provides a time-varying measure of within-person face–body cross-modal consistency [3], whereas PMCI provides a time-varying measure of probabilistic agreement between modality-specific affective distributions [4]. Their joint trajectories across both interaction partners are treated as a multivariate temporal representation from which session-level dynamic features can be derived. This representation makes it possible to examine whether the temporal organization of CIAS and PMCI contains reproducible dyad-associated structure across repeated interaction sessions.
Dyadic-dynamics analysis requires datasets that extend beyond isolated emotion-labeled clips. They must preserve the identities of both interaction partners, contain repeated sessions of the same dyad and provide temporally aligned multimodal information. The Seamless Interaction dataset satisfies these requirements by providing repeated, improvised and naturalistic interactions, together with participant and session identifiers [28]. Its structure enables differentiation among sessions from the same dyad, sessions sharing one participant, and sessions involving entirely different participants.
Recent multimodal dyadic corpora reinforce the importance of preserving partner-specific and temporally structured nonverbal information. For example, the NoXi+J corpus combines speech acoustics, facial expressions, backchannel behavior, and gestures in multilingual dyadic conversations and demonstrates systematic relationships between nonverbal interaction patterns, cultural context, and engagement [29].
Importantly, interpersonal synchrony is conceptually distinct from within-person multimodal consistency and is typically investigated through temporal relationships between the two members of a dyad. A recent systematic review and meta-analysis distinguishes behavioral, physiological, and neural forms of interpersonal synchrony and examines their relationships across human dyads [30]. Direct behavioral-coordination studies likewise quantify temporal dependence between interlocutors explicitly; for example, cross-recurrence quantification analysis has been applied to head and body movements of both partners during dyadic conversation [31]. In contrast, CIAS in the present study operates within each participant’s facial and bodily streams. The primary finding is therefore interpreted as a reproducible joint profile of the two partners’ multimodal dynamics rather than as direct evidence of interpersonal synchrony.
Relative to raw RGB frames, pose-based representations may reduce sensitivity to background, clothing, lighting, image texture, and some appearance-related shortcut features. Therefore, such representations are useful for experiments in which the goal is not appearance-based matching but analysis of the relational structure between modalities.
The statistical analysis of session similarity requires methods that account for dependence among pairwise relational observations. The Quadratic Assignment Procedure provides a permutation-based method for analyzing matrices of dyadic relations without treating matrix entries as independent observations [32]. Its multiple-regression extension, MRQAP, supports the simultaneous examination of several relational predictors, including same-dyad identity, shared-participant structure, and technical similarity [33]. These methods were therefore used to determine whether similarity between session-level dynamic representations was associated with repeated dyad membership after accounting for shared-participant and measured technical structure.
Overall, previous work supports treating multimodal consistency as an explicit analytical object rather than merely as an internal property of a fusion model [3,4,17,18]. However, less is known about how multiple time-varying consistency signals can be integrated into a higher-level representation of repeated interaction and whether this representation contains reproducible dyad-associated structure. The present study addresses this gap by examining whether the temporal organization of CIAS and PMCI can be transformed into a reproducible session-level dynamic representation across repeated interactions.
3. Materials and Methods
3.1. Study Design and Research Questions
The present study examined whether repeated interaction sessions involving the same two participants exhibit a reproducible organization of multimodal affective dynamics. Rather than treating each temporal observation as an independent classification instance, the analysis considered the complete dyadic session as the primary analytical unit. Each session was represented through the temporal trajectories of two complementary measures: CIAS, which describes within-person temporal face–body cross-modal consistency, and PMCI, which describes probabilistic agreement between facial and bodily affect-related representations [3,4].
For dyad and session s, the complete analytical transformation can be expressed as:
where denotes the original dyadic session, denotes the joint time-varying observation constructed from partners A and B, denotes the resulting sequence of joint dynamic states, and denotes the session-level dynamic representation.
Four primary research questions were examined:
- Are repeated sessions from the same dyad more similar than sessions involving different dyads?
- Does same-dyad similarity persist after excluding absolute CIAS and PMCI levels?
- Does dyad-specific similarity remain after accounting for shared participants and measured technical characteristics?
- Does CIAS provide complementary dyad-specific information beyond PMCI alone, and can any observed improvement be distinguished from the effect of adding extra representation dimensions?
The analyses were not preregistered; “primary” and “sensitivity” describe the analytical hierarchy, whereas “frozen” indicates transfer without parameter re-estimation. The analysis therefore distinguished three relational cases: sessions involving the same dyad, sessions from different dyads that shared one participant, and sessions involving entirely different participants. This distinction was required because similarity between sessions could reflect either the complete interaction pair or stable participant-associated structure carried by one repeatedly observed individual. This design follows a dynamic dyadic systems perspective, according to which interaction is represented as an evolving structure jointly produced by both partners rather than as a collection of independent individual observations [34,35]. Because pairwise distances between sessions share observations and are therefore statistically dependent, relational permutation procedures were required rather than conventional regression methods that assume independent observations [32,33].
For clarity, the pair-cross-fitted dynamics-without-levels MRQAP analyses under the 2 s and 3 s minimum-dwell specifications constituted the principal relational analyses, whereas the remaining ablation, transfer, and robustness analyses were used to assess sensitivity, generalization, or model dependence. The overall analytical workflow is summarized in Figure 1.
Figure 1.
Overview of the joint CIAS–PMCI dynamic representation pipeline, from time-varying CIAS and PMCI trajectories to state segmentation, session-level dynamic representations, and relational validation.
3.2. Dataset and Sample Construction
The analysis used repeated improvised and naturalistic interactions from the Seamless Interaction dataset [28]. The Seamless Interaction Dataset was introduced as a large-scale resource for modeling dyadic audiovisual behavior and contains face-to-face interactions collected across multiple conversational contexts [28]. Improvised interactions were scenario-guided, whereas naturalistic interactions used prompted conversations [28]. Prompt identity was included as a technical control. Coarse source relationship-category metadata were available and were matched to all 41 discovery interactions. All 24 improvised interactions were labeled as involving strangers, whereas the naturalistic sample comprised 14 familiar interactions (12 friends and 2 family-generic) and 3 stranger interactions.
These relationship categories were examined descriptively but were not included as inferential predictors; longitudinal relationship history beyond these source categories was not available. Repeated observations of the same dyad corresponded to distinct interaction identifiers within the same source recording-session identifier. The available time variable was local to each interaction and reset across interaction identifiers; therefore, exact elapsed intervals between repeated observations could not be reconstructed. The present study did not use the complete corpus; it used only sessions satisfying the specified requirements for participant identity, dyadic completeness, repeated-pair structure, and availability of the multimodal variables required for CIAS and PMCI computation. The dataset provides participant-level multimodal feature files together with participant, interaction, and session identifiers. These identifiers enabled reconstruction of complete dyadic sessions and repeated interactions involving the same pair of participants.
A session was retained only when usable multimodal files were available for both interaction partners. Individual participant files were grouped by interaction identifier, and a session was considered complete when exactly two distinct participants were represented. The pair identifier was defined independently of participant order:
For the discovery analyses, each participant file contributed the 24-dimensional facial action-unit vector (movement_v4:FAUValue), the x–y coordinates of 133 facial, body, and hand keypoints (boxes_and_keypoints:keypoints), and the available eight-dimensional affective score vector (movement_v4:emotion_scores). The keypoint confidence coordinate was excluded, and the remaining coordinates were centered using the frame-wise median body position and normalized by the median keypoint distance from that center. Within each participant file, the facial, keypoint, affective, and validity arrays were aligned by truncating them to the shortest available frame sequence. Participant-level observations were calculated in 2.0 s windows at an assumed rate of 30 frames/s, with a 0.5 s stride. The two participant streams were subsequently aligned through an inner join on the exact session identifier and window start time (time_sec), so windows without a corresponding observation from both partners were removed. Windows were discarded when the proportion of valid frames, derived from movement_v4:is_valid, was non-finite or below 0.70. Remaining non-finite facial and bodily feature values were replaced with zero before CIAS and PMCI calculation. A session was retained only when exactly two usable participant files representing exactly two distinct participants were available; a participant file was considered usable when facial features, bodily keypoints, and either affective scores or both valence and arousal variables were present.
The discovery sample contained 24 improvised sessions from 12 dyads involving 17 participants. The naturalistic discovery sample contained 17 sessions from 12 dyads involving 21 participants. The improvised and naturalistic conditions shared no participants, dyads, or sessions.
With respect to repeated-session information, the improvised discovery sample contained nine repeated dyads contributing 21 sessions and three singleton dyads, whereas the naturalistic discovery sample contained four repeated dyads contributing nine sessions and eight singleton dyads. Singleton dyads contributed to between-dyad comparisons but could not contribute same-dyad repeated-session pairs. Session counts per dyad were 3, 3, 3, 2, 2, 2, 2, 2, 2, 1, 1, and 1 in the improvised sample, and 3, 2, 2, 2, 1, 1, 1, 1, 1, 1, 1, and 1 in the naturalistic sample.
A participant-disjoint naturalistic sample contained 73 sessions from six repeated dyads involving 11 participants who did not occur in the discovery samples. One participant appeared in two dyads; consequently, the six dyads formed five participant-connected components. This sample was reserved for frozen participant-disjoint transfer and was not used to estimate the original state model or the cross-attention architecture. For exact reproduction of the historical participant-disjoint robustness analysis, the corresponding legacy Seamless Interaction WebDataset fields (movement:FAUValue, movement:emotion_scores, and movement:is_valid) were used together with the same keypoint representation. The frozen CIAS–PMCI pipeline reproduced all 73 previously selected naturalistic sessions from the six dyads before application of the learned C/D extension. No state-model or architecture parameters were re-estimated on this sample. The composition of the discovery and participant-disjoint samples is summarized in Table 2.
Table 2.
Composition of the discovery and final analytical samples used in the primary and robustness analyses.
3.3. CIAS Temporal Cross-Modal Consistency Signal
The temporal component was based on the Consistency Index for Affective Synchrony introduced in the previous UVS/CIAS study [3]. Let
denote the facial feature sequence and
denote the corresponding bodily or gestural feature sequence. Separate modality encoders map these sequences into a common latent representation:
CIAS combines cross-modal latent similarity with a dynamic stability component:
where denotes cosine similarity. The local dynamic component is defined as
where is a short temporal offset and denotes the local latent trajectory. Thus, high CIAS values require both cross-modal similarity and short-term temporal stability.
CIAS was calculated as a time-varying signal over consecutive observations. In the present study, it was not treated as an emotion classifier or as a psychological state indicator. It formed the within-person temporal cross-modal consistency dimension of the joint dyadic representation. The published UVS/CIAS formulation describes CIAS as a continuous time-resolved synchrony measure constructed through a shared latent representation. In the present dyadic analysis, however, this signal is interpreted operationally as within-person face–body consistency and not as a direct measure of temporal coupling between the two interaction partners.
3.4. PMCI Probabilistic Agreement Signal
The probabilistic component followed the PMCI formulation [4]. Let
denote a shared set of M affective prototypes. Facial and bodily embeddings were independently projected onto this prototype space. For the facial stream, the probability assigned to prototype m was
and for the bodily stream,
where is a temperature parameter controlling the concentration of the distributions. In the present analysis, the PMCI model used 16 affective prototypes and a softmax temperature of = 0.25. To distinguish the number of affective prototypes from the number of dynamic states introduced below, the prototype count is denoted by M = 16, whereas K = 3 is reserved for the dynamic-state model.
The PMCI score was defined as
where denotes Jensen–Shannon divergence [36]. Higher PMCI values indicate greater agreement between modality-specific prototype distributions.
PMCI does not assign an emotion category and does not replace an exact pair-matching model. A low PMCI value only indicates reduced agreement between the two probability distributions. It does not identify whether that reduction was caused by behavioral variability, ambiguity, occlusion, data quality, or another psychological or technical factor.
3.5. Joint CIAS–PMCI Dynamic State Space
For each temporally aligned observation t, the dyadic interaction was represented by a four-dimensional vector containing CIAS and PMCI values for both partners:
Before state estimation, each dimension was standardized using the mean and standard deviation estimated from the discovery data:
A three-state clustering solution was applied to the standardized observations. The state label was assigned according to the nearest centroid:
where is the centroid of state k. The scaler and centroid definitions were held fixed in the primary downstream analyses.
The joint state space was partitioned using k-means clustering with K = 3, k-means++ initialization, n_init = 20, random_state = 42, and a maximum of 500 iterations. The three-state solution was used as a fixed low-complexity descriptive partition of the joint signal space rather than as a claim that three represents the uniquely correct number of latent interaction states. Importantly, K = 3 was not selected to optimize downstream same-dyad coefficients, dyad-specificity, identification accuracy, or statistical significance. The scaler and state centroids estimated from the discovery data were subsequently held fixed across interaction conditions, minimum-dwell constraints, representation specifications, and technical-control analyses.
Because the CIAS-only, PMCI-only, and joint representations differ in dimensionality and geometry, each representation used its own three-state model for the ablation analyses described in Section 3.9; within each representation, the corresponding state model was fixed before downstream dyadic analysis. Sensitivity analyses evaluated whether the findings depended on these modeling choices. State-count sensitivity analyses were conducted for K = 2, 3, 4, and 5 under both 2 s and 3 s minimum-dwell specifications using the same downstream relational evaluation protocol. The stability of the K = 3 solution was assessed across ten random initializations using the pairwise adjusted Rand index (ARI). Additional robustness analyses tested invariance to a global reversal of partner order and repeated the temporal pipeline using 1.5, 2.0, and 3.0 s analysis windows.
For descriptive interpretation, two additional quantities were calculated. Joint consistency was defined as the mean of the four CIAS–PMCI coordinates:
whereas cross-partner PMCI asymmetry was defined as
These quantities were not used to assign the dynamic state labels. They were used for descriptive interpretation of the state centroids and were subsequently summarized at the session level as derived representation variables. State labels remained neutral because the dataset did not contain psychological annotations supporting terms such as conflict, regulation, repair, or ambiguity.
The centroids show that CIAS values varied only slightly across the three states, whereas PMCI produced the principal separation.
3.6. State Segmentation and Minimum-Dwell Constraint
The initial state assignments may contain short isolated changes resulting from local noise or small variations around a centroid boundary. To obtain temporally interpretable state runs, consecutive identical labels were grouped into contiguous segments.
For run r, its duration was defined as
where tr denotes the inferred temporal sampling interval for run r, with a nominal value of 0.5 s when required.
Real discontinuities in the observation timeline were treated as segment boundaries and were not counted as transitions. Runs shorter than a specified minimum duration were reassigned according to the surrounding temporal state structure:
where denotes the minimum-dwell smoothing operator.
The minimum-dwell procedure was implemented as a deterministic temporal post-processing rule rather than as a separately trained sequence model. Short runs were merged according to the immediately adjacent state configuration, while discontinuities in the original time axis were preserved as segment boundaries. Thus, a gap between two observations could not be interpreted as an observed state transition. The procedure reduced isolated boundary fluctuations while preserving longer periods of stable state occupancy.
Two minimum-dwell thresholds, 2 s and 3 s, were evaluated as parallel sensitivity specifications. Neither threshold was selected on the basis of the resulting same-dyad coefficients, dyad-specificity, identification accuracy, or statistical significance. The two specifications were analyzed in parallel, and their agreement was used to assess whether the findings depended on a particular temporal-smoothing parameter.
3.7. Session-Level Dynamic Representation
Each smoothed session sequence was converted into a session-level dynamic representation designed to capture four complementary aspects of temporal organization: state usage, transition topology, absolute signal levels, and within-session signal dynamics. The complete representation contained 32 features. The primary dynamics-without-levels representation contained 26 features and excluded the six absolute-level variables, allowing dyad-associated temporal structure to be tested without relying on persistent differences in mean CIAS or PMCI levels. For state k, occupancy was defined as:
where Lr denotes the duration of run r, and qr is its assigned state. Mean and median dwell durations were calculated from the set of runs assigned to state k:
where Rk denotes the number of contiguous runs assigned to state k.
With three states, occupancy, mean dwell and median dwell produced nine state-usage features.
Transition-topology features. For each pair of states and j, transition probabilities were calculated as
where is the number of observed transitions from state to state j.
The switching rate per minute was
where is session duration in seconds. Normalized transition entropy was calculated from the empirical transition distribution:
For sessions with fewer than two observed transition types, H_trans was set to 0.
The transition-topology group contained nine features, including transition probabilities, switching rate, transition entropy, and return probability.
Absolute signal levels. Absolute-level features were session means of the available CIAS–PMCI-derived signals:
This group contained six features.
Signal dynamics. Temporal variability was represented through session-level standard deviations:
Time-aware joint velocity was calculated from consecutive joint vectors:
and summarized at the session level by
The signal-dynamics group contained eight features.
The resulting representations were:
The dynamics-without-levels representation was
whereas the complete representation was
For K = 3, the state-usage group comprised three state-occupancy proportions, three state-specific mean dwell durations, and three state-specific median dwell durations. The transition-topology group comprised the six possible directed off-diagonal transition probabilities, switching rate per minute, normalized transition entropy, and return probability. Return probability was defined as the proportion of three-run sequences of the form i→j→i, where ij. The six absolute-level features were the session means of CIAS for partners A and B, PMCI for partners A and B, joint consistency, and cross-partner PMCI asymmetry. The eight signal-dynamics features comprised the session standard deviations of these six signals, the standard deviation of time-aware joint velocity, and the mean time-aware joint velocity. Consequently, the dynamics-without-levels representation contained 26 features, whereas the full representation contained all 32 features.
3.8. Session Similarity and Statistical Analysis
Session representations were analyzed separately for each interaction condition, minimum-dwell constraint, and feature representation. Before calculating session similarity, representation variables were median-imputed when necessary, zero-variance features were removed, and the remaining variables were standardized:
where is the feature r of session i, and and are the corresponding sample mean and standard deviation. Pairwise session dissimilarity was represented using Euclidean distance:
Smaller values indicated more similar session-level dynamic representations.
Three relational matrices were constructed. The same-dyad matrix identified sessions involving the same two participants. The shared-participant matrix identified sessions from different dyads containing one common participant.
A technical-distance matrix represented differences in measured recording and data-quality characteristics. The technical variables comprised session duration, the number of retained temporal observations, the number of continuous temporal segments, partner-specific means and standard deviations of valid-frame fraction, and the session mean and standard deviation of nearest-state distance when available. Vendor, recording-session identifier, and interaction-prompt identifier were additionally included as categorical technical variables when they contained between 2 and 30 observed levels. Numeric technical variables were median-imputed and standardized, whereas categorical variables were one-hot encoded before construction of the technical-distance matrix.
Because entries in a session-distance matrix are relationally dependent, statistical significance was evaluated using the Quadratic Assignment Procedure and Multiple Regression Quadratic Assignment Procedure [32,33]. The principal model was
where denotes same-dyad identity, denotes a shared participant, and denotes technical distance. A negative indicated that sessions from the same dyad were closer than other sessions.
Dyad-specificity was evaluated using the contrast
A negative value indicated that sessions from the same complete dyad were closer than sessions that only shared one participant.
As an additional technical-control analysis, measured technical effects were removed from the representation variables using ridge regression with pair-group cross-fitting. Ridge regression was used to stabilize the technical-adjustment model in the presence of correlated technical predictors [37]. The regularization parameter was fixed at = 10. Cross-fitting was performed at the dyad level using a leave-one-dyad-out procedure: all sessions belonging to the evaluated dyad were excluded from model fitting, and the fitted technical model was then used to obtain residuals for the held-out dyad. This procedure was repeated until residualized representation variables were obtained for all dyads. The residualized variables were standardized, Euclidean session-distance matrices were reconstructed, and MRQAP and nearest-dyad analyses were repeated using the same relational definitions as in the primary analysis.
Nearest-dyad identification was used as a secondary validation measure. For each session, the closest other session was identified from the distance matrix, and the prediction was considered correct when both sessions belonged to the same dyad:
Permutation probabilities for the MRQAP and nearest-dyad analyses were estimated using 5000 permutations. Relational permutation testing preserved the matrix structure through simultaneous row-and-column permutations. For the primary representation analyses, Benjamini–Hochberg correction was applied separately within each combination of interaction condition, minimum-dwell constraint, and technical-adjustment specification across the defined representation feature groups. The resulting FDR-adjusted permutation probabilities are reported as q-values [38].
In practical terms, a negative same-dyad coefficient indicates that repeated sessions of the same pair are more similar than sessions involving different pairs. The dyad-specificity contrast further tests whether this similarity exceeds the similarity expected merely from sharing one participant.
To assess dyad-level influence on the principal MRQAP estimates, a leave-one-dyad-out sensitivity analysis was performed. In each iteration, one dyad was removed and the complete technical-residualization and MRQAP effect-estimation procedure was recomputed; the resulting ranges of and were used as sensitivity summaries.
3.9. Representation Ablation and Same-Dimensional Controls
To determine whether the dyad-associated structure was driven primarily by one of the two indices, the complete analytical pipeline was repeated using three alternative observation spaces: CIAS-only , PMCI-only , and the joint CIAS–PMCI representation . State estimation was performed separately for each representation because the three observation spaces differed in dimensionality and geometry. In the improvised condition, three-state labels were estimated under pair-group cross-fitting. Naturalistic observations were subsequently projected into the corresponding representation-specific model fitted on the complete improvised sample. The remaining analysis was identical across representations and included minimum-dwell constraints of 2 and 3 s, run and transition extraction, session-level representation construction, pair-group cross-fitted technical residualization, MRQAP analysis with shared-participant adjustment, and nearest-dyad identification.
Two representation sets were evaluated. The dynamics-without-levels set excluded absolute session means of CIAS and PMCI and retained state occupancy, dwell, transition, variability, asymmetry, and velocity-related features. The full representation additionally included the available representation-specific signal levels. Only features applicable to the active representation were included. Because the joint representation produced slightly more representation dimensions than the single-index representations, MRQAP coefficients used for direct cross-representation comparisons were normalized by the square root of the effective number of representation features:
where denotes the number of features entering the corresponding distance matrix. More negative normalized coefficients indicated stronger same-dyad similarity or dyad-specificity. Direct added-value tests compared the joint and PMCI-only distance structures through permutation-based differences in normalized same-dyad coefficients, normalized dyad-specificity coefficients, and nearest-dyad accuracy.
Two additional same-dimensional controls were used to test whether any improvement in the joint representation could be explained merely by adding two CIAS-related dimensions. In the shuffled-CIAS control, the original PMCI trajectories were retained, whereas each CIAS trajectory was replaced by a real trajectory sampled from a session belonging to another dyad within the same interaction condition. Donor trajectories were resampled to the duration of the recipient session, and the assignment of the two interaction partners was randomly reversed with probability 0.5. This procedure preserved realistic CIAS ranges and temporal structure while disrupting their association with the target dyad.
In the matched-noise control, CIAS was replaced by bivariate autoregressive noise. The generated signals were approximately matched to the condition-specific CIAS means, standard deviations, lag-one autocorrelations, and cross-partner correlation and were restricted to the original range. PMCI trajectories remained unchanged. For each control type, 199 empirical realizations were generated and evaluated individually rather than averaged into a single control distance matrix. Statistical comparisons were based on the resulting empirical control distributions, with multiplicity correction applied across the corresponding comparison family. Statistical significance was evaluated using 5000 permutations, and Benjamini–Hochberg correction was applied within each corresponding comparison family. The shuffled-CIAS analysis served as the primary test of dyad-linked CIAS information, whereas matched noise provided an additional control for dimensionality and temporal smoothness.
3.10. Exploratory Bidirectional Cross-Attention Consistency–Discrepancy Extension
To test whether a learned cross-modal interaction mechanism could provide information beyond the interpretable CIAS–PMCI representation, an auxiliary bidirectional cross-attention module was evaluated. This extension was not used to replace the primary CIAS–PMCI state space and did not alter the frozen primary analysis. Instead, it generated additional consistency and discrepancy trajectories that were evaluated as complementary session-level features.
Facial and bodily sequences were first projected independently into a 64-dimensional latent space using linear projection, layer normalization, and GELU activation. Bidirectional cross-attention was then implemented using four-head multi-head attention with a dropout rate of 0.10, following the general cross-modal attention principle used in multimodal transformer architectures [20]. A single cross-attention block was used; no stacked Transformer layers were employed. The face-to-body branch was defined as
whereas the body-to-face branch was defined as
Residual connections and layer normalization were applied to both attended representations, followed by temporal mean pooling to obtain the latent vectors. Two separate heads were then used. The consistency head operated on the concatenation of the two pooled representations and their element-wise interaction,
whereas the discrepancy head operated on their absolute difference,
The consistency head used a 192→64→1 multilayer projection with GELU activation and dropout of 0.10, whereas the discrepancy head used a 64→64→1 projection with the same activation and dropout; sigmoid activation converted the resulting logits into consistency and discrepancy scores.
Training and architecture-level evaluation used a dyad-disjoint split of the improvised discovery data, with 30% of dyads assigned to the held-out test partition. The learned consistency and discrepancy quantities are operational cross-modal correspondence scores defined by the matched-versus-mismatched training objective. They should not be interpreted as direct measures of psychological agreement, conflict, emotional compatibility, or interpersonal regulation. The model was trained for 10 epochs using AdamW with a learning rate of 7 * 10−4, a weight decay of 10−4, and equal weighting of two binary-cross-entropy objectives: consistency was trained against matched face–body pairs, whereas discrepancy was trained against the complementary mismatch target. Four architectural configurations were evaluated: no cross-attention, face-to-body attention, body-to-face attention, and bidirectional cross-attention. The bidirectional configuration was retained for the downstream extension. Attention weights were not interpreted as behavioral importance scores.
For session-level analysis, five dynamics features were calculated separately from the consistency and discrepancy trajectories: the standard deviation for partner A, the standard deviation for partner B, the standard deviation of the cross-partner gap, mean time-aware velocity, and the standard deviation of time-aware velocity. Thus, ten learned dynamics features were added to the 26-feature dynamics-without-levels CIAS–PMCI representation, producing a 36-feature extended representation. Absolute means of the learned consistency and discrepancy signals were not included in the primary architecture-extension analysis. The same pair-cross-fitted technical residualization, MRQAP procedure, nearest-dyad analysis, 5000 permutations, and FDR correction were then applied. A 20-repeat session-shuffled consistency/discrepancy control preserved dimensionality while disrupting the association of the learned features with the target session.
Facial and bodily standardization parameters were estimated from the training partition only and subsequently applied to the held-out data. Balanced matched and mismatched face–body examples were generated separately within the training and test partitions: matched examples preserved the original face–body pairing, whereas mismatched examples were generated by randomly permuting bodily windows while preventing a window from being paired with itself.
Architecture-level robustness was assessed through three additional analyses. First, the bidirectional model was evaluated across five matched training seeds. Second, no-attention, face-to-body, body-to-face, and bidirectional variants were compared using exact dyad-block randomization on the four held-out test dyads. Third, consistency-only, discrepancy-only, and combined consistency–discrepancy readouts were evaluated separately. The downstream contribution of the learned consistency/discrepancy features was further evaluated against a 20-repeat session-shuffled C/D control that preserved representation dimensionality while disrupting the association between learned features and session identity.
4. Results
4.1. Joint Dynamic States and Session-Level Representations
The three-state solution was characterized by limited variation in CIAS and substantially greater variation in PMCI. CIAS values ranged from approximately 0.523 to 0.528 across the state centroids, whereas the principal state separation reflected differences between the partners’ PMCI values. State 0 showed high PMCI for both partners, while States 1 and 2 showed reduced PMCI for one partner and higher PMCI for the other. The states were therefore interpreted as descriptive regions of the joint signal space rather than as psychological categories.
The session-level analysis produced four feature groups: nine state-usage features, nine transition-topology features, six absolute-level features, and eight signal-dynamics features. The dynamics-without-levels representation contained 26 features, whereas the complete representation contained 32 features.
As shown in Table 3 and Figure 2A, the PMCI coordinates clearly separated the three centroids. State 0 was characterized by high PMCI for both partners, whereas States 1 and 2 showed complementary cross-partner asymmetry. In contrast, Figure 2B shows that all CIAS coordinates remained within a narrow interval of approximately 0.523–0.528. Thus, the estimated state structure was jointly represented by CIAS and PMCI but was predominantly separated along the PMCI dimensions.
Table 3.
State centroids in the original CIAS–PMCI units.
Figure 2.
Data-driven centroids of the joint CIAS–PMCI state space. (A) PMCI coordinates provide the principal separation among the three states. (B) CIAS coordinates occupy a substantially narrower range.
Table 4 summarizes the resulting representation composition. The full representation contained 32 variables, whereas the dynamics-without-levels representation retained 26 variables after excluding the six absolute CIAS- and PMCI-level features. This distinction was used in Section 4.3 to determine whether the observed dyad-associated structure reflected temporal organization rather than mean signal levels. State-count sensitivity analyses showed that the principal same-dyad pattern was not restricted to the K = 3 partition. Across K = 2–5, same-dyad coefficients remained negative across all tested interaction-condition and minimum-dwell combinations in the primary dynamics-without-levels analysis, although their magnitude and FDR-adjusted significance varied across K. Pair-held-out silhouette values did not identify K = 3 as an optimal state count (mean silhouette: 0.2775 for K = 2, 0.2635 for K = 3, 0.2834 for K = 4, and 0.2935 for K = 5), supporting its interpretation as a fixed low-complexity descriptive partition rather than a uniquely optimal clustering solution. The K = 3 partition was essentially invariant to random initialization (mean pairwise ARI = 0.9989; minimum = 0.9974). No inconsistencies were observed across the 24 audited comparisons following a global reversal of partner order. Results were also directionally and statistically consistent across 1.5, 2.0, and 3.0 s windows.
Table 4.
Session-level dynamic representation feature groups.
4.2. Same-Dyad Similarity and Dyad-Specificity
In the naturalistic condition, pair-cross-fitted dynamics-without-levels representations showed significant same-dyad similarity and dyad-specificity (, FDR ). The dyad-specificity contrast was also significant (, FDR ), indicating that the observed similarity was stronger than the similarity associated with sharing only one participant. Nearest-dyad identification accuracy was 0.2353 and exceeded its permutation null (FDR ).
The result remained significant under the 3 s constraint. The same-dyad coefficient was (FDR ), the dyad-specificity contrast was (FDR ), and nearest-dyad identification accuracy was 0.3529 (FDR ).
In the improvised condition, same-dyad similarity was significant under both 2 s (, FDR ) and 3 s (, FDR ) specifications. However, dyad-specificity did not reach FDR-adjusted significance (, FDR , and , FDR , respectively).
The complete numerical results are reported in Table 5. The same-dyad and dyad-specificity estimates are visualized in Figure 3, whereas nearest-dyad identification accuracy is summarized in Figure 4.
Table 5.
Primary same-dyad similarity, dyad-specificity, and nearest-dyad identification results.
Figure 3.
Pair-cross-fitted MRQAP estimates for same-dyad association and dyad-specificity across conditions and dwell constraints. (A) Same-dyad coefficient . (B) Dyad-specificity contrast . Line segments connect the estimates to zero for visual reference and do not represent confidence intervals.
Figure 4.
Nearest-dyad identification accuracy across conditions and dwell constraints.
Leave-one-dyad-out sensitivity retained the negative direction of both and under every single-dyad omission. The ranges were −2.0782 to −1.2241 (improvised, 2 s), −1.8630 to −1.0729 (improvised, 3 s), −2.3645 to −1.4126 (naturalistic, 2 s), and −2.8928 to −1.9899 (naturalistic, 3 s); the corresponding ranges were −2.3227 to −0.9231, −1.6209 to −0.3132, −3.4990 to −2.1083, and −3.3951 to −2.0136, respectively. Overall, the naturalistic condition provided the strongest evidence: both same-dyad similarity and dyad-specificity were retained across the 2 s and 3 s specifications, whereas the improvised condition showed significant same-dyad similarity but did not reach FDR-adjusted significance for dyad-specificity.
4.3. Effects After Removing Absolute Signal Levels
The naturalistic same-dyad and dyad-specificity effects remained significant in the dynamics-without-levels representation, which excluded the six absolute CIAS- and PMCI-level features. Under the 2 s specification, the same-dyad coefficient was = −1.6149 (FDR q = 0.0288), and the dyad-specificity contrast was = −2.5497 (FDR q = 0.0024). Under the 3 s specification, the corresponding estimates were = −2.0180 (FDR q = 0.0114) and = −2.4670 (FDR q = 0.0033). These results indicate that the detected dyad-associated structure was not explained solely by consistently high or low mean CIAS or PMCI values. The full representation produced convergent results. Under the naturalistic 2 s specification, the residualized full-representation same-dyad coefficient was = −2.0336 (FDR q = 0.0246), and the dyad-specificity contrast was = −2.9868 (FDR q = 0.0024). Thus, absolute signal levels were not required for detecting dyad-associated structure, although the full representation produced numerically stronger relational effects under this specification.
4.4. Feature-Group Robustness Results
For each representation feature group, robustness was summarized across eight analysis configurations defined by two interaction conditions (improvised and naturalistic), two minimum-dwell constraints (2 and 3 s), and two technical-adjustment specifications (raw representations with technical-distance adjustment and pair-cross-fitted technical residualization).
Same-dyad similarity reached FDR-adjusted significance in six of eight dynamics-without-levels analyses and six of eight full-representation analyses. State usage reached significance in four of eight analyses. Transition topology did not reach FDR-adjusted significance for either same-dyad similarity or dyad-specificity in any of the eight configurations.
Dyad-specificity reached FDR-adjusted significance in four dynamics-without-levels analyses, four full-representation analyses, and three state-usage analyses. Median nearest-dyad accuracy was 0.2635 for dynamics without levels, 0.3137 for the full representation, 0.1924 for state usage, and 0.1213 for transition topology.
These results indicate that the most consistent dyad-associated structure was captured by the broader organization of state use, dwell behavior, and signal dynamics rather than by transition topology alone. The feature-group robustness results are summarized in Table 6.
Table 6.
Robustness across feature groups.
4.5. Participant-Disjoint Transfer Robustness
Participant-disjoint robustness was evaluated using the frozen joint CIAS–PMCI state representation. The primary transfer specification used the dynamics-without-levels representation, a 3 s minimum-dwell constraint, pair-cross-fitted technical residualization, and 5000 permutations. The transfer sample contained 73 naturalistic sessions from six dyads and 11 participants who were absent from the discovery sample, providing a participant-disjoint evaluation of the learned representation.
In the participant-disjoint naturalistic sample, significant dyad-associated structure was retained. Same-dyad sessions were closer than other sessions ( = −0.5638, FDR q = 0.0152), while the dyad-specificity contrast remained strong ( = −2.2751, FDR q = 0.0002). Session-weighted nearest-dyad identification accuracy was 0.8356 (FDR q = 0.0002). To account for unequal numbers of sessions across dyads, macro-dyad accuracy was additionally calculated and reached 0.8512, with a dyad-bootstrap 95% confidence interval of 0.7540–0.9365.
Because one participant occurred in two dyads, the six dyads formed five participant-connected components. Equal weighting across these components yielded a macro accuracy of 0.8595 (component-bootstrap 95% CI: 0.7429–0.9548). All five components showed accuracy above their corresponding chance levels, and leave-one-component-out analyses retained session-weighted accuracy between 0.8197 and 0.9153. These analyses indicate that the participant-disjoint result was not driven by a single participant-connected component. The session-weighted accuracy of 0.8356 was also well above the permutation-null mean of 0.1628 (95% interval: 0.0685–0.2603).
Per-dyad nearest-identification accuracy was 0.7857 for P0116–P0410 (14 sessions), 0.8333 for P0410–P0665 (12 sessions), 0.6429 for P0467–P0697 (14 sessions), 0.9167 for P0484–P0485 (12 sessions), 1.0000 for P0671–P0672 (7 sessions), and 0.9286 for P0779–P0780 (14 sessions). The similarity of the session-weighted and macro-dyad estimates, together with this per-dyad distribution, indicates that the transfer result was not driven solely by one high-session or high-accuracy dyad. The participant-disjoint robustness results are summarized in Table 7.
Table 7.
Participant-disjoint robustness under the 3 s dynamics-without-levels specification.
4.6. Representation Ablation and Control Analyses
Representation ablation showed that PMCI was the principal source of separation in the dynamic-state representation and the primary contributor to the observed dyad-associated structure. PMCI-only representations produced significant same-dyad effects in both interaction conditions and strong dyad-specificity effects in the naturalistic condition. CIAS-only representations were weaker: they retained some same-dyad structure in the improvised condition but did not produce significant nearest-dyad identification and did not yield robust naturalistic effects after FDR correction.
Despite this unequal contribution, the joint CIAS–PMCI representation numerically outperformed PMCI alone across the evaluated interaction-condition, minimum-dwell, and representation specifications. Direct permutation tests showed that these improvements were selective rather than uniform. Under the improvised 2 s specification, nearest-dyad accuracy increased from 0.2500 for the PMCI-only dynamics representation to 0.4583 for the joint representation ( = 0.2083, FDR q = 0.0112). Under the improvised 3 s specification, the joint representation significantly strengthened dyad-specificity relative to PMCI alone (q = 0.0176 for the dynamics-without-levels representation). In the naturalistic condition, the clearest incremental effect occurred under the 3 s minimum-dwell constraint, where nearest-dyad accuracy increased from 0.2353 for PMCI alone to 0.3529 for the joint representation ( = 0.1176, FDR q = 0.0112). Direct improvements in the same-dyad coefficient did not survive FDR correction. Thus, the contribution of CIAS was modest and condition-dependent rather than uniformly significant across relational outcomes. Under the 2 s constraint, the corresponding accuracy improvements were positive but did not survive correction. Direct improvements in the same-dyad coefficient also did not remain significant after FDR correction in either interaction condition.
Same-dimensional controls provided a stricter test of whether the joint representation benefited from CIAS-specific content rather than merely from the addition of two extra dimensions. Real CIAS trajectories were compared with empirical shuffled-CIAS and matched autoregressive-noise controls while retaining the original PMCI trajectories. Across the evaluated conditions, dwell specifications, and representation sets, the real joint representation did not show uniform superiority over the control distributions. Thus, the control analyses do not support a general claim that CIAS consistently improves all relational outcomes beyond dimensionality and temporal-structure controls. Instead, they support a more limited interpretation in which CIAS provides complementary information under selected conditions, while PMCI remains the dominant contributor to the discrete state structure. The direct added-value comparisons are summarized in Table 8.
Table 8.
Direct added value of the joint CIAS–PMCI representation over PMCI alone for the dynamics-without-levels representation.
4.7. Bidirectional Cross-Attention Consistency–Discrepancy Extension
Architecture ablation on the held-out dyads showed modest differences among the four evaluated configurations. The no-attention model achieved ROC–AUC = 0.7621, PR–AUC = 0.7409, and accuracy = 0.6751. Face-to-body attention achieved ROC–AUC = 0.7782, PR–AUC = 0.7490, and accuracy = 0.6942, whereas body-to-face attention achieved ROC–AUC = 0.7707, PR–AUC = 0.7456, and accuracy = 0.6865. In the single primary run, bidirectional cross-attention yielded ROC–AUC = 0.7795, PR–AUC = 0.7523, and accuracy = 0.6907. Thus, cross-attention produced modest improvements in ranking performance relative to the no-attention configuration, without establishing a uniformly superior directional architecture.
Across five matched training seeds, performance differences among the attention variants remained small. Mean ROC–AUC was 0.7753 ± 0.0034 for no attention, 0.7767 ± 0.0066 for face-to-body attention, 0.7828 ± 0.0079 for body-to-face attention, and 0.7814 ± 0.0084 for bidirectional attention. Mean accuracies were similarly close (0.6925–0.6963). These results indicate that the advantage of bidirectional attention observed in the primary run was modest and not uniformly reproduced across seeds.
Consistency and discrepancy trajectories learned by the model were mostly uncorrelated with CIAS. Spearman rank correlations with CIAS were = 0.0003 (consistency) and = −0.0010 (discrepancy) in improvised interaction, and = −0.0361 and 0.0519, respectively, in naturalistic interaction. Associations with PMCI were stronger and opposite-signed. Consistency was positively correlated with PMCI ( = 0.3952 improvised, 0.3258 naturalistic) while discrepancy was negatively correlated ( = −0.4084 and −0.3105). Thus, the learned signals are not identical to CIAS or PMCI, but instead correlate with probabilistic affective agreement to varying degrees. Under the 3 s dynamics-without-levels specification, adding the ten learned C/D dynamics features strengthened the relational representation in the improvised condition. The normalized same-dyad coefficient changed from RMS = −0.2625 for the CIAS–PMCI baseline to −0.4232 for CIAS–PMCI+C/D, while dyad-specificity changed from −0.1940 to −0.3127. Direct comparison supported both improvements after FDR correction (RMS = −0.1607, q = 0.0045 for same-dyad similarity; RMS = −0.1187, q = 0.0486 for dyad-specificity), whereas nearest-dyad accuracy remained unchanged at 0.2917. In the naturalistic discovery sample, the same-dyad coefficient changed from RMS = −0.3958 to −0.4791, and dyad-specificity changed from −0.4838 to −0.5980, while nearest-dyad accuracy remained 0.3529. However, the direct incremental comparisons did not reach FDR-adjusted significance (q = 0.2042 for same-dyad similarity and q = 0.0694 for dyad-specificity). Thus, the learned C/D extension provided incremental relational information that reached statistical significance, but its added value was not consistently demonstrated across interaction conditions. In improvised interaction, real C+D features produced stronger same-dyad structure than shuffled C+D features (RMS = −0.1797, p = 0.0006) and stronger dyad-specificity (RMS = −0.1588, p = 0.0138). In naturalistic interaction, the corresponding comparisons were not significant (p = 0.0754 and p = 0.4977). Therefore, the evidence that the learned C/D features contain session-linked information beyond added dimensionality was strongest in the improvised interaction.
RMS denotes the pairwise-distance regression coefficient normalized by the square root of the effective representation dimensionality. More negative values indicate greater similarity among repeated sessions of the same dyad or stronger dyad-specificity. FDR q-values were obtained using Benjamini–Hochberg correction. The effects of the cross-attention consistency–discrepancy extension are summarized in Table 9.
Table 9.
Effect of the bidirectional cross-attention consistency/discrepancy extension on the dynamics-without-levels representation.
4.8. Unsupported Dynamic Analyses
Transition-topology features did not reach FDR-adjusted significance for same-dyad similarity or dyad-specificity across the eight primary analysis configurations. More specific analyses of gateway states, transition precursors, and leader–follower organization likewise did not provide statistically reliable evidence sufficient to support the principal claim. Exploratory direct cross-person temporal-coupling analysis also failed to show robust coupling beyond circular-shift null distributions for either CIAS or PMCI. Accordingly, these analyses were not used as evidence of interpersonal synchrony and are reported, where applicable, as supplementary or exploratory results.
5. Discussion
The results of this study support a reproducible joint profile of the two partners’ multimodal dynamics at the session level rather than direct interpersonal synchrony, psychological compatibility, or a stable relational trait. In the conventional psychological sense, interpersonal synchrony refers to temporal coordination between the members of a dyad [22,23,30], whereas CIAS and PMCI primarily quantify within-person cross-modal organization. The present results therefore concern the reproducibility of the partners’ joint multimodal dynamic profile across sessions.
The strongest evidence was observed in the naturalistic condition. Same-dyad sessions remained closer under both 2 s and 3 s minimum-dwell constraints, and the same-pair effect remained significant after pair-cross-fitted technical residualization. Dyad-specificity was also stronger than shared-participant similarity, indicating that the result was not reducible to the repeated presence of one individual. In the participant-disjoint transfer analysis, session-weighted nearest-dyad accuracy reached 0.8356 and macro-dyad accuracy reached 0.8512, showing that the observed structure transferred to previously unseen participants.
The findings also reinforce the distinction between affective prediction and multimodal consistency. Reliable affective prediction does not necessarily imply that a representation preserves the temporally structured cross-modal information required for analyzing interaction dynamics [3,4,17,18]. In the present study, CIAS and PMCI were therefore treated as time-varying observables of within-person cross-modal organization rather than as direct labels of emotion or interpersonal state. If an emotion is primarily expressed through the face and voice, and the body remains relatively static, then a face–body consistency score has limited relational information to exploit. This is especially relevant because affective meaning is not always carried by facial expression alone; body posture, motion, and kinematics can substantially shape emotional interpretation, particularly under ambiguity or high emotional intensity [10,11,12,13,14,15,16]. Thus, a suitable evaluation of affective consistency requires datasets in which the relevant modalities actually contain expressive and temporally structured information.
An important aspect of the result is that it persisted in the dynamics-without-levels representation. The detected structure therefore did not depend solely on one dyad exhibiting consistently higher or lower mean CIAS or PMCI values. Instead, information remained in state occupancy, dwell structure, switching behavior, entropy, and signal variability. This supports the interpretation of multimodal consistency as a temporally organized component of the session-level interaction representation rather than only as a scalar signal level [34,35].
The difference between the naturalistic and improvised conditions also requires a cautious interpretation. Both conditions showed significant same-dyad similarity, but only the naturalistic condition showed FDR-adjusted evidence that same-dyad similarity exceeded shared-participant similarity. This difference may reflect contextual or behavioral differences between the conditions; however, the present sample does not support a specific psychological explanation [28,34].
The joint states were primarily differentiated by PMCI, whereas CIAS varied within a narrower range. PMCI therefore provided the dominant structure for the discrete state partition. CIAS was not redundant with PMCI, but its incremental contribution was more modest and condition dependent. Accordingly, the two measures should be interpreted as complementary time-varying observables rather than as equally weighted determinants of the session-level representation.
The findings are consistent with dynamic dyadic approaches in which interaction is analyzed as an evolving system jointly constructed by both partners [34,35]. They are also compatible with the broad process logic of the Mutual Regulation Model, which emphasizes recurrent coordination and miscoordination [39]. However, the present study does not directly validate that model. The data-driven states cannot be interpreted as repair, rupture, regulation, or psychological attractors without external behavioral annotations.
Representation ablation clarified the unequal but potentially complementary roles of CIAS and PMCI. PMCI provided the principal source of state separation and independently retained substantial dyad-associated structure. CIAS-only representations were weaker, particularly in naturalistic interaction. Adding CIAS to PMCI produced improvements in selected identification and dyad-specificity analyses, but these improvements were not uniformly significant across conditions, outcomes, or dyad-level tests. Same-dimensional shuffled-CIAS and matched-noise controls likewise did not establish universal superiority of real CIAS. The evidence therefore supports a limited complementary contribution of CIAS rather than an equally dominant or uniformly additive role.
The cross-attention extension provides a complementary architectural perspective. In the primary architecture run, bidirectional attention achieved the highest ROC–AUC and PR–AUC among the evaluated configurations, although the differences were modest. Across five matched training seeds, however, no single attention direction consistently dominated: body-to-face attention had the highest mean ROC–AUC, whereas accuracy differed only minimally among the attention variants. Bidirectionality should therefore be interpreted as a competitive architectural option rather than as a universally superior configuration.
The downstream analyses further showed that the contribution of the learned consistency/discrepancy features was condition dependent. In improvised interaction, adding C+D significantly strengthened both same-dyad similarity and dyad-specificity relative to the CIAS–PMCI baseline, while nearest-dyad accuracy remained unchanged. In naturalistic interaction, the corresponding relational coefficients changed in the same direction, but the direct incremental comparisons did not reach FDR-adjusted significance. Thus, the learned extension provided incremental relational information that reached statistical significance in the improvised condition.
The lack of robust transition-topology, gateway, precursor, leader–follower, and direct cross-person coupling effects narrows the interpretation of the contribution. In particular, exploratory CIAS and PMCI cross-person coupling measures did not exceed circular-shift null expectations after FDR correction. The results therefore support reproducible session-level organization of multimodal dynamics, but they do not demonstrate direct interpersonal synchrony or a causal mechanism through which one partner regulates the other.
The main contribution of the current work is the demonstration that the joint temporal structure of CIAS and PMCI can be transformed into a reproducible session-level dynamic representation associated with repeated dyadic interaction. Thus, the bidirectional cross-attention module is best understood as an auxiliary learned extension of this framework rather than as a substitute for the interpretable CIAS–PMCI representation. It was used to assess whether explicit cross-modal interaction and learned consistency–discrepancy decomposition could recover relational information not already captured by the primary representation.
From an application perspective, the principal value of the proposed representation lies in longitudinal comparison of repeated interactions rather than in single-session psychological classification. Importantly, dyad-associated structure remained detectable after absolute CIAS and PMCI levels were removed, indicating that the informative component was not simply whether a dyad exhibited generally high or low cross-modal consistency. Instead, information was retained in the temporal organization of the signals, including state occupancy, dwell structure, switching behavior, and within-session variability. In human–computer interaction and socially interactive robotics, such a representation could therefore support comparison of a current interaction with previous sessions and provide an auxiliary indication of whether multimodal organization remains stable or changes over time [26,27]. Because PMCI provided the dominant contribution to the discrete state structure and the resulting regimes have no external behavioral labels, the representation is therefore best interpreted as a descriptive interaction-monitoring signal rather than an estimator of psychological state, relationship quality, or compatibility.
Clinical use would require independent behavioral or clinical validation. HCI and collaborative-robot deployment would additionally require robustness to occlusion, missing modalities, domain shift, latency, and privacy constraints. The main limitation of the discovery analysis relates to the number and nature of repeated observations. Although there were 12 dyads in each interaction condition, repeated-session data were available for only 9 dyads in the improvised condition and 4 dyads in the naturalistic condition, with singletons only informing between-dyad differences. Hence, while the permutation, bootstrap, and leave-one-dyad-out analyses help assess uncertainty in our estimates, they cannot substitute for a larger sample of independent repeated dyads. Because participants were reused across some dyads, there is another source of dependency that we addressed with the shared-participant and participant-component controls, but this still leaves open the need for replicated dyads from a larger pool of independent interactions. Coarse source relationship-category metadata were available and are reported descriptively, but they were not included as inferential predictors; longitudinal relationship history beyond these categories was unavailable. Pre-existing familiarity may therefore contribute to dyad-associated structure and cannot be fully separated from dyad identity in the present discovery sample. Exact elapsed intervals between repeated observations could not be reconstructed from the available interaction-local timing variables and therefore could not be evaluated as a contributor to reproducibility. A second limitation relates to the interpretation of the representation. The discrete state space was primarily organized by PMCI, with CIAS providing more secondary and context-specific information. Furthermore, the data-driven regimes were not associated with externally coded behaviors or psychological states. As such, they describe recurring regions of the multimodal signal space, but do not necessarily correspond to specific forms of engagement, conflict, regulation, repair, or interpersonal synchrony. Finally, the lack of significant effects on the transition topology and the fact that we were unable to detect robust effects in our exploratory direct cross-person coupling analyses suggest that our findings represent consistent patterns of organization at the session level, but do not identify any specific interpersonal coordination process.
The participant-disjoint analysis substantially strengthens the transfer evidence, but it should not be interpreted as an independent external-dataset replication because the 73-session sample originates from the same broader Seamless Interaction corpus and acquisition framework. Moreover, recording-session context was structurally associated with dyad identity in this sample: six recording-session levels were present, and none was shared across multiple dyads. Consequently, the present design does not allow dyad-associated structure to be completely separated from recording-context effects. Future validation should therefore use independently collected repeated-dyad datasets in which multiple dyads are recorded across shared contexts and the same dyads are observed across multiple recording contexts.
Finally, the cross-attention analysis should be interpreted as an auxiliary architecture experiment rather than as a separately powered model-comparison study. The number of held-out dyads was small, and the five-seed evaluation showed only modest differences among directional attention variants. These results support the feasibility of extracting additional learned cross-modal interaction features, but they do not establish a uniquely preferred attention direction or demonstrate that the learned extension should replace the more interpretable CIAS–PMCI representation.
6. Conclusions
This study integrated CIAS and PMCI as complementary time-varying measures of within-person face–body cross-modal organization and examined their joint temporal dynamics across repeated dyadic interaction sessions. The resulting session-level representation combined descriptive multimodal regimes with measures of state occupancy, dwell structure, transition characteristics, and continuous temporal variability. The identified regimes were treated strictly as data-driven descriptors and were not assigned psychological or behavioral labels.
In the naturalistic condition, repeated sessions involving the same dyad were more similar than sessions involving different dyads. This pattern remained after removal of absolute CIAS and PMCI levels and after pair-cross-fitted adjustment for measured technical characteristics. The effect was also retained in the participant-disjoint transfer analysis comprising 73 sessions from six dyads and 11 previously unseen participants, with a session-weighted nearest-dyad accuracy of 0.8356 and a macro-dyad accuracy of 0.8512. Thus, the principal supported conclusion is the reproducibility of dyad-associated multimodal dynamic organization across repeated interactions, rather than the identification of a specific psychological or interpersonal coordination mechanism.
The representation ablation analyses showed that PMCI provided the principal source of separation within the discrete state partition. The contribution of CIAS was more modest and also depended on the condition and outcome measure. In the auxiliary cross-attention consistency–discrepancy analysis, the results suggested that the learned cross-modal interaction features might enhance certain relational effects, especially in the improvised condition. No direction of attention showed a consistent overall benefit, and the learned cross-modal interaction did not replace the more directly interpretable CIAS–PMCI representation.
Future studies should evaluate the proposed framework in larger cohorts with repeated observations of the same dyads, in independently collected interaction datasets, and under designs that more clearly separate dyad identity from recording context. External behavioral annotations will also be required to determine whether the observed multimodal dynamic regimes correspond to specific and independently interpretable interaction processes.
Author Contributions
Conceptualization, Y.S. and S.K.; methodology, Y.S. and S.K.; software, Y.S.; validation, Y.S. and O.S.; formal analysis, Y.S.; investigation, Y.S. and S.K.; data curation, Y.S.; visualization, Y.S.; writing—original draft preparation, Y.S.; writing—review and editing, S.K. and O.S.; supervision, S.K.; project administration, Y.S. and S.K. All authors have read and agreed to the published version of the manuscript.
Funding
This research received no external funding.
Institutional Review Board Statement
The present study involved a secondary analysis of the previously collected and publicly released Seamless Interaction Dataset. The authors did not recruit new participants or collect new human-participant data. Ethical and privacy procedures for the original data collection are described by the creators of the Seamless Interaction Dataset.
Informed Consent Statement
No new participants were recruited for the present study. The original Seamless Interaction data collection was conducted with informed consent from the participants, as reported by the dataset creators.
Data Availability Statement
The data analyzed in this study are available through the publicly released Seamless Interaction Dataset described in Ref. [28]. The code and analysis scripts supporting the findings of this study are available in the authors’ GitHub repository: https://github.com/ernar65019920/Dyad-Specific-Multimodal-Affective-Dynamics (accessed on 10 August 2026).
Conflicts of Interest
The authors declare no conflicts of interest.
References
- Hazmoune, S.; Bougamouza, F. Using Transformers for Multimodal Emotion Recognition: Taxonomies and State of the Art Review. Eng. Appl. Artif. Intell. 2024, 133, 108339. [Google Scholar] [CrossRef] [Scilit]
- Ryumina, E.; Ryumin, D.; Axyonov, A.; Ivanko, D.; Karpov, A. Multi-Corpus Emotion Recognition Method Based on Cross-Modal Gated Attention Fusion. Pattern Recognit. Lett. 2025, 190, 192–200. [Google Scholar] [CrossRef] [Scilit]
- Kudubayeva, S.; Seksenbayev, Y.; Yerimbetova, A.; Daiyrbayeva, E.; Sakenov, B.; Telman, D.; Turdalyuly, M. Unified Visual Synchrony: A Framework for Face–Gesture Coherence in Multimodal Human–AI Interaction. Big Data Cogn. Comput. 2026, 10, 88. [Google Scholar] [CrossRef] [Scilit]
- Seksenbayev, Y.; Kudubayeva, S.; Baimankulov, A.; Yerimbetova, A.; Daiyrbayeva, E.; Berzhanova, U.; Sakenov, B. PMCI: A Prototype-Based Diagnostic Index for Cross-Modal Affective Agreement. Technologies 2026, 14, 445. [Google Scholar] [CrossRef] [Scilit]
- Darwin, C. The Expression of the Emotions in Man and Animals; John Murray: London, UK, 1872. [Google Scholar]
- Ekman, P.; Friesen, W.V. The Repertoire of Nonverbal Behavior: Categories, Origins, Usage, and Coding. Semiotica 1969, 1, 49–98. [Google Scholar] [CrossRef] [Scilit]
- Plutchik, R. Emotion: A Psychoevolutionary Synthesis; Harper & Row: New York, NY, USA, 1980. [Google Scholar]
- Russell, J.A. A Circumplex Model of Affect. J. Personal. Soc. Psychol. 1980, 39, 1161–1178. [Google Scholar] [CrossRef] [Scilit]
- Scherer, K.R. What Are Emotions? And How Can They Be Measured? Soc. Sci. Inf. 2005, 44, 695–729. [Google Scholar] [CrossRef] [Scilit]
- de Gelder, B. Towards the Neurobiology of Emotional Body Language. Nat. Rev. Neurosci. 2006, 7, 242–249. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Meeren, H.K.M.; van Heijnsbergen, C.C.R.J.; de Gelder, B. Rapid Perceptual Integration of Facial Expression and Emotional Body Language. Proc. Natl. Acad. Sci. USA 2005, 102, 16518–16523. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Aviezer, H.; Trope, Y.; Todorov, A. Body Cues, Not Facial Expressions, Discriminate between Intense Positive and Negative Emotions. Science 2012, 338, 1225–1229. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Lecker, M.; Dotsch, R.; Bijlstra, G.; Aviezer, H. Bidirectional Contextual Influence between Faces and Bodies in Emotion Perception. Emotion 2020, 20, 1154–1164. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Reisenzein, R.; Studtmann, M.; Horstmann, G. Coherence between Emotion and Facial Expression: Evidence from Laboratory Experiments. Emot. Rev. 2013, 5, 16–23. [Google Scholar] [CrossRef] [Scilit]
- Kleinsmith, A.; Bianchi-Berthouze, N. Affective Body Expression Perception and Recognition: A Survey. IEEE Trans. Affect. Comput. 2013, 4, 15–33. [Google Scholar] [CrossRef] [Scilit]
- Poyo Solanas, M.; Vaessen, M.J.; de Gelder, B. The Role of Computational and Subjective Features in Emotional Body Expressions. Sci. Rep. 2020, 10, 6202. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Baltrušaitis, T.; Ahuja, C.; Morency, L.-P. Multimodal Machine Learning: A Survey and Taxonomy. IEEE Trans. Pattern Anal. Mach. Intell. 2019, 41, 423–443. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Atrey, P.K.; Hossain, M.A.; El Saddik, A.; Kankanhalli, M.S. Multimodal Fusion for Multimedia Analysis: A Survey. Multimed. Syst. 2010, 16, 345–379. [Google Scholar] [CrossRef] [Scilit]
- Zadeh, A.; Chen, M.; Poria, S.; Cambria, E.; Morency, L.-P. Tensor Fusion Network for Multimodal Sentiment Analysis. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, Copenhagen, Denmark, 7–11 September 2017; pp. 1103–1114. [Google Scholar] [CrossRef] [Scilit]
- Tsai, Y.-H.H.; Bai, S.; Liang, P.P.; Kolter, J.Z.; Morency, L.-P.; Salakhutdinov, R. Multimodal Transformer for Unaligned Multimodal Language Sequences. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Florence, Italy, 28 July–2 August 2019; pp. 6558–6569. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Barros, P.; Wermter, S. Developing Crossmodal Expression Recognition Based on a Deep Neural Model. Adapt. Behav. 2016, 24, 373–396. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Ramseyer, F.; Tschacher, W. Nonverbal Synchrony in Psychotherapy: Coordinated Body Movement Reflects Relationship Quality and Outcome. J. Consult. Clin. Psychol. 2011, 79, 284–295. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Mayo, O.; Gordon, I. In and Out of Synchrony—Behavioral and Physiological Dynamics of Dyadic Interpersonal Coordination. Psychophysiology 2020, 57, e13574. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Bloch, C.; Tepest, R.; Koeroglu, S.; Feikes, K.; Jording, M.; Vogeley, K.; Falter-Wagner, C.M. Interacting with Autistic Virtual Characters: Intrapersonal Synchrony of Nonverbal Behavior Affects Participants’ Perception. Eur. Arch. Psychiatry Clin. Neurosci. 2024, 274, 1585–1599. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Miao, G.Q.; Dale, R.; Galati, A. (Mis)align: A Simple Dynamic Framework for Modeling Interpersonal Coordination. Sci. Rep. 2023, 13, 18325. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Chen, H.; Alghowinem, S.; Jang, S.J.; Breazeal, C.; Park, H.W. Dyadic Affect in Parent-Child Multimodal Interaction: Introducing the DAMI-P2C Dataset and Its Preliminary Analysis. IEEE Trans. Affect. Comput. 2023, 14, 3345–3361. [Google Scholar] [CrossRef] [Scilit]
- Javed, H.; Wang, W.; Usman, A.B.; Jamali, N. Modeling Interpersonal Perception in Dyadic Interactions: Towards Robot-Assisted Social Mediation in the Real World. Front. Robot. AI 2024, 11, 1410957. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Agrawal, V.; Akinyemi, A.; Alvero, K.; Behrooz, M.; Buffalini, J.; Carlucci, F.M.; Chen, J.; Chen, J.; Chen, Z.; Cheng, S.; et al. Seamless Interaction: Dyadic Audiovisual Motion Modeling and Large-Scale Dataset. arXiv 2025, arXiv:2506.22554. [Google Scholar]
- Funk, M.; Okada, S.; André, E. Multilingual Dyadic Interaction Corpus NoXi+J: Toward Understanding Asian-European Non-Verbal Cultural Characteristics and Their Influences on Engagement. In Proceedings of the 26th International Conference on Multimodal Interaction (ICMI ’24), San José, Costa Rica, 4–8 November 2024; pp. 224–233. [Google Scholar] [CrossRef] [Scilit]
- Ohayon, S.; Gordon, I. Multimodal Interpersonal Synchrony: Systematic Review and Meta-Analysis. Behav. Brain Res. 2025, 480, 115369. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Kodama, K.; Shimizu, D.; Fujiwara, K. Different Effects of Visual Occlusion on Interpersonal Coordination of Head and Body Movements during Dyadic Conversations. Front. Psychol. 2024, 15, 1296521. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Krackhardt, D. Predicting with Networks: Nonparametric Multiple Regression Analysis of Dyadic Data. Soc. Netw. 1988, 10, 359–381. [Google Scholar] [CrossRef] [Scilit]
- Dekker, D.; Krackhardt, D.; Snijders, T.A.B. Sensitivity of MRQAP Tests to Collinearity and Autocorrelation Conditions. Psychometrika 2007, 72, 563–581. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Solomon, D.H.; Brinberg, M.; Bodie, G.D.; Jones, S.; Ram, N. A Dynamic Dyadic Systems Approach to Interpersonal Communication. J. Commun. 2021, 71, 1001–1026. [Google Scholar] [CrossRef] [Scilit]
- Fogel, A.; Nwokah, E.; Dedo, J.Y.; Messinger, D.; Dickson, K.L.; Matusov, E.; Holt, S.A. Social Process Theory of Emotion: A Dynamic Systems Approach. Soc. Dev. 1992, 1, 122–142. [Google Scholar] [CrossRef] [Scilit]
- Lin, J. Divergence Measures Based on the Shannon Entropy. IEEE Trans. Inf. Theory 1991, 37, 145–151. [Google Scholar] [CrossRef] [Scilit]
- Hoerl, A.E.; Kennard, R.W. Ridge Regression: Biased Estimation for Nonorthogonal Problems. Technometrics 1970, 12, 55–67. [Google Scholar] [CrossRef]
- Benjamini, Y.; Hochberg, Y. Controlling the False Discovery Rate: A Practical and Powerful Approach to Multiple Testing. J. R. Stat. Soc. Ser. B (Methodol.) 1995, 57, 289–300. [Google Scholar] [CrossRef] [Scilit]
- Gianino, A.; Tronick, E.Z. The Mutual Regulation Model: The Infant’s Self and Interactive Regulation and Coping and Defensive Capacities. In Stress and Coping Across Development; Field, T.M., McCabe, P.M., Schneiderman, N., Eds.; Lawrence Erlbaum Associates: Hillsdale, NJ, USA, 1988; pp. 47–68. [Google Scholar]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.



