Next Article in Journal
DEPART: Multi-Task Interpretable Depression and Parkinson’s Disease Detection from In-the-Wild Video Data
Previous Article in Journal
An Intelligent Evaluation Method for Slope Stability Based on a Database Integrating Real Cases and Numerical Simulations
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Unified Visual Synchrony: A Framework for Face–Gesture Coherence in Multimodal Human–AI Interaction

by
Saule Kudubayeva
1,2,
Yernar Seksenbayev
2,*,
Aigerim Yerimbetova
1,3,*,
Elmira Daiyrbayeva
1,4,
Bakzhan Sakenov
1,
Duman Telman
1,5 and
Mussa Turdalyuly
1,3
1
Institute of Information and Computational Technologies CS MSHE RK, Almaty 050000, Kazakhstan
2
Faculty of Digital Sciences and Artificial Intelligence, L. N. Gumilyov Eurasian National University, Astana 010008, Kazakhstan
3
School of Engineering and Information Technology, META University, Almaty 050012, Kazakhstan
4
Department of Software Engineering, Satbayev University, Almaty 050010, Kazakhstan
5
School of Information Technologies and Applied Mathematics, SDU University, Kaskelen 040901, Kazakhstan
*
Authors to whom correspondence should be addressed.
Big Data Cogn. Comput. 2026, 10(3), 88; https://doi.org/10.3390/bdcc10030088
Submission received: 15 January 2026 / Revised: 28 February 2026 / Accepted: 9 March 2026 / Published: 12 March 2026

Abstract

Multimodal human–AI systems generally consider facial expressions and body motions as separate input streams, leading to disjointed interpretations and diminished emotional coherence. To overcome this issue, we offer the Engagement-Safe Expressive Alignment (ESEA) paradigm and the Unified Visual Synchrony (UVS) framework as its computational implementation. UVS models the coherence between facial expressions and gestures, offering an interpretable visual synchrony signal that can function as adaptive feedback in human–AI interactions. The framework’s key component is the Consistency Index for Affective Synchrony (CIAS), which correlates brief visual segments with scalar synchrony scores through a common latent representation. Facial and gestural signals are processed by modality-specific projection networks into a unified latent space, and CIAS is derived from the similarity and short-term temporal consistency of these latent trajectories. The synchrony index is regarded as an estimation of affective visual coherence within the ESEA paradigm. We formalize the UVS/CIAS framework and conduct a comparative experimental evaluation utilizing matched and mismatched face–gesture segments derived from rendered dialog footage. Utilizing ROC analysis, score distribution comparisons, temporal visualizations, and negative control tests, we illustrate that CIAS effectively captures structured face–gesture alignment that surpasses similarity-based baselines, while also delivering a persistent, time-resolved synchronization signal. These findings establish CIAS as a principled and interpretable feedback signal for future affect-aware, engagement-focused multimodal agents.
Keywords: multimodal fusion; synchrony; affective computing; gesture analysis; facial expression; human–AI interaction; adaptive systems multimodal fusion; synchrony; affective computing; gesture analysis; facial expression; human–AI interaction; adaptive systems

Share and Cite

MDPI and ACS Style

Kudubayeva, S.; Seksenbayev, Y.; Yerimbetova, A.; Daiyrbayeva, E.; Sakenov, B.; Telman, D.; Turdalyuly, M. Unified Visual Synchrony: A Framework for Face–Gesture Coherence in Multimodal Human–AI Interaction. Big Data Cogn. Comput. 2026, 10, 88. https://doi.org/10.3390/bdcc10030088

AMA Style

Kudubayeva S, Seksenbayev Y, Yerimbetova A, Daiyrbayeva E, Sakenov B, Telman D, Turdalyuly M. Unified Visual Synchrony: A Framework for Face–Gesture Coherence in Multimodal Human–AI Interaction. Big Data and Cognitive Computing. 2026; 10(3):88. https://doi.org/10.3390/bdcc10030088

Chicago/Turabian Style

Kudubayeva, Saule, Yernar Seksenbayev, Aigerim Yerimbetova, Elmira Daiyrbayeva, Bakzhan Sakenov, Duman Telman, and Mussa Turdalyuly. 2026. "Unified Visual Synchrony: A Framework for Face–Gesture Coherence in Multimodal Human–AI Interaction" Big Data and Cognitive Computing 10, no. 3: 88. https://doi.org/10.3390/bdcc10030088

APA Style

Kudubayeva, S., Seksenbayev, Y., Yerimbetova, A., Daiyrbayeva, E., Sakenov, B., Telman, D., & Turdalyuly, M. (2026). Unified Visual Synchrony: A Framework for Face–Gesture Coherence in Multimodal Human–AI Interaction. Big Data and Cognitive Computing, 10(3), 88. https://doi.org/10.3390/bdcc10030088

Article Metrics

Back to TopTop