Multimodal Sensing for Live-Stand Analytics: A Design-Oriented Literature Synthesis and Reference Architecture
Abstract
1. Introduction
2. Multimodal Behaviors and Satisfaction: A Review by Sensor Type
2.1. Methodology
2.2. Behavioral Sensing Approaches
2.3. Satisfaction Sensing Approaches
2.4. Sensing Advantages, Deployment Limitations and Suitability for Live-Stand Analytics
2.5. Datasets
2.6. Constructs (KPIs)
2.7. Discussion
2.8. Design Implications for SMBSA
3. Stand Multimodal Behavior and Satisfaction Analysis Architecture
3.1. Layers
- -
- Continuous KPIs. Engagement (En) and Social Interactions (So) are proposed as scalar estimates in [0, 1] by dedicated regressors that exploit the instance-level features provided by the perception modules. Conversion (Con), Commitment (Com), and Retention (Re) are included as candidate operational-outcome heads; before they can be estimated as calibrated, utility-oriented scalar values, they require validated labels from transaction records, repeat-visit identifiers, consented surveys, or staff-validated logs. Engagement, for example, fuses head-pose, facial dynamics, acoustic cues, and body movement into a continuous attentiveness score [11,29], while the Social Interactions head would estimate proxemic/F-formation social groups from spatiotemporal position, orientation, and interpersonal-distance cues [71,75]. On the satisfaction side, Attention (At), calibrated Dwell Time proxy (Dw), and Emotion (Em) may be normalized and oriented to [0, 1]. Dwell Time is retained only as a contextual proxy because it can indicate interest, congestion, or confusion depending on zone and task.
- -
- Mapped KPIs. Feedback (Fe), Odd behavior (Od), Person counting (Pe), Sentiment (Se), and Conversation (Co) are initially returned as labels, counts, or event outputs (e.g., “applause”, “running”, “positive”, “discussion”). These outputs require validated utility mappings or explicitly documented provisional expert-defined utility tables before contributing to the mapped behavior and satisfaction components, and , defined below. The Feedback head is treated as a proxy-feedback head for observable response events from sound-event cues such as speech, pauses, and applause, optionally fused with spatial context in future microphone-array deployments [35]. Validated feedback labels or staff/survey confirmation are still required before interpreting it as a stand-level feedback KPI. The Odd Behavior head combines an autoencoder-based reconstruction score with a discriminative classifier [17,27].
- -
- Spatial map KPIs. Density Maps (Dm) and spatialized Motion Patterns (Mo) are represented as 2D spatial arrays. The Density Map head employs CSRNet-style convolutions with count regression to generate continuous density surfaces [8]. The Motion-Pattern head performs multi-label activity classification on temporally extended features [23,34]; it enters the spatial component only after its outputs are converted into a declared spatial activity-intensity map . In deployments without a declared , Mo is excluded from the displayed composite unless a deployment-specific mapped Motion Patterns variant and corresponding candidate set are explicitly declared. These maps are not used directly as scalar KPIs; they are reduced by predefined map summaries and combined into the spatial behavior component defined below.
- -
- Metrics general computation. The following equations define a bounded linear aggregation scheme for the proposed architecture, not an empirically validated scoring model. Let W denote a time window and z a spatial scope, such as the full stand or a specific zone. For each calibrated scalar KPI, let denote the corresponding primitive estimate. For categorical, count, or event-label outputs, let denote a validated or explicitly provisional utility mapping from the raw head output to a scalar score. For spatial maps, let denote a predefined map summary, such as a zone-weighted mean, percentile, or normalized integral. Every , , and must be oriented so that larger values carry more favorable evidence for the target construct; adverse or non-monotonic raw outputs require an explicit transformation before aggregation.
3.2. Pipeline
4. Conclusions and Future Work
Supplementary Materials
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| 1D/2D/3D | One/two/three-dimensional |
| AI | Artificial Intelligence |
| CNN | Convolutional Neural Network |
| FOA | Focus of Attention |
| HOTA | Higher Order Tracking Accuracy |
| GDPR | General Data Protection Regulation |
| IEC | International Electrotechnical Commission |
| ISO | International Standardization Organization |
| KPI | key performance indicators |
| PCA | Principal Component Analysis |
| RGB-D | Red, Green, Blue - Depth sensor/camera |
| ROI | Return on Investment |
| RoI | Region of Interest |
| SMBSA | Stand Multimodal Behavior and Satisfaction Analysis |
References
- Bloch, P.H.; Gopalakrishna, S.; Crecelius, A.T.; Murarolli, M.S. Exploring booth design as a determinant of trade show success. J. Bus. Bus. Mark. 2017, 24, 237–256. [Google Scholar] [CrossRef]
- Ramos, C.M.Q.; Rodrigues, J.M.F. SNUX2.0: A Social Network Model for Cohort Behaviour Analysis as Support for Purchasing Tourism Products and Services. J. Relatsh. Mark. 2023, 22, 132–151. [Google Scholar] [CrossRef]
- Sun, D.; Xue, M.; Zhao, Y. Identifying and controlling key factors in exhibition effect: A hybrid method combining Best-Worst Method and regression models. Front. Commun. 2025, 10, 1670964. [Google Scholar] [CrossRef]
- European Parliament and Council of the European Union. Regulation (EU) 2016/679 of the European Parliament and of the Council of 27 April 2016 on the protection of natural persons with regard to the processing of personal data and on the free movement of such data (General Data Protection Regulation). Off. J. Eur. Union 2016, L 119, 1–88. Available online: http://data.europa.eu/eli/reg/2016/679/oj (accessed on 10 July 2026). [CrossRef]
- European Parliament and Council of the European Union. Regulation (EU) 2024/1689 of the European Parliament and of the Council of 13 June 2024 laying down harmonised rules on artificial intelligence (Artificial Intelligence Act). Off. J. Eur. Union 2024, L Series, 2024/1689. Available online: http://data.europa.eu/eli/reg/2024/1689/oj (accessed on 10 July 2026).
- Pei, G.; Li, H.; Lu, Y.; Wang, Y.; Hua, S.; Li, T. Affective computing: Recent advances, challenges, and future trends. Intell. Comput. 2024, 3, 0076. [Google Scholar] [CrossRef]
- Morshed, M.G.; Sultana, T.; Alam, A.; Lee, Y.K. Human action recognition: A taxonomy-based survey, updates, and opportunities. Sensors 2023, 23, 2182. [Google Scholar] [CrossRef] [PubMed]
- Wang, M.; Zhou, X.; Chen, Y. A comprehensive survey of crowd density estimation and counting. IET Image Process. 2025, 19, e13328. [Google Scholar] [CrossRef]
- Rodrigues, J.M.F.; Pereira, J.A.R.; Sardo, J.D.P.; Freitas, M.A.G.; Cardoso, P.J.S.; Gomes, M.; Bica, P. Adaptive Card Design UI Implementation for an Augmented Reality Museum Application. In Proceedings of the Universal Access in Human–Computer Interaction. Design and Development Approaches and Methods. UAHCI 2017; Lecture Notes in Computer Science; Springer International Publishing: Cham, Switzerland, 2017; Volume 10277, pp. 433–443. [Google Scholar] [CrossRef]
- Standard ISO 9241-11:2018; Ergonomics of Human-System Interaction—Part 11: Usability: Definitions and Concepts. International Organization for Standardization: Geneva, Switzerland, 2018. Available online: https://www.iso.org/standard/63500.html (accessed on 10 July 2026).
- Abdelrahman, A.A.; Strazdas, D.; Khalifa, A.; Hintz, J.; Hempel, T.; Al-Hamadi, A. Multimodal Engagement Prediction in Multiperson Human-Robot Interaction. IEEE Access 2022, 10, 61980–61991. [Google Scholar] [CrossRef]
- Navarro, R.C.; Ruiz, A.R.; Molina, F.J.; Romero, M.J.; Chaparro, J.D.; Alises, D.V.; Lopez, J.C. Indoor occupancy estimation for smart utilities: A novel approach based on depth sensors. Build. Environ. 2022, 222, 109406. [Google Scholar] [CrossRef]
- Khaire, P.; Kumar, P. A semi-supervised deep learning based video anomaly detection framework using RGB-D for surveillance of real-world critical environments. Forensic Sci. Int. Digit. Investig. 2022, 40, 301346. [Google Scholar] [CrossRef]
- Kajendran, K.; Mayan, J.A. Recognition and detection of unusual activities in ATM using dual-channel capsule generative adversarial network. Expert Syst. Appl. 2024, 247, 122987. [Google Scholar] [CrossRef]
- Abed, A.; Akrout, B.; Amous, I. Convolutional Neural Network for Head Segmentation and Counting in Crowded Retail Environment Using Top-view Depth Images. Arab. J. Sci. Eng. 2024, 49, 3735–3749. [Google Scholar] [CrossRef]
- Mishima, Y.; Matsui, T.; Matsuda, Y.; Suwa, H.; Yasumoto, K. Micro Activity Recognition Using Multi-View 3D Point Clouds. In Proceedings of the 2024 IEEE International Conference on Pervasive Computing and Communications Workshops and Other Affiliated Events, PerCom Workshops 2024; Institute of Electrical and Electronics Engineers Inc.: Piscataway, NJ, USA, 2024; pp. 453–456. [Google Scholar] [CrossRef]
- Cao, J.; Zhou, K.; Du, J. HyPCV-Former: Hyperbolic spatio-temporal transformer for 3D point cloud video anomaly detection. Adv. Eng. Inform. 2026, 73, 104537. [Google Scholar] [CrossRef]
- Kim, G.Y.; Ko, E.S.; Kim, D.R.; Kim, J.E.; Hindsley, D.; Matson, E.T. Real-Time Crowd Density Estimation and Stampede Risk Assessment System Using Thermal Camera. In Proceedings of the Proceedings—2024 8th IEEE International Conference on Robotic Computing, IRC 2024; IEEE: Piscataway, NJ, USA, 2024. [Google Scholar] [CrossRef]
- Hishida, R.; Iryo, M.; Alhamdani, R.; Mano, K.; Yamaguchi, Y. Development of Methods for High-Density Crowd Measurement and Tracking in Railway Station Concourses. Int. J. Intell. Transp. Syst. Res. 2026; early access. [CrossRef]
- Wu, Z.; Cao, Z.; Yu, X.; Zhu, J.; Song, C.; Xu, Z. A Novel Multiperson Activity Recognition Algorithm Based on Point Clouds Measured by Millimeter-Wave MIMO Radar. IEEE Sens. J. 2023, 23, 19509–19523. [Google Scholar] [CrossRef]
- Canil, M.; Pegoraro, J.; Shastri, A.; Casari, P.; Rossi, M. ORACLE: Occlusion-Resilient and Self-Calibrating mmWave Radar Network for People Tracking. IEEE Sens. J. 2024, 24, 3157–3171. [Google Scholar] [CrossRef]
- Martin-Martin, A.; Verona-Almeida, M.; Padial-Allue, R.; Saez, B.; Mendez, J.; Castillo, E.; Parrilla, L. Spiking Neural Networks for People Counting Based on FMCW Radar. IEEE Access 2025, 13, 60846–60858. [Google Scholar] [CrossRef]
- Zhang, F.; Sun, H.; Peng, J.; Wang, H. LPBS-Net: A Lightweight Network for Human Activity Recognition from Sparse Millimeter-Wave Radar Point Clouds. IEEE Sens. Lett. 2025, 9, 6012704. [Google Scholar] [CrossRef]
- Zhang, N.; Li, H.; Zahid, A.; Tian, Y.; Li, W. A Robust mmWave Radar Framework for Accurate People Counting and Motion Classification. Sensors 2026, 26, 1289. [Google Scholar] [CrossRef] [PubMed]
- Roche, J.; De-Silva, V.; Hook, J.; Moencks, M.; Kondoz, A. A Multimodal Data Processing System for LiDAR-Based Human Activity Recognition. IEEE Trans. Cybern. 2022, 52, 10027–10040. [Google Scholar] [CrossRef] [PubMed]
- Zhang, X.; Li, Z.; Zhang, J. Synthesized Millimeter-Waves for Human Motion Sensing. In Proceedings of the SenSys 2022—Proceedings of the 20th ACM Conference on Embedded Networked Sensor Systems; Association for Computing Machinery: New York, NY, USA, 2022. [Google Scholar] [CrossRef]
- Nguyen, V.A.; Kong, S.G. Multimodal feature fusion for illumination-invariant recognition of abnormal human behaviors. Inf. Fusion 2023, 100, 101949. [Google Scholar] [CrossRef]
- Cheng, P.; Xiong, Z.; Bao, Y.; Zhuang, P.; Zhang, Y.; Blasch, E.; Chen, G. A Deep Learning-Enhanced Multi-Modal Sensing Platform for Robust Human Object Detection and Tracking in Challenging Environments. Electronics 2023, 12, 3423. [Google Scholar] [CrossRef]
- Tu, V.N.; Huynh, V.T.; Yang, H.J.; Kim, S.H.; Nawaz, S.; Nandakumar, K.; Zaheer, M.Z. DCTM: Dilated Convolutional Transformer Model for Multimodal Engagement Estimation in Conversation. In Proceedings of the MM 2023—Proceedings of the 31st ACM International Conference on Multimedia; Association for Computing Machinery: New York, NY, USA, 2023. [Google Scholar] [CrossRef]
- Shafizadegan, F.; Naghsh-Nilchi, A.R.; Shabaninia, E. Multimodal vision-based human action recognition using deep learning: A review. Artif. Intell. Rev. 2024, 57, 178. [Google Scholar] [CrossRef]
- Kamra, V.; Vaishnav, A.; Verma, A.; Khan, R.; Singh, S. A Novel Approach for Crowd Analysis and Density Estimation by Using Machine Learning Techniques. In Proceedings of the 2024 International Conference on Intelligent Systems for Cybersecurity, ISCS 2024; Institute of Electrical and Electronics Engineers Inc.: Piscataway, NJ, USA, 2024. [Google Scholar] [CrossRef]
- Mu, B.; Shao, F.; Xie, Z.; Chen, H.; Jiang, Q.; Ho, Y.S. Visual Prompt Multibranch Fusion Network for RGB-Thermal Crowd Counting. IEEE Internet Things J. 2024, 11, 31758–31775. [Google Scholar] [CrossRef]
- She, X.; Xu, Z. Human Abnormal Behavior Detection Based on Multimodal Data Fusion. In Proceedings of the 2nd IEEE International Conference on Data Science and Network Security, ICDSNS 2024; IEEE: Piscataway, NJ, USA, 2024. [Google Scholar] [CrossRef]
- Tsiktsiris, D.; Lalas, A.; Dasygenis, M.; Votis, K. Multimodal Abnormal Event Detection in Public Transportation. IEEE Access 2024, 12, 133469–133480. [Google Scholar] [CrossRef]
- Vrochidis, A.; Dimitriou, N.; Krinidis, S.; Panagiotidis, S.; Parcharidis, S.; Tzovaras, D. A Deep Learning Framework for Monitoring Audience Engagement in Online Video Events. Int. J. Comput. Intell. Syst. 2024, 17, 124. [Google Scholar] [CrossRef]
- Shin, J.; Hassan, N.; Miah, A.S.M.; Nishimura, S. A Comprehensive Methodological Survey of Human Activity Recognition Across Diverse Data Modalities. Sensors 2025, 25, 4028. [Google Scholar] [CrossRef] [PubMed]
- Chen, X.; Yang, J. X-FI: A Modality-Invariant Foundation Model for Multimodal Human Sensing. In Proceedings of the 13th International Conference on Learning Representations, ICLR 2025; OpenReview: Alameda, CA, USA, 2025; Available online: https://openreview.net/forum?id=b42wmsdwmB (accessed on 10 July 2026).
- Lei, L. The artificial intelligence technology for immersion experience and space design in museum exhibition. Sci. Rep. 2025, 15, 27317. [Google Scholar] [CrossRef] [PubMed]
- Xu, B. IoT-based multimodal learning framework for predicting student engagement in English education. In Proceedings of the Second International Conference on Intelligent Transportation and Smart Cities (ICITSC 2025); Li, Y., Mezhuyev, V., Li, Z., Eds.; SPIE: Bellingham, WA, USA, 2025; p. 109. [Google Scholar] [CrossRef]
- Liu, X. AI-driven real-time responsive design of urban open spaces based on multi-modal sensing data fusion. Sci. Rep. 2025, 15, 41255. [Google Scholar] [CrossRef] [PubMed]
- Jeon, M.; Woo, S. A Lightweight Radar–Camera Fusion Deep Learning Model for Human Activity Recognition. Sensors 2026, 26, 894. [Google Scholar] [CrossRef] [PubMed]
- Devi, H.; Kumar, P.; Govindarajan, V.; Kumar, S.; Lohano, R.; Hitesh, H.; Shiwlani, A. A Comparative Study of Classical Machine Learning and Deep Learning Approaches for Human Behavior Detection Using Multisensor Data. IEEE Access 2026, 14, 25311–25325. [Google Scholar] [CrossRef]
- Vaz, P.J.; Rodrigues, J.M.F.; Cardoso, P.J.S. Affective Computing Emotional Body Gesture Recognition: Evolution and the Cream of the Crop. IEEE Access 2025, 13, 192871–192890. [Google Scholar] [CrossRef]
- S, N.P.; Koti, M.S.; G, T.K.; Anwar, S.; J, G.B.; Thinakaran, R. Real Time Customer Satisfaction Analysis using Facial Expressions and Headpose Estimation. Int. J. Adv. Comput. Sci. Appl. 2022, 13, 231–238. [Google Scholar] [CrossRef]
- Yusupova, N.I.; Bogdanova, D.R.; Nuriakhmetov, A.I. Assessing the Quality of Customer Service Based on the Emotional Satisfaction of Clients Using Artificial Immune System Technologies. Pattern Recognit. Image Anal. 2023, 33, 544–554. [Google Scholar] [CrossRef]
- Kwon, D.H.; Yu, J.M. Real-time Multi-CNN-based Emotion Recognition System for Evaluating Museum Visitors’ Satisfaction. J. Comput. Cult. Herit. 2024, 17, 1–18. [Google Scholar] [CrossRef]
- Gangan, H.A.; Rohani, M.; Hosseini, S.A.; Mansouri, A. Detection of Customer Satisfaction in In-Person Telecommunication Services Using Automated Facial Image Analysis with Artificial Intelligence. In Proceedings of the 11th International Symposium on Telecommunication: Communication in the Age of Artificial Intelligence, IST 2024; Institute of Electrical and Electronics Engineers Inc.: Piscataway, NJ, USA, 2024; pp. 485–490. [Google Scholar] [CrossRef]
- Harianto, D.; Filbert, S.; Cahyakusuma, A.B.; Zakiyyah, A.Y. Analyzing Customer Satisfaction Through Face Emotion Recognition: A Comparative Study of Convolutional Neural Networks (CNN) and Long Short Term Memory (LSTM). In Proceedings of the 2024 10th International Conference on Smart Computing and Communication, ICSCC 2024; Institute of Electrical and Electronics Engineers Inc.: Piscataway, NJ, USA, 2024; pp. 50–54. [Google Scholar] [CrossRef]
- Alhasson, H.F.; Alsaheel, G.M.; Alsalamah, A.A.; Alharbi, N.S.; Alhujilan, J.M.; Alharbi, S.S. Integration of machine learning bi-modal engagement emotion detection model to self-reporting for educational satisfaction measurement. Int. J. Inf. Technol. 2024, 16, 3633–3647. [Google Scholar] [CrossRef]
- Zhang, J.; Sato, W.; Kawamura, N.; Shimokawa, K.; Tang, B.; Nakamura, Y. Sensing emotional valence and arousal dynamics through automated facial action unit analysis. Sci. Rep. 2024, 14, 19563. [Google Scholar] [CrossRef] [PubMed]
- Karthikayani; Nithya, A.R.; Bhattacharya, C.; Saraswathi, C.; Prasanna, S.S.; Nandhini, M. Predicting Faculty Emotions and Job Satisfaction from Facial Expressions in Higher Education Using CNN-SVM Models. In Proceedings of the 2025 Global Conference in Emerging Technology, GINOTECH 2025; Institute of Electrical and Electronics Engineers Inc.: Piscataway, NJ, USA, 2025. [Google Scholar] [CrossRef]
- Lin, K.C.; Lin, Y.H.; Chen, M.Y. A Realtime Classroom Assessment System for Analysis of Students’ Evaluation of Teaching Through a Deep Learning and Emotional Contagion Mechanism. Int. J. Interact. Multimed. Artif. Intell. 2025, 9, 51–59. [Google Scholar] [CrossRef]
- Sang, T.V.D.; Hoan, N.N.; Tung, L.C.; Thu, N.T.A.; Tuan, P.V. Proposed Machine Learning Approach To Measure Nonverbal Clues-Based Student Satisfaction In STEAM Maker Innovation Space. In Proceedings of the 2025 IEEE International Conference on Artificial Intelligence and Mechatronics Systems, AIMS 2025; Institute of Electrical and Electronics Engineers Inc.: Piscataway, NJ, USA, 2025. [Google Scholar] [CrossRef]
- Ma, C.; Zhao, S.; Zhou, D.; Pei, Y.; Luo, Z.; Xie, L.; Yan, Y.; Yin, E. MPFNet: A Multi-Prior Fusion Network with a Progressive Training Strategy for Micro-Expression Recognition. IEEE Trans. Affect. Comput. 2026, 17, 348–365. [Google Scholar] [CrossRef]
- Shi, M.; Zheng, W. Towards Identity-Independent Facial Action Unit Detection: Integrating Decoupled 3D Geometry with Textural Features. IEEE Trans. Affect. Comput. 2026, 17, 482–496. [Google Scholar] [CrossRef]
- Zhu, Q.; Mao, Q.; Dong, W.; Shao, X.; Huang, X.; Zheng, W. Adaptive Key Role Guided Hierarchical Relation Inference for Enhanced Group-level Emotion Recognition. IEEE Trans. Affect. Comput. 2026, 17, 366–378. [Google Scholar] [CrossRef]
- Lemos, M.; Cardoso, P.J.S.; Rodrigues, J.M.F. From Cues to Engagement: A Comprehensive Survey and Holistic Architecture for Computer Vision-Based Audience Analysis in Live Events. Multimodal Technol. Interact. 2026, 10, 8. [Google Scholar] [CrossRef]
- Assiri, B.; Hossain, M.A. Face emotion recognition based on infrared thermal imagery by applying machine learning and parallelism. Math. Biosci. Eng. 2023, 20, 913–929. [Google Scholar] [CrossRef] [PubMed]
- Rashmi, R.; Snekhalatha, U.; Salvador, A.L.; Raj, A.N.J. Facial emotion detection using thermal and visual images based on deep learning techniques. Imaging Sci. J. 2024, 72, 153–166. [Google Scholar] [CrossRef]
- Parra-Gallego, L.F.; Orozco-Arroyave, J.R. Classification of emotions and evaluation of customer satisfaction from speech in real world acoustic environments. Digit. Signal Process. Rev. J. 2022, 120, 103286. [Google Scholar] [CrossRef]
- Ko, Y.H.; Hsu, P.Y.; Liu, Y.C.; Yang, P.C. Confirming Customer Satisfaction with Tones of Speech. IEEE Access 2022, 10, 83236–83248. [Google Scholar] [CrossRef]
- Chawla, K.; Clever, R.; Ramirez, J.; Lucas, G.M.; Gratch, J. Towards Emotion-Aware Agents for Improved User Satisfaction and Partner Perception in Negotiation Dialogues. IEEE Trans. Affect. Comput. 2024, 15, 433–444. [Google Scholar] [CrossRef]
- Ganesan, S. Deep learning model for identification of customers satisfaction in business. J. Auton. Intell. 2024, 7, 1–11. [Google Scholar] [CrossRef]
- Parra-Gallego, L.F.; Arias-Vergara, T.; Orozco-Arroyave, J.R. Multimodal evaluation of customer satisfaction from voicemails using speech and language representations. Digit. Signal Process. Rev. J. 2025, 156, 104820. [Google Scholar] [CrossRef]
- Gan, C.; Zhou, D.; Zhu, Q.; Wang, X.; Jain, D.K.; Struc, V. Improving Emotion Recognition from Ambiguous Speech via Spatio-Temporal Spectrum Analysis and Real-Time Soft-Label Correction. IEEE Trans. Affect. Comput. 2026, 17, 1058–1073. [Google Scholar] [CrossRef]
- Perez-Toro, P.A.; Vasquez-Correa, J.C.; Bocklet, T.; Noth, E.; Orozco-Arroyave, J.R. User State Modeling Based on the Arousal-Valence Plane: Applications in Customer Satisfaction and Health-Care. IEEE Trans. Affect. Comput. 2023, 14, 1533–1546. [Google Scholar] [CrossRef]
- Maris, L.; Matsuda, Y.; Sadre, R.; Yasumoto, K. Towards Cheaper Tourists’ Emotion and Satisfaction Estimation with PCA and Subgroup Analysis. In Proceedings of the 2023 IEEE International Conference on Pervasive Computing and Communications Workshops and other Affiliated Events, PerCom Workshops 2023; Institute of Electrical and Electronics Engineers Inc.: Piscataway, NJ, USA, 2023; pp. 502–508. [Google Scholar] [CrossRef]
- Maheen, S.M.; Sultana, I.; Kshetri, N.; Zim, M.N.F. emoAIsec: Fortifying Real-Time Customer Experience Optimization with Emotion AI and Data Security. In Proceedings of the 2nd International Conference on Machine Learning and Autonomous Systems, ICMLAS 2025—Proceedings; Institute of Electrical and Electronics Engineers Inc.: Piscataway, NJ, USA, 2025; pp. 793–798. [Google Scholar] [CrossRef]
- Luo, Y.; Liu, W.; Sun, Q.; Li, S.; Li, J.; Wu, R.; Tang, X. TriagedMSA: Triaging Sentimental Disagreement in Multimodal Sentiment Analysis. IEEE Trans. Affect. Comput. 2025, 16, 1557–1569. [Google Scholar] [CrossRef]
- Brscic, D.; Kanda, T.; Ikeda, T.; Miyashita, T. Person Tracking in Large Public Spaces Using 3-D Range Sensors. IEEE Trans. Hum.-Mach. Syst. 2013, 43, 522–534. [Google Scholar] [CrossRef]
- Alameda-Pineda, X.; Staiano, J.; Subramanian, R.; Batrinca, L.; Ricci, E.; Lepri, B.; Lanz, O.; Sebe, N. SALSA: A novel dataset for multimodal group behavior analysis. IEEE Trans. Pattern Anal. Mach. Intell. 2016, 38, 1707–1720. [Google Scholar] [CrossRef] [PubMed]
- Zhang, Z.; Rhim, J.; TaherAhmadi, M.; Yang, K.; Lim, A.; Chen, M. SFU-store-nav: A multimodal dataset for indoor human navigation. Data Brief 2020, 33, 106539. [Google Scholar] [CrossRef] [PubMed]
- Ehsanpour, M.; Saleh, F.; Savarese, S.; Reid, I.; Rezatofighi, H. JRDB-Act: A Large-scale Dataset for Spatio-temporal Action, Social Group and Activity Detection. In Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition; IEEE Computer Society: Piscataway, NJ, USA, 2022; pp. 20951–20960. [Google Scholar] [CrossRef]
- An, S.; Li, Y.; Ogras, U. mRI: Multi-modal 3D Human Pose Estimation Dataset using mmWave, RGB-D, and Inertial Sensors. In Proceedings of the Advances in Neural Information Processing Systems, New Orleans, LA, USA, 28 November–9 December 2022; Volume 35. Available online: https://proceedings.neurips.cc/paper_files/paper/2022/hash/af9c9c6d2da701da5a0acf91ec217815-Abstract-Datasets_and_Benchmarks.html (accessed on 10 July 2026).
- Su, J.; Huang, J.; Qing, L.; He, X.; Chen, H. A new approach for social group detection based on spatio-temporal interpersonal distance measurement. Heliyon 2022, 8, e11038. [Google Scholar] [CrossRef] [PubMed]
- Martin-Martin, R.; Patel, M.; Rezatofighi, H.; Shenoi, A.; Gwak, J.Y.; Frankel, E.; Sadeghian, A.; Savarese, S. JRDB: A Dataset and Benchmark of Egocentric Robot Visual Perception of Humans in Built Environments. IEEE Trans. Pattern Anal. Mach. Intell. 2023, 45, 6748–6765. [Google Scholar] [CrossRef] [PubMed]
- Vendrow, E.; Le, D.T.; Cai, J.; Rezatofighi, H. JRDB-Pose: A Large-Scale Dataset for Multi-Person Pose Estimation and Tracking. In Proceedings of the 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: Piscataway, NJ, USA, 2023; pp. 4811–4820. [Google Scholar] [CrossRef]
- Piadyk, Y.; Rulff, J.; Brewer, E.; Hosseini, M.; Ozbay, K.; Sankaradas, M.; Chakradhar, S.; Silva, C. StreetAware: A High-Resolution Synchronized Multimodal Urban Scene Dataset. Sensors 2023, 23, 3710. [Google Scholar] [CrossRef] [PubMed]
- Le, D.T.; Gou, C.; Datta, S.; Shi, H.; Reid, I.; Cai, J.; Rezatofighi, H. JRDB-PanoTrack: An Open-World Panoptic Segmentation and Tracking Robotic Dataset in Crowded Human Environments. In Proceedings of the 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: Piscataway, NJ, USA, 2024; pp. 22325–22334. [Google Scholar] [CrossRef]
- Jahangard, S.; Cai, Z.; Wen, S.; Rezatofighi, H. JRDB-Social: A Multifaceted Robotic Dataset for Understanding of Context and Dynamics of Human Interactions Within Social Groups. In Proceedings of the 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: Piscataway, NJ, USA, 2024; pp. 22087–22097. [Google Scholar] [CrossRef]
- Zhou, Y.; Song, N.; Ma, J.; Man, K.L.; López-Benítez, M.; Yu, L.; Yue, Y. RAV4D: A Radar-Audio-Visual Dataset for Indoor Multi-Person Tracking. In Proceedings of the IEEE Radar Conference; Institute of Electrical and Electronics Engineers Inc.: Piscataway, NJ, USA, 2024. [Google Scholar] [CrossRef]
- Gucsi, B.; Tuyen, N.T.V.; Chu, B.; Tarapore, D.; Tran-Thanh, L. HRI-SENSE: A Multimodal Dataset on Social and Emotional Responses to Robot Behaviour. In Proceedings of the ACM/IEEE International Conference on Human-Robot Interaction; IEEE: Piscataway, NJ, USA, 2025. [Google Scholar] [CrossRef]
- Abdrakhmanova, M.; Kuzdeuov, A.; Jarju, S.; Khassanov, Y.; Lewis, M.; Varol, H.A. Speakingfaces: A large-scale multimodal dataset of voice commands with visual and thermal video streams. Sensors 2021, 21, 3465. [Google Scholar] [CrossRef] [PubMed]
- Fard, A.P.; Hosseini, M.M.; Sweeny, T.D.; Mahoor, M.H. AffectNet+: A Database for Enhancing Facial Expression Recognition with Soft-Labels. IEEE Trans. Affect. Comput. 2026, 17, 784–800. [Google Scholar] [CrossRef]
- Gao, X.; Bansal, S.; Gowda, K.; Li, Z.; Nayak, S.; Kumar, N.; Coler, M. AMuSeD: An Attentive Deep Neural Network for Multimodal Sarcasm Detection Incorporating Bi-modal Data Augmentation. IEEE Trans. Affect. Comput. 2026, 17, 900–912. [Google Scholar] [CrossRef]
- Ji, Y.; Wang, S.; Xu, R.; Chen, J.; Quan, Y.; Jiang, X.; Deng, Z.; Liu, J. Hugging Rain Man: A Novel Facial Action Units Dataset for Analyzing Atypical Facial Expressions in Children with Autism Spectrum Disorder. IEEE Trans. Affect. Comput. 2025, 16, 2287–2302. [Google Scholar] [CrossRef]
- Guo, X.; Rodriguez, A.C.I.; Wang, C.; Rundensteiner, E.A.; Liu, S. DEPRESS: Dataset on Emotions, Performance, Responses, Environment, and Satisfaction during COVID-19. Sci. Data 2026, 13, 331. [Google Scholar] [CrossRef] [PubMed]
- Vaz, P.J.; Rodrigues, J.M.F.; Cardoso, P.J.S. Affective Computing Databases: In-Depth Analysis of Systematic Reviews and Surveys. IEEE Trans. Affect. Comput. 2025, 16, 537–554. [Google Scholar] [CrossRef]
- Wang, S.; Wu, W.; Li, Y.; Xu, Y.; Lyu, Y. MIANet: Bridging the Gap in Crowd Density Estimation with Thermal and RGB Interaction. IEEE Trans. Intell. Transp. Syst. 2025, 26, 254–267. [Google Scholar] [CrossRef]
- Sadiq, T.; Omlin, C.W. Sensing in Smart Cities: A Multimodal Machine Learning Perspective. Smart Cities 2026, 9, 3. [Google Scholar] [CrossRef]
- Altuwairqi, K.; Jarraya, S.K.; Allinjawi, A.; Hammami, M. Student behavior analysis to measure engagement levels in online learning environments. Signal Image Video Process. 2021, 15, 1387–1395. [Google Scholar] [CrossRef] [PubMed]
- Lemos, M.; Cardoso, P.J.; Rodrigues, J.M. MiE: A microscopic model for real-time group engagement estimation using gaze and posture. J. Comput. Sci. 2026, 96, 102856. [Google Scholar] [CrossRef]


| Sensor Type | Behavior | Satisfaction |
|---|---|---|
| RGB-D | 7 | 14 |
| Thermal Cameras | 1 | 2 |
| LiDAR | 1 | 0 |
| mmWave Radar | 5 | 0 |
| Microphone (Arrays) | 0 | 6 * |
| Multimodal | 20 | 4 |
| Total | 34 | 26 |
| Constructs (KPIs) | References |
|---|---|
| Commitment (proxy) (V) | [38,40] |
| Conversion (proxy) (V) | [38,40] |
| Density Maps (D) | [8,18,19,31,32,40,89] |
| Engagement (P) | [11,29,35,38,39,40,57] |
| Feedback (single-source proxy) (V) | [35] |
| Motion Patterns (D) | [16,19,20,21,23,24,25,26,28,30,31,34,35,36,37,38,40,41,90] |
| Odd Behavior (P) | [13,14,17,27,31,33,34,36,90] |
| Person Counting (D) | [8,12,15,18,19,20,21,22,24,28,31,32,40,89,90] |
| Retention (proxy) (V) | [38,40] |
| Social Interactions (P) | [11,28,30,35,90] |
| Attention (P) | [44,49,53,57] |
| Conversation (P) | [29,35,39,62,64,68] |
| Dwell Time (D) | [12,15,19,38,40] |
| Emotion (P) | [44,45,46,47,48,49,50,51,52,53,54,55,56,57,58,59,60,61,62,63,64,65,66,67,68] |
| Sentiment (P) | [45,46,47,48,49,52,53,57,60,61,62,63,64,66,67,68,69] |
| KPI | RGB-D | Thermal | LiDAR | mmWave | Audio | Multimodal |
|---|---|---|---|---|---|---|
| Commitment | – | – | – | – | – | [38,40] |
| Conversion | – | – | – | – | – | [38,40] |
| Density Maps | [8,31,32,40,89] | [8,18,31,32,89] | [19,31] | – | – | [8,31,32,40,89] |
| Engagement | [11,29,35,38,39,40,57] | – | – | – | [29,35,39] | [29,35,38,39,40] |
| Feedback | [35] | – | – | – | [35] | [35] |
| Motion Patterns | [16,25,28,30,31,34,35,36,37,38,40,41,90] | [28,30,31,90] | [19,25,31,36,37,90] | [20,21,23,24,28,36,37,41,90] | [30,34,35,38,40,90] | [25,26,28,30,31,34,35,36,37,38,40,41,90] |
| Odd Behavior | [13,14,17,27,31,33,34,36,90] | [27,31,36,90] | [31,36,90] | [36,90] | [34,90] | [27,31,34,36,90] |
| Person Counting | [8,12,15,28,31,32,40,89,90] | [8,18,28,31,32,89,90] | [19,31,90] | [20,21,22,24,28,90] | [40,90] | [8,28,31,32,40,89,90] |
| Retention | – | – | – | – | – | [38,40] |
| Social Interactions | [11,28,30,35,90] | [28,30,90] | [90] | [28,90] | [35,90] | [28,30,35,90] |
| Attention | [44,49,53,57] | – | – | – | – | – |
| Conversation | [29,35,39,68] | – | – | – | [29,35,39,62,64,68] | [29,35,39,68] |
| Dwell Time | [12,15,40] | – | [19] | – | – | [38,40] |
| Emotion | [44,45,46,47,48,49,50,51,52,53,54,55,56,57,59,67,68] | [58,59] | – | – | [60,61,62,63,64,65,66,67,68] | [58,59,66,67,68] |
| Sentiment | [45,46,47,48,49,52,53,57,67,68,69] | – | – | – | [60,61,62,63,64,66,67,68,69] | [67,68,69] |
| Construct | RGB-D | Thermal | LiDAR | mmWave | Audio | Multimodal |
|---|---|---|---|---|---|---|
| Commitment (V) | – | – | – | – | – | ∘ |
| Conversion (V) | – | – | – | – | – | ∘ |
| Density Maps (D) | • | • | ∘ | – | – | • |
| Engagement (P) | • | – | – | – | • | • |
| Feedback (V) | ∘ | – | – | – | ∘ | ∘ |
| Motion Patterns (D) | • | ∘ | ∘ | • | ∘ | • |
| Odd Behavior (P) | • | • | ∘ | ∘ | ∘ | • |
| Person Counting (D) | • | • | ∘ | • | ∘ | • |
| Retention (V) | – | – | – | – | – | ∘ |
| Social Interactions (P) | • | ∘ | ∘ | ∘ | ∘ | • |
| Attention (P) | • | – | – | – | – | – |
| Conversation (P) | ∘ | – | – | – | • | ∘ |
| Dwell Time (D) | • | – | • | – | – | ∘ |
| Emotion (P) | • | • | – | – | ∘ | • |
| Sentiment (P) | ∘ | – | – | – | • | • |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Walid, A.; Solá, D.; Martins, J.A.; Cardoso, P.J.S.; Rodrigues, J.M.F. Multimodal Sensing for Live-Stand Analytics: A Design-Oriented Literature Synthesis and Reference Architecture. Appl. Sci. 2026, 16, 7286. https://doi.org/10.3390/app16147286
Walid A, Solá D, Martins JA, Cardoso PJS, Rodrigues JMF. Multimodal Sensing for Live-Stand Analytics: A Design-Oriented Literature Synthesis and Reference Architecture. Applied Sciences. 2026; 16(14):7286. https://doi.org/10.3390/app16147286
Chicago/Turabian StyleWalid, Abdellah, David Solá, Jaime A. Martins, Pedro J. S. Cardoso, and João M. F. Rodrigues. 2026. "Multimodal Sensing for Live-Stand Analytics: A Design-Oriented Literature Synthesis and Reference Architecture" Applied Sciences 16, no. 14: 7286. https://doi.org/10.3390/app16147286
APA StyleWalid, A., Solá, D., Martins, J. A., Cardoso, P. J. S., & Rodrigues, J. M. F. (2026). Multimodal Sensing for Live-Stand Analytics: A Design-Oriented Literature Synthesis and Reference Architecture. Applied Sciences, 16(14), 7286. https://doi.org/10.3390/app16147286

