Next Article in Journal
Physicochemical and Gamma-Spectrometric Characterization of Legacy Liquid Radioactive Waste from the BN-350 Reactor Facility
Previous Article in Journal
Dimension-Constrained Organizational Relay for Long-Context LLMs
Previous Article in Special Issue
RBD-YOLOv10: A Lightweight Small-Object Detector for Laser-Tracking Cooperative Targets
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Review

The HEART Framework for LLM-Enabled Socially Assistive Robots in Healthcare: A PRISMA-Informed Structured Review

by
Tihomir Orehovački
Faculty of Informatics, Juraj Dobrila University of Pula, Zagrebačka 30, 52100 Pula, Croatia
Appl. Sci. 2026, 16(16), 7904; https://doi.org/10.3390/app16167904
Submission received: 22 June 2026 / Revised: 31 July 2026 / Accepted: 5 August 2026 / Published: 7 August 2026
(This article belongs to the Special Issue Artificial Intelligence and Its Application in Robotics, 2nd Edition)

Abstract

Large language models (LLMs) are expanding the capabilities of socially assistive robots (SARs) through natural dialogue, personalisation, multimodal reasoning, retained interaction context, and adaptive behaviour in healthcare. Integrating generative language models into robots, however, complicates evaluation because fluent output may exaggerate perceived competence and increase the risks of hallucination, overtrust, privacy exposure, relationship dependency, and unsafe reliance on advice or actions. This PRISMA-informed review synthesises healthcare robotics, human–robot interaction, LLM-enabled systems, ethics, implementation, and care delivery. Database searches returned 128 records, of which 110 were unique after deduplication. Supplementary retrieval and assessment yielded 85 substantive sources spanning background mapping, primary analysis, and governance. Studies focused mainly on feasibility, usability, acceptability, dialogue quality, and short-term engagement, whereas longitudinal safety, governance of retained interaction context, comparative effectiveness, workflow integration, and sustained healthcare value received limited attention. These gaps indicate that evaluation of LLM-enabled SARs must account for physical presence, social role, interaction memory, and potential actions rather than focus on conversational performance alone. The review therefore proposes HEART, a healthcare-specific evaluative architecture comprising Human-Centred Communication, Ethical and Trustworthy Deployment, Adaptive and Embodied Intelligence, Relationship Continuity, and Translational Healthcare Value. HEART uses boundary rules, operational indicators, qualitative labels, and non-additive deployment gates to separate evaluative domains, define assessable outcomes, summarise reported support, and prevent strengths in one area from masking critical safety or governance failures. Future research should validate HEART through longitudinal and comparative assessment of hallucination severity, language-to-action safety, long-term effects, equity, and post-deployment monitoring.

1. Introduction

Socially assistive robots (SARs) are embodied systems designed to support users through social interaction, guidance, companionship, coaching, monitoring, or motivation. They are studied in institutional, community, rehabilitation, and home-based care settings, where their value depends on technical reliability, verbal and non-verbal communication, personalisation, engagement, and fit with care relationships [1,2,3].
Despite this breadth, evidence of healthcare value remains uneven. The literature commonly emphasises usability, acceptability, engagement, perceived usefulness, comfort, or short-term interaction quality, whereas sustained clinical, psychosocial, relational, organisational, and implementation outcomes are less well established [4,5,6]. Trust is also underexamined when assessment relies on subjective measures and brief or controlled encounters rather than calibration in everyday care [7]. These gaps are especially consequential for users whose age, health status, or cognitive, sensory, or functional impairments may increase vulnerability, the risk of dependency or privacy loss, and asymmetries of knowledge and authority.
LLM integration broadens the capabilities of socially assistive robots while making their evaluation more complex. An LLM-enabled SAR incorporates an LLM or another generative language model that performs a comparable role in one or more functions, including dialogue generation, language understanding, personalisation, retention of interaction context, multimodal interpretation, planning, or action selection. Integration may be limited to conversational functions or extend to system-level memory, multimodal reasoning, planning, and action-linked autonomy.
These functions may support open-ended conversation, contextual interpretation, multimodal interaction, and more flexible adaptation than earlier scripted or rule-based designs [8,9,10]. At the same time, fluent output may make a robot appear more knowledgeable, empathic, socially competent, or clinically capable than its demonstrated performance warrants. Risks of hallucination, misleading capability claims, overtrust, privacy exposure, unclear accountability, and unsafe reliance therefore become central evaluation concerns in healthcare [11,12]. When generated language is coupled with social presence and possible physical action, failures can affect behaviour, care relationships, and safety rather than communication alone.
Prior reviews cover healthcare social robot characteristics, technical requirements, deployment settings, LLM-driven HRI capabilities, trust, acceptance, ethics, and implementation [2,8,9]. Their contributions remain dispersed across the technical, clinical, relational, and governance literatures. Taken together, these reviews do not yet provide an integrated approach for judging systems in which generative communication, memory, adaptation, social credibility, physical agency, and healthcare accountability interact.
In response, this study proposes HEART as a healthcare-specific interpretive architecture rather than a psychometric scale or composite score. It organises evaluation across five distinct but interacting dimensions: Human-Centred Communication, Ethical and Trustworthy Deployment, Adaptive and Embodied Intelligence, Relationship Continuity, and Translational Healthcare Value. Explicit boundary rules assign primary evaluative objects, operational indicators translate the dimensions into assessable outcomes, and non-additive gates prevent strengths in one domain from masking critical safety or governance deficits in another.
Three research questions guide the analysis:
RQ1. What evaluative challenges emerge across the literature on healthcare robotics, socially assistive robots, human–robot interaction, and LLM-enabled robotic systems?
RQ2. How do LLMs and generative AI change the evaluation of socially assistive robots in healthcare, particularly in relation to communication, trust, embodiment, autonomy, continuity, and healthcare value?
RQ3. How can these evaluative challenges be synthesised into an integrative framework for assessing LLM-enabled socially assistive robots in healthcare?
To address these questions, the study combines PRISMA-informed evidence identification and selection with framework-oriented analysis of retained sources. The evidence base is organised into three functional corpora that distinguish background mapping, primary analytical studies, and governance or regulatory interpretation. This structure supports the development of an evaluative framework rather than pooled estimation of intervention effects. The resulting HEART architecture is intended to inform study design, comparative assessment, stakeholder involvement, implementation decisions, and future empirical validation.

2. Related Work

The discussion of prior work moves from healthcare robotics and HRI to population-specific, ethical, evaluative, and generative AI perspectives. This sequence identifies the unresolved problems that motivate HEART. Table A1, Table A2, Table A3, Table A4 and Table A5 in Appendix A provide detailed gap maps.

2.1. Healthcare Robotics, Human–Robot Interaction, and Embodied AI

Prior work in healthcare robotics indicates that value arises from situated interaction rather than technical capability alone. HRI reviews emphasise social perception, behavioural adaptation, communication, morphology, interaction modes, and service roles across hospital, residential, home, rehabilitation, and professional-support contexts [2,3,13]. Broader intelligent robotics research covers AI, sensing, privacy, ethics, and responsible implementation, but gives less attention to healthcare-specific vulnerability, continuity, and clinical value [14]. Work on embodied AI foregrounds perception, planning, decision-making, and action, whereas clinical HRI contributes evidence on trust, therapeutic alliance, personalisation, and responsiveness while identifying persistent methodological gaps and difficulties in translation to sustained clinical practice [15,16]. This foundation does not yet explain how evaluation changes when generative language, social presence, autonomy, and care-related risk converge.

2.2. Older Adults, Dementia, Frailty, and Long-Term Care Contexts

In age-related and long-term care, robotic support must be judged within everyday routines rather than isolated encounters. Care robots operate within networks of users, families, professionals, institutions, homes, and technical services [17], while related studies emphasise ageing in place, cognitive and daily living support, monitoring, companionship, autonomy, caregiver burden, and workflow integration [18,19]. Reviews of dementia and frailty care commonly report feasibility and acceptability, but effects on cognition, neuropsychiatric symptoms, quality of life, and frailty outcomes remain limited or inconclusive [20,21]. Heterogeneity in design, duration, measures, and data collection further weakens claims of sustained efficacy or effectiveness [5]. Uptake also depends on reliability, individual needs, organisational context, maintenance, staffing, and integration into routine care [18,19,22]. For HEART, the implication is that personalisation, memory, and adaptation require longitudinal assessment; short-term acceptance alone cannot establish durable benefit.

2.3. Paediatric Care, Clinical Workflows, and Nursing Practice

Paediatric and professional care settings raise a shared requirement: robotic support must remain aligned with clinical roles, emotional sensitivity, and staff responsibility. Child-facing interventions span autism therapy, procedural pain, emotional regulation, and hospital-based psychological support, using robots for social learning, communication, imitation, distraction, motivation, and developmentally appropriate engagement [23,24,25]. Reported reductions in pain, anxiety, distress, or negative affect vary by outcome, procedure, design, and measurement approach [24,25,26]. Active participation, guided breathing, motivational support, and physical comfort also shape responses, showing that robot presence alone is insufficient [27]. From the professional perspective, potential benefits for workload, efficiency, and patient well-being are balanced by concerns about reliability, role boundaries, training, privacy, consent, professional identity, and equitable access [28,29]. Obstetric and neonatal nurses additionally prioritise safety, sterilisation, precision, alarms, data transfer, navigation, usable guidance, and clear responsibility [30]. Overall, claims of value are most credible when linked to a defined care problem rather than inferred from novelty or general approval.

2.4. Trust, Acceptance, Ethics, Autonomy, and Evaluation

Positive user reactions are difficult to interpret when a system can influence trust, autonomy, and care relationships. Variation in autonomy, interaction mode, responsibility, and care role complicates comparison across deployments [1]. Technology-acceptance models explain willingness to use a robot but cannot establish durable clinical, psychosocial, relational, or organisational value [4]. Trust is dynamic and potentially repairable, yet it is often measured through short-term self-report rather than calibrated against actual performance in vulnerable populations [7,31,32]. Ethical analyses similarly warn that social presence and efficiency claims do not justify use when users may misunderstand limitations, adapt to device constraints, or lose control over care relationships [33,34]. Longer-term consequences for agency, behaviour, participation, and well-being remain underexamined [35]. Ethical concerns also include consent, privacy, replacement of human care, inequitable access, infantilisation, dignity, dependency, deception, and data protection [36,37]. Measurement instruments show uneven psychometric quality and limited validation for specific clinical populations [6]. Fluent generative dialogue intensifies these problems by introducing risks of hallucination, opacity, deceptive behaviour, unsafe reliance, and diffuse responsibility [11]. The ethical baseline for HEART is therefore clear: positive attitudes become safety-relevant when language is delivered by a socially credible robot in a care setting.

2.5. Large Language Models, Generative AI, and Robotic Autonomy

Integrating generative models into healthcare SARs couples dialogue, reasoning, planning, personalisation, and physical action more closely than scripted systems do. LLM-driven HRI supports natural language understanding, multimodal input, contextual sensing, high-level reasoning, plan generation, adaptive alignment, and task execution [8,10]. These capabilities can make assistance more flexible, but may also blur the distinction between demonstrated ability and persuasive social performance, raising concerns about bias, trust, and responsible adoption [10,38]. In robotic autonomy, language commands can be connected to navigation, manipulation, semantic planning, voice interaction, and adaptive execution, with trade-offs in latency, scalability, privacy, reliability, and real-time responsiveness [8,39]. Protocol-level work accordingly treats the intersection of generative AI and healthcare social robots as an emerging area requiring systematic mapping of models, platforms, outcomes, implementation, and risk [9]. Because misleading capability claims or credible misinformation can affect patient behaviour, safety, and accountability, evaluation must consider communication, planning, physical agency, privacy, and governance together [11,38].

2.6. From Existing Evidence to the HEART Framework

No single literature stream covers the full evaluative problem. Healthcare HRI contributes evidence on communication, physical agency, and clinical fit; care studies add perspectives on continuity, workflow, and professional responsibility; ethics and measurement scholarship addresses autonomy and calibrated trust; and LLM research introduces open-ended generation, persistent memory, planning, and persuasive social performance.
Adjacent approaches help define the remaining gap. Biometric-wearable research links continuous sensing with behavioural effects but does not address socially credible generative dialogue or language-to-action coupling [40]. A user-centred mental healthcare SAR framework emphasises autonomy, emotional experience, participation, personalisation, and long-term trust, but does not operationalise LLM hallucination or execution risk [41]. Clinical LLM-safety work contributes error taxonomies and severity-sensitive assessment of hallucinations and omissions that can be adapted to patient-facing robotic dialogue [42].
HEART brings these contributions together within a single evaluative architecture. Table 1 maps each literature stream to the relevant dimensions and identifies the unresolved problems addressed by the framework.

3. Methods

3.1. Review Design

This review used a PRISMA-informed structured design that combined transparent source identification with framework-oriented analysis of LLM-enabled SARs. The aim was to organise a rapidly developing and heterogeneous literature rather than estimate pooled intervention effects. Established guidance on scoping, structured mapping, and review reporting informed the design [43,44,45,46]. The source base included reviews, empirical and technical studies, pilots, protocols, design work, preprints, and normative material across healthcare robotics, HRI, generative AI, ethics, implementation, and clinical evaluation.
PRISMA 2020 and PRISMA-ScR informed reporting of identification, deduplication, source-level assessment, inclusion, and flow documentation [45,46]. The qualifier PRISMA-informed reflects adaptation of these principles to a structured, framework-oriented review rather than a registered systematic review or meta-analysis. Accordingly, no pooled effect estimates or formal risk-of-bias meta-assessment were undertaken.

3.2. Information Sources

Web of Science Core Collection and Scopus supplied the controlled database coverage because both index the literature across healthcare, robotics, HRI, computer science, engineering, and social science. Google Scholar supported targeted supplementary identification of methodological references, arXiv material, governance or regulatory documents, and selected non-LLM SAR and healthcare-robot studies that were insufficiently represented in the controlled results. The purposive comparative baseline was not intended as an exhaustive search of the pre-LLM SAR literature. Publisher, arXiv, institutional, and governance websites served only as retrieval platforms after discovery. Citation tracking was not undertaken.
Sources identified through the supplementary route were interpreted according to evidentiary function: preprints informed emerging claims, whereas governance documents supported normative and regulatory interpretation rather than empirical effectiveness claims (Section 3.9).
Exploratory queries conducted between 18 and 24 March 2026 calibrated terminology across healthcare robotics, HRI, LLMs, generative AI, ethics, implementation, and clinical evaluation. The main Web of Science and Scopus searches followed on 25–27 March. Duplicate removal and initial relevance checking took place from 28 March to 12 April, and consolidated source-level assessment and structured charting continued from 13 April to 5 May. Thematic analysis and framework development were conducted from 6 to 28 May, followed by manuscript drafting and revision between 29 May and 9 June. A targeted Google Scholar update on 10 June captured recent LLM-HRI, methodological, arXiv, governance, and regulatory sources. The literature was locked on 14 June 2026.
Coverage focused primarily on publications from 2021 to 2026, reflecting the recent emergence of generative-language and embodied-AI robotic systems. Earlier material was retained only when it supplied methodological, ethical, conceptual, or review-design foundations, including scoping-review methods, PRISMA reporting, and thematic synthesis.

3.3. Search Strategy

Four concept blocks structured the controlled database search: (1) social, socially assistive, and healthcare robots; (2) large language models, generative AI, and multimodal AI; (3) healthcare contexts; and (4) evaluation, ethics, trust, autonomy, implementation, safety, and clinical translation. Equivalent Boolean logic was implemented in Web of Science Core Collection Advanced Search and Scopus, with field syntax adapted to each platform. Broader WoS Smart Search and Scopus All Fields variants were tested during calibration but excluded from the numerical flow. Shorter Google Scholar queries targeted the supplementary categories described in Section 3.2 and selected non-LLM comparators. This iterative formulation followed scoping and evidence-mapping approaches for heterogeneous fields [43,44,47].
(“socially assistive robot*” OR “social robot*” OR “healthcare robot*” OR “care robot*” OR “human-robot interaction”) AND (“large language model*” OR LLM OR “generative AI” OR ChatGPT OR “multimodal AI”) AND (healthcare OR clinical OR hospital OR “older adult*” OR dementia OR pediatric OR nursing OR rehabilitation) AND (trust OR ethics OR acceptance OR autonomy OR evaluation OR implementation OR safety).
Database interfaces, field configurations, search dates, displayed record yields, and the supplementary Google Scholar queries are reported in Supplementary Table S5. Audited counts derive only from the two controlled database configurations. Displayed Google Scholar totals were treated as approximate platform estimates; screening was targeted rather than exhaustive, and exact route-level additions and overlap were not retained.
Platform-specific syntax and wildcard conventions were applied to this expression in the audited searches. Materials located through Google Scholar were retrieved from the relevant publisher, arXiv, institutional, or regulatory website. Zotero Web Library was used for reference management and duplicate checking.

3.4. Study Selection and Evidence Flow

The selection flow combines an audited database-search branch with targeted Google Scholar retrieval and a preserved consolidated candidate-source log. Figure 1 distinguishes these pathways and marks the transparency boundary created by the absence of exact counts for supplementary additions, overlap, and attrition.
The two database searches yielded 128 records: 47 from Web of Science Core Collection via Advanced Search without a field restriction and 81 from Scopus through title, abstract, and keyword queries. After 18 duplicates were removed, 110 unique items remained. The supplementary route added sources from the categories described in Section 3.2, including selected non-LLM comparators used as a purposive baseline. Consolidation across routes produced a preserved coding log of 110 candidate sources. The matching totals are coincidental and do not indicate identical membership because additions and removals occurred without a retained route-level transition log. Source-level assessment excluded 18 candidates and retained 92 references. Of the exclusions, nine were peripheral to the healthcare SAR/LLM scope, three were overly broad or insufficiently specific to physical healthcare systems, four had unsuitable publication types or limited evidentiary contributions, and two were superseded or redundant. Supplementary Table S6 provides the source-level reasons.
Seven of the 92 retained references were methodological sources used to justify review design, PRISMA-informed reporting, scoping logic, data charting, and thematic synthesis. They were cited but not included in substantive coding. This left 85 sources distributed across Corpus A (n = 35), Corpus B (n = 35), and Corpus C (n = 15). Three contextual references added after the literature lock for framework comparison and metric operationalisation [40,41,42] were also excluded from corpus counts.
Table 2 differentiates the controlled searches that generated the audited counts from exploratory interface variants, targeted Google Scholar identification, source-platform retrieval, and post-lock contextual updating.

3.5. Corpus Assembly and Classification

Analytical role determined the assignment of the 85 substantive sources to three complementary corpora. Corpus A contained reviews, surveys, and background or gap-mapping studies; Corpus B contained primary empirical, technical, pilot, protocol, and design-oriented work; and Corpus C contained governance, ethics, privacy, accountability, and regulatory material. The three groups supported landscape mapping, system-level analysis, and interpretation of trustworthy deployment, oversight, safety, and healthcare implementation, respectively.
Each source received one primary corpus assignment based on the function used most directly in the analysis. Mixed contributions retained secondary relevance through HEART tagging, allowing a paper to inform several dimensions without being counted in more than one corpus.

3.6. Eligibility Criteria

A source was eligible when it made a substantive contribution to healthcare robotics, socially assistive or generative AI-enabled robotic interaction, HRI, embodied autonomy, implementation, evaluation, safety, ethics, trust, or regulatory governance. Both topical relevance and analytical function were considered because the review combined heterogeneous source types.
Reviews and broad mappings of relevant research were assigned to Corpus A. Physical-robot studies involving care, generative AI, adaptation, multimodality, autonomy, or evaluation were assigned to Corpus B. Normative or regulatory material on trustworthy healthcare deployment, data protection, safety, or accountability was assigned to Corpus C.
Exclusions covered sources outside the healthcare, HRI, or robotics scope; disembodied chatbot papers without relevance to physical systems; industrial automation without human-facing care implications; and publications that did not contribute to the analytical objectives. Purely technical LLM work was retained only when it informed robotic autonomy, multimodal grounding, privacy, safety, or responsible deployment.

3.7. Data Extraction

A structured template informed by scoping and evidence-synthesis methods [43,47] was used to chart the retained sources. Recorded fields covered bibliographic details, study design, robot configuration, care or interaction context, target population, AI or LLM component, interaction modality, duration, evaluation methods, outcomes, ethical concerns, implementation issues, limitations, and implications for real-world healthcare use.
Empirical-study extraction emphasised characteristics that shaped interpretation: laboratory, simulated, clinical, educational, residential, or home setting; single-session, multi-session, or longitudinal interaction; participant group; and measurement through self-report, observation, behaviour, system performance, clinical indicators, or qualitative feedback.
Technical studies were charted according to the role of the LLM or generative component, including dialogue, planning, perception, action selection, memory, personalisation, multimodal grounding, and autonomous execution. Review-level extraction recorded the literature stream, principal conclusions, methodological limitations, gaps, and relevance to healthcare robots. Governance and ethics sources were examined for principles, risks, responsibilities, oversight mechanisms, and implications for trustworthy deployment.
The author performed duplicate checking, source-level assessment, corpus assignment, data extraction, HEART tagging, retrospective profiling, and thematic synthesis. No independent duplicate screening or coding was conducted. Explicit decision rules and the source-level crosswalks in Supplementary Tables S1 and S2 were used to document these judgements.

3.8. Descriptive Mapping of the Evidence Base

Rather than applying formal bibliometric methods, the review used variables tailored to corpus function. Review type and gap-mapping contribution described Corpus A; publication year, primary source orientation, LLM status, context, and HEART relevance described Corpus B; and governance, ethical, privacy, accountability, and regulatory role described Corpus C. HEART relevance was an overlapping thematic tag, whereas publication year, source orientation, AI/LLM status, and primary context were mutually exclusive within Corpus B.
Corpus B formed the primary analytical set of 35 sources. Thirty-one (88.6%) were published between 2024 and 2026. Fifteen (42.9%) explicitly involved an LLM, foundation model, small language model, GPT-based component, or generative AI, while 20 (57.1%) provided non-LLM SAR, healthcare-robotics, AI-enabled, or HRI baselines. These comparators were selected purposively from the controlled results or targeted Google Scholar searching. Their role was to supply established knowledge on care relationships, workflow, longitudinal use, trust, and implementation rather than represent an exhaustive review of pre-LLM SAR research.
Table 3 reports mutually exclusive descriptive variables for Corpus B. Overlapping HEART tags are reported separately in Supplementary Table S1 and compared by AI/LLM status in Section 4.7. Across the 35 sources, T was represented in 22, R in 17, E in 15, H in 12, and A in 10; these values describe the coded analytical corpus rather than the prevalence of topics in the wider field.

3.9. Evidence Status and Claim Function

Because the source base was heterogeneous, a pooled risk-of-bias assessment was unsuitable. Interpretation instead used two descriptors. Evidence status distinguishes among completed peer-reviewed empirical studies, peer-reviewed technical or design work, reviews and syntheses, protocols or preprints, and governance or regulatory sources. Claim function differentiates observed user or clinical outcomes, observed technical performance, proposed designs or methods, planned evaluations, and normative requirements. Section 4 and Section 5 identify protocols, preprints, technical demonstrations, and governance material by function whenever maturity affects interpretation; Supplementary Table S2 provides source-level descriptors.
This distinction preserves evidentiary differences without forcing heterogeneous materials onto a single quality hierarchy. Completed empirical studies support observed outcomes, technical work supports capability and failure-mode claims, protocols describe planned evaluation, preprints provide preliminary evidence, and governance documents establish normative requirements rather than effectiveness.

3.10. Thematic Synthesis

Recurring evaluative concerns were identified through an iterative process informed by qualitative evidence synthesis and thematic analysis [48,49]. Descriptive codes were expanded into broader analytical themes while preserving corpus function. Corpus A contributed landscape and gap codes; Corpus B contributed study-, system-, technical-, and interaction-level material; and Corpus C informed ethical, privacy, governance, and regulatory interpretation. This procedure enabled comparison without treating review-level, empirical, technical, and normative sources as equivalent.
Sources were initially coded by contribution, application area, evaluation focus, and limitations. Related codes were then grouped into themes such as communication quality, trustworthiness, physical interaction, adaptive behaviour, personalisation, autonomy, continuity, implementation, and healthcare value. Cross-corpus comparison was used to identify fragmentation and gaps relevant to LLM-enabled SAR evaluation.
Analysis centred on the changes introduced when an LLM is incorporated into a socially assistive robot. Acceptance, usability, and engagement were considered alongside social credibility, trust calibration, autonomy, responsibility, safety, and integration into care. Particular attention was paid to tensions between apparent capability and demonstrated reliability, short-term engagement and sustained value, personalisation and user control, conversational fluency and clinical appropriateness, and technical autonomy and professional accountability.

3.11. Framework Refinement and Methodological Boundaries

Initial HEART categories were compared with the literature streams in Table 1 and the cross-corpus themes derived from the analysis. This comparison assessed whether the proposed dimensions captured recurring concerns in both the related-work mapping and the source-based review.
A dimension was retained when it addressed a recurring limitation, appeared across more than one corpus function, and captured an issue made more complex by generative language in a socially assistive healthcare robot. Applying these criteria refined HEART from an initial literature map into an integrative interpretive framework. The resulting framework is an evaluative structure rather than a validated measurement instrument.
The method prioritises transparent framework development over effect estimation or exhaustive coverage. Documentation preserves the exact controlled-database counts, a consolidated source-level exclusion log, retained methodological references, substantive corpus construction, and an explicit boundary around supplementary routes. The non-LLM subset was assembled purposively as a comparative baseline, not as a comprehensive review of the earlier SAR literature. Section 6 addresses the remaining limitations.

4. The HEART Framework

HEART provides a healthcare-specific evaluative architecture for LLM-enabled socially assistive robots. Its five dimensions—abbreviated H, E, A, R, and T—are examined below through their primary objects, evidentiary requirements, and deployment implications.
The framework treats generated language, social credibility, memory, physical agency, governance, and care outcomes as interacting features of a healthcare intervention. HEART complements technical validation, clinical evaluation, ethical governance, implementation science, and regulatory oversight, but remains interpretive rather than a psychometrically validated scale or additive score.

4.1. Framework Logic, Boundaries, and Deployment Gates

The five dimensions are complementary but not interchangeable. Their primary objects are generated communication and user understanding (H), justified reliance and accountable deployment (E), coordination of language with perception and action (A), persistence and adaptation across encounters (R), and patient-, workflow-, organisational-, or system-level outcomes (T). Table 4 specifies the boundaries, LLM-specific failure modes, and deployment implications.
Potential overlap is resolved through a primary-object rule. Multimodal coherence belongs to H when the outcome is comprehension or communicative consistency and to A when the outcome is technical synchronisation of speech, sensing, gesture, or movement. Personalisation belongs to R when continuity and adaptation are assessed and to E when consent, data minimisation, access, deletion, or privacy are assessed. Secondary cross-links may represent causal or governance relationships, but the same observation is not counted twice in a HEART profile.
Hallucinated or unsupported content is assigned to H when the evaluated outcome is factual content, comprehensibility, communicative repair, or user understanding. It is assigned to E when the evaluated outcome is potential clinical harm, justified reliance, oversight, escalation, or deployment responsibility. The same event may create a secondary cross-link, but its primary coding follows the outcome under evaluation.
Figure 2 visualises this logic by linking the system under evaluation to the five primary objects specified in Table 4. The GATE markers show that E and A function as non-additive constraints: a critical failure in either dimension can block deployment even when other dimensions are favourable. The R and T rows additionally indicate evidence thresholds, requiring repeated interaction and outcomes beyond usability or acceptance, respectively.

4.2. H—Human-Centred Communication

Communication quality in healthcare cannot be inferred from fluency or conversational naturalness. H evaluates whether generated content is understandable, emotionally appropriate, context-sensitive, clinically bounded, and coherent with non-verbal behaviour. The central LLM-specific risk is that persuasive language can increase perceived competence despite weak factual support, absent clinical authority, or unreliable repair.
User and field studies document both potential and limitations. Studies with older participants indicate that users value continuity of relevant information, privacy controls, reminders, social connection, and empathic interaction. They also document interruptions, latency, repetition, superficiality, hallucinations, outdated information, confusion, frustration, and worry [50,51]. Comparative LLM-powered HRI research shows that robots can support connection-building and deliberative interaction, while participants also expect coherent non-verbal cues and identify weaknesses in logical communication [52].
Geriatric evaluations report fewer comprehension failures and improved interaction success, acceptability, and usability after LLM integration [53,54]. Technical and user-evaluation studies additionally identify turn-taking, back-channelling, speech quality, gesture synchronisation, persona conditioning, multilingual adaptation, and source validation as relevant design features [55,56,57,58,59]. Appropriate measures therefore include unsupported-claim rates, safe-response rates, repair success, explanation quality, and user understanding of limitations rather than fluency alone.

4.3. E—Ethical and Trustworthy Deployment

Safe and accountable deployment requires patients, caregivers, clinicians, and institutions to understand a robot’s capabilities and limits and to have defined means of governance, contestation, escalation, and oversight. E examines whether reliance is justified by demonstrated performance and whether privacy, consent, accountability, and human control are established before use. Trust should be calibrated to actual capability rather than maximised as an acceptance outcome [7,32].
LLM-enabled robots intensify this problem because fluent and confident outputs may be incorrect, incomplete, outdated, unsupported, or outside the robot’s intended role. A published ethical case study documents misleading capability claims in LLM-based care robots [11]. For operational evaluation, the clinical-safety framework of Asgari et al. distinguishes hallucination and omission frequency from potential clinical severity and classifies errors such as fabrication, negation, contextual distortion, and causal error [42]. Adapted to patient-facing SARs, E should therefore report rates of unsupported clinical claims and omissions, rates of major or critical errors, appropriate abstention and referral, and the gap between demonstrated reliability and user reliance.
Peer-reviewed ethical reviews and qualitative studies treat autonomy, consent, dignity, social presence, substitution of human care, dependency, equitable access, and stakeholder involvement as deployment requirements rather than optional design values [33,34,36,37,60,61]. These sources support normative and interpretive claims; they should not be read as evidence that a particular LLM-SAR has achieved safe deployment.
Privacy and governance are equally central because LLM-enabled robots may process speech, images, interaction histories, preferences, behavioural data, and health-related information. Technical and governance sources support data minimisation, privacy-aware memory, lifecycle monitoring, human oversight, change control, transparency, and accountable update procedures [12,62,63,64,65,66,67,68]. Privacy-sensitive personalisation belongs to E when the question concerns consent, collection, retention, access, deletion, or inference risk, and to R when the question concerns the accuracy and usefulness of personalisation across time.

4.4. A—Adaptive and Embodied Intelligence

When generated language can influence sensing, planning, movement, proximity, reminders, or action, the risk profile extends beyond dialogue quality. A evaluates the technical correctness and recoverability of that coupling. Fluent dialogue may mask an incorrectly grounded plan that produces inappropriate physical behaviour, delayed support, or unclear responsibility.
Studies of robotic autonomy show that LLMs can translate high-level language commands into lower-level control processes, support semantic planning, enable voice-based interaction, and integrate vision, speech, or proprioception. These capabilities also introduce problems of grounding, latency, reliability, error recovery, and safety [39]. Humanoid platforms illustrate how instructions, environmental observations, execution feedback, and human input may be combined for adaptive behaviour [69], while text-to-motion research demonstrates both the promise and limits of grounding GPT-based systems in movement [70]. Coordinated text and gesture generation adds a further requirement for verbal and non-verbal coherence [56].
User research adds an interactional requirement: language should align with gaze, gesture, timing, and physical presence because mismatches can reduce perceived competence or create confusion [52]. The Nadine system illustrates how affective behaviour and human-like recall can be incorporated into a social robot while raising governance questions about memory and user expectations [71]. Other LLM-integrated systems combine speech processing, summarisation of prior exchanges, persona conditioning, and multilingual adaptation, linking adaptive communication more closely to physical interaction and perception [57].
Clinical settings impose a higher threshold. A GPT-reinforced patient communication robot illustrates the need for validated knowledge sources, dependability, and interaction safeguards when a physical system addresses health-related topics [59]. Approaches based on clinical supervision likewise support personalisation without removing professional oversight or accountability [72]. Adaptive and Embodied Intelligence therefore requires evidence that language, perception, planning, movement, and action are grounded, reliable, recoverable after errors, and appropriate to the care context.

4.5. R—Relationship Continuity

Appropriateness over time cannot be inferred from a single encounter. R examines whether stored information and adaptation remain accurate, controllable, responsive to changing needs, and useful after novelty declines. Persistent context may support reminders and shared history, but it can also preserve false information, expose sensitive data, or imply a stable understanding that the system cannot sustain.
Older participants expect companion robots to recall previous exchanges, tailor interaction, protect privacy, provide control over learned data, and adapt to social context [50]. Open-domain systems may nevertheless fail to preserve coherent context or stable adaptation when interruptions, outdated responses, or hallucinations undermine confidence across sessions [51]. Multi-session studies suggest that repeated interaction can support engagement but still requires assessment of accuracy, adaptation, and stability over time [73]. Features such as persona conditioning and human-like recall may foster continuity while complicating authenticity, boundaries, and dependence [57,71].
Continuity also concerns cognitive support, well-being, and integration into daily routines. LLM-powered SARs are being explored for cognitive-health promotion, companionship, and adaptive engagement in elder care [55,58]. Collaborative drawing suggests that personalised creative activity may support social participation, although its value must persist beyond novelty [74]. Home and real-world studies report varied trajectories of adoption, scepticism, use, constraints, and outcomes rather than a uniform pattern [75,76,77].
Co-design work indicates that continuity should be shaped with users rather than imposed by designers. Older adults identify priorities involving companionship, autonomy, support, social engagement, and control over interaction [78,79], while personalised cognitive-support research emphasises alignment between robot functions, individual needs, and care goals [80]. A preprint review of privacy-preserving LLM methods further supports data minimisation, transparency, user control, and protection against inference risks in systems that retain or reuse personal information [81].

4.6. T—Translational Healthcare Value

Positive interaction outcomes do not by themselves establish healthcare value. T asks whether a robot addresses a defined care need without shifting hidden burdens to patients, caregivers, professionals, or institutions; usability, acceptance, engagement, and technical novelty remain necessary but insufficient indicators.
Evidence from geriatric and clinical settings illustrates this distinction. LLM-integrated robots may reduce comprehension failures, increase successful interactions, improve perceived usefulness, and support companionship, but short-term acceptability does not establish durable benefit, workflow fit, reliability, or safety [53,54]. Review-level work in age-related care reinforces the caution: feasibility is common, but benefits remain inconsistent across clinical and psychosocial outcomes [20,21]. Paediatric results likewise vary across procedures, measures, and intervention designs [24,25,26].
Organisational fit is equally important. Professional perspectives indicate that adoption depends on patient benefit, role clarity, workload, safety, accountability, training, privacy, and integration into clinical routines [28]. Nursing and care studies add staff capacity, maintenance, sterilisation, alarms, precision, navigation, communication, and responsibility boundaries [19,29,30]. Work in geriatric and hospital settings further highlights institutional readiness, technical reliability, contextual fit, and sustained support [3,82].
Two peer-reviewed protocols describe planned effectiveness and cost-effectiveness evaluations rather than observed outcomes [83,84]. Additional intervention and implementation studies suggest possible contributions to cognitive support, residential care, group activities, physical activity, clinical nursing, and patient engagement [85,86,87,88]. Claims of translational value therefore require stronger comparative, longitudinal, economic, workload, and implementation evidence.
Appendix A, Table A6 summarises the detailed evidence themes, representative sources, and unresolved evaluative gaps underpinning each HEART dimension.

4.7. Evidence Landscape Across the HEART Dimensions

Once the dimensions and their evaluative scope have been established, Corpus B reveals how analytical coverage is distributed. The 15 LLM-explicit and 20 non-LLM sources served complementary functions. LLM-explicit studies primarily informed generative communication, social credibility, stored context, multimodal coordination, and language-to-action risks, whereas baseline work contributed more strongly to longitudinal use, workflow integration, implementation, and care outcomes.
Figure 3 compares the overlapping HEART tags within the two subsets. H appeared in 10 of 15 LLM-explicit sources (66.7%) and 2 of 20 baseline sources (10.0%), whereas T appeared in five LLM-explicit sources (33.3%) and 17 baseline sources (85.0%). E and R were more evenly distributed, while A was more common in the LLM-explicit subset. These coding patterns describe the analytical corpus rather than topic prevalence across the wider field.
The comparison reveals three concentrations: recent LLM-enabled work on communication, social credibility, and embodied adaptation; age-related and geriatric studies of continuity, companionship, and repeated use; and non-LLM baseline research on implementation, workflow, and translational healthcare value. Few LLM-explicit studies combine generative-system testing with longitudinal, organisational, or outcome evaluation, leaving the connection between technical capability and real-world benefit underdeveloped.

4.8. Operationalising HEART

To make HEART usable in study design, assessment must operate at model, interaction, physical-system, longitudinal, and healthcare levels. Calibrated trust is not equivalent to a high trust score; it is the correspondence between demonstrated reliability, the user’s belief about that reliability, and willingness to act on the robot’s output. Hallucination assessment should report both frequency and potential clinical severity, distinguishing fabricated, negated, contextual, or causal errors and clinically important omissions [42]. Table 5 translates the dimensions into assessable constructs, possible methods, relevant stakeholders, temporal levels, and illustrative outcomes.
User-facing work with older adults shows why model-level metrics must be combined with comprehension, privacy, and control over retained data [50,51]. Comparative LLM-powered HRI also indicates that social credibility and non-verbal cues shape perceived competence [52].
A worked prospective example in Supplementary Table S3 presents a 12-week evaluation of a medium-autonomy LLM-enabled SAR in residential care. The scenario limits the robot to bounded non-clinical dialogue, routine reminders, and companionship; clinical recommendations, care-plan modification, medication-related action, and unsupervised physical assistance are prohibited. The comparison is between an LLM-enabled SAR plus usual care and a scripted SAR plus usual care. The study design links dialogue safety, trust and privacy safeguards, controlled memory, hazard and override testing, repeated use, workflow, and healthcare-value outcomes to explicit decision gates. As a prospective template, it does not provide evidence of effectiveness. Supplementary Table S4 offers a reusable record for study planning, evidence profiling, and deployment review.

4.9. Retrospective Application to Two Empirical LLM-SAR Evidence Strands

The two selected evidence strands permit an illustrative retrospective application of the operational scheme. Table 6 compares companion-robot studies with older adults by Irfan et al. [50,51] and two-wave geriatric-unit evaluations by Blavette et al. [53,54]. The labels—supported, partial, concern, and not assessed—summarise the reported results and scope rather than numerical HEART scores.
The resulting profiles differ in ways relevant to evaluation design. Geriatric-unit studies provide stronger support for communication and early feasibility in a care setting, whereas companion-robot research reveals more explicit concerns involving hallucination, privacy, memory, and user expectations. Neither strand establishes translational effectiveness or comprehensive language-to-action safety; “not assessed” denotes an evidence gap rather than poor system performance.

4.10. Research Question Synthesis

The reviewed literature answers RQ1 by identifying a multidimensional evaluation problem spanning communication, justified reliance, language-to-action coordination, continuity, and healthcare value. With respect to RQ2, LLM integration connects generative language and retained context with social credibility, adaptation, and possible physical action under healthcare accountability. HEART addresses RQ3 by providing a structured basis for determining whether evidence is sufficient, incomplete, or blocked by a critical safety or governance concern. Across the three questions, responsible use depends on the combined evidence profile rather than performance on any single measure.

5. Discussion

5.1. Implications for Evaluation, Implementation, and Longitudinal Evidence

The central evaluative implication is that early interactional promise does not constitute evidence sufficient for deployment decisions. Positive attitudes, perceived usefulness, comfort, and engagement indicate feasibility but do not establish safety, trustworthiness, clinical value, or sustainability [4,5,6,7,35,89,90,91,92]. For LLM-enabled SARs, socially responsive dialogue may inflate perceived competence, so acceptance results require complementary assessment of factual reliability, calibrated reliance, physical safety, continuity, and care outcomes.
Responsible implementation depends on more than technical performance. Role clarity, staff workload, training, maintenance, escalation, privacy governance, institutional readiness, equitable access, cost, and sustainable service models shape whether a system can be incorporated into care [3,19,28,29,82,84,87,88,93]. Because generative dialogue can blur the boundaries between information, emotional support, clinical advice, and autonomous action, evaluation must also capture supervision, updates, workflow effects, and burdens transferred to professionals or caregivers.
Longitudinal assessment is therefore indispensable. Single-session or tightly controlled studies cannot reveal whether inaccurate stored context, hallucination, misleading capability claims, privacy exposure, overreliance, or emotional attachment accumulate through repeated use.
Evidence from real-world use indicates that outcomes vary with routine integration, individual trajectories, caregiver burden, staff needs, frailty, sensory or cognitive impairment, novelty decline, personal goals, and service sustainability [75,76,77,83,84,91,94,95]. Longitudinal designs should combine repeated measures, interaction logs, qualitative follow-up, professional and caregiver input, adverse-event monitoring, implementation outcomes, and durable care-value measures.

5.2. HEART and Adjacent Evaluation Approaches

No single adjacent approach addresses the full combination of generative communication, physical agency, memory, and healthcare accountability. HEART therefore complements rather than replaces HRI evaluation, technology-acceptance models, trustworthy-AI governance, implementation science, medical-device assessment, wearable-sensing approaches, user-centred SAR design, or clinical LLM-safety frameworks. Table 7 compares their established contributions with the residual problems that arise when these concerns converge.
Across these approaches, HEART contributes a common boundary logic for tracing how failures in generated communication may affect reliance, physical action, longitudinal personalisation, and care delivery. Its healthcare grounding reflects vulnerable users, sensitive data, professional accountability, and the possibility of clinical harm. The novelty lies not in each criterion individually, but in assigning criteria to non-redundant primary objects, connecting them through causal and governance pathways, and constraining deployment through E and A gates.
HEART is presented as a healthcare-specific framework. Similar concerns may occur in other high-stakes applications, but transfer is not assumed. Any adaptation would require domain-specific definitions of value, stakeholders, accountability, and gate criteria, followed by separate empirical validation. Portability remains a research question rather than a claimed contribution.

5.3. Research Priorities for LLM-Enabled SARs

Five research priorities follow from the review. First, studies should specify the robot platform, model configuration and deployment mode, dialogue architecture, stored context, autonomy, intended users, care function, and prohibited claims [9]. Second, LLM-enabled systems require comparison with scripted robots, disembodied chatbots, telehealth tools, human-administered interventions, and usual care [8,10]. Third, physical-system testing should assess grounding, latency, language-to-action execution, recovery, escalation, and regression after updates [39].
Fourth, personalisation and memory-supported functions should be examined longitudinally through accuracy audits, user control over data, dependency indicators, repeated measures, and observation after novelty declines [57,73,74]. Fifth, safety, privacy, governance, and healthcare value should remain core outcomes throughout the lifecycle [11,12,42]. Recommended designs include adversarial prompts, vulnerable-user and clinical-boundary scenarios, severity-sensitive assessment of hallucinations and omissions, appropriate abstention and escalation, data minimisation, deletion testing, post-update validation, and monitoring of harmful or misleading outputs. Table 8 consolidates these implications and research priorities.

5.4. Staged Evaluation and Translation Roadmap

To support practical decisions, the roadmap translates HEART into a staged process that extends from use-case definition to post-deployment monitoring. Figure 4 arranges this process into seven sequential but revisitable stages, with arrows indicating the intended progression between them and advancement contingent on the evidence and risk criteria specified at each stage. Coloured H, E, A, R, and T markers indicate the primary HEART emphases at each stage, whereas grey markers denote dimensions that remain relevant but are not the primary focus. Stage 1 is marked H · E · A · R · T because intended use and claims define the communication role, governance boundaries, physical autonomy, memory and continuity assumptions, and healthcare-value criteria for the entire evaluation. It specifies target users, care setting, model and robot configuration, level of autonomy, memory functions, comparator, prohibited claims, and measurable success criteria. Stage 2 evaluates dialogue and model safety before user exposure through scenario-based testing of unsupported claims, clinically consequential omissions, severity of errors, appropriate abstention and escalation, adversarial prompts, privacy risks, and consistency with the robot’s stated role. At Stage 3, the physical system is verified through tests of language-to-action grounding, sensing and movement coordination, latency, interruption handling, recovery, action constraints, logging, and clinician or caregiver override. Unresolved critical E or A failures at these stages prevent progression.
Stage 4 introduces controlled user evaluation only after the pre-user safety gates are met. It examines comprehension of system limits, trust calibration, communication quality, task success, usability, emotional appropriateness, adverse responses, and the adequacy of repair and escalation, using representative users and relevant care professionals. Stage 5 extends evaluation across repeated encounters in a longitudinal care pilot. The focus shifts to memory accuracy and controllability, personalisation, changing needs, novelty decline, dependency, privacy-sensitive inference, caregiver and staff burden, and the stability of benefits and risks over time. Predefined progression criteria should specify which findings require redesign, restricted use, additional evidence, or termination.
Stage 6 evaluates implementation and healthcare value under pragmatic conditions and against a meaningful comparator. Outcomes should include patient or psychosocial benefit, workflow fit, staff and caregiver effects, cost, maintenance, equity, sustainability, residual risk, and organisational accountability. Although T is the primary translational decision focus, evidence from H, E, A, and R is considered jointly, as shown by H · E · A · R · T in the figure. Stage 7 introduces post-deployment monitoring of incidents, complaints, harmful or misleading outputs, model or data drift, memory errors, changes in user reliance, software and model updates, and emerging workflow effects. Each stage ends with an explicit decision to proceed, redesign, constrain, collect additional evidence, or stop. Major updates or adverse findings return the system to the relevant earlier stage rather than being treated as automatic progression along a simple maturity ladder.

5.5. Empirical Validation of HEART

Empirical testing should proceed in stages rather than begin with a composite score. First, a multidisciplinary content-validity and feasibility study should assess whether the dimension definitions, boundary rules, indicators, status labels, and deployment gates are relevant, comprehensive, understandable, and workable. Participants should include patients, caregivers, clinicians, HRI researchers, LLM-safety specialists, privacy and ethics experts, healthcare organisations, and regulators. A modified Delphi process combined with cognitive interviews and pilot use of Supplementary Table S4 would support item refinement and identify missing or redundant criteria.
Second, independent raters should apply HEART to standardised evidence packages for a heterogeneous sample of LLM-SAR systems and scenarios. Inter-rater agreement should be evaluated separately for dimension assignment, evidence-status labels, risk identification, and deployment decisions. Construct validity can then be examined through preregistered hypotheses. For example, systems with documented memory leakage should receive less favourable E profiles, systems with language-to-action failures less favourable A profiles, and longitudinal real-world studies more complete R and T evidence than single-session laboratory studies.
Third, prospective multi-site studies should test utility and decision validity by comparing HEART-guided evaluation with usual study planning or governance review. Relevant outcomes include additional risks detected, changes to outcome selection, agreement among stakeholders, time and burden required, safety incidents identified before deployment, and correspondence between pre-deployment HEART profiles and subsequent real-world events. Reassessment after model, memory, or robot-platform updates would test responsiveness to known changes. Numerical weighting or an overall HEART score should be considered only after the reliability, validity, and practical consequences of the individual dimensions and decision rules have been established.

6. Limitations

The scope and design impose several limits on interpretation. This PRISMA-informed structured review is neither a systematic review nor a meta-analysis, and it does not estimate pooled effects or establish clinical effectiveness. Its conclusions concern the organisation and interpretation of a heterogeneous, rapidly developing source base.
The audited database branch comprises 128 records, 18 duplicate removals, and 110 unique items. Targeted Google Scholar searching added methodological, arXiv, governance, regulatory, and selected non-LLM comparative sources. After consolidation, a separate coding log contained 110 candidates, of which 18 were excluded and 92 retained. Seven methodological references were then separated from substantive coding, leaving 85 sources. Because route-level additions, overlap, and attrition for the supplementary pathway were not preserved, the consolidated set cannot be reconstructed as a database-only chain. Figure 1 separates the branches and marks this transparency boundary; the design should therefore be interpreted as PRISMA-informed rather than as a systematic review with a fully reproducible search chain.
Study design, population, setting, robot type, technological maturity, and outcome measures varied substantially. Supplementary Table S2 reports evidence status and claim function, but empirical studies, technical demonstrations, protocols, reviews, and normative documents remain methodologically non-equivalent.
Direct evidence on LLM-enabled healthcare SARs remains limited. Fifteen of the 35 primary analytical sources explicitly involved an LLM or generative AI component, and few assessed longitudinal or translational value. The purposively selected non-LLM subset established requirements for care, workflow, continuity, and implementation; it was not treated as evidence of LLM safety or effectiveness, or as an exhaustive account of earlier SAR research.
Corpus assignment, HEART tagging, retrospective profiling, and thematic synthesis involved interpretive judgement. One author performed all screening and coding, so inter-rater reliability could not be assessed. Boundary rules and source-level crosswalks strengthen traceability but cannot eliminate subjectivity. HEART has not yet been prospectively tested for reliability, validity, or predictive utility.
Publication, language, database, and grey-literature constraints also apply, and machine-exported interface histories were not retained. Supplementary Tables S5 and S6 document the preserved Boolean expression, selected interfaces and fields, displayed yields, exploratory variants, Google Scholar queries, and 18 source-level exclusions, but they cannot reconstruct unavailable route-level transitions or imply exhaustive screening. The three post-lock contextual sources [40,41,42] remained outside corpus counts. HEART does not replace formal risk management, clinical validation, data-protection assessment, medical-device regulation, or institutional governance. Its applicability outside healthcare was not examined.

7. Conclusions

Research on LLM-enabled socially assistive robots is advancing faster than the evidence for durable healthcare benefit. Current studies provide stronger support for feasibility, interaction quality, and early acceptability than for longitudinal safety, governance of retained user context, comparative effectiveness, workflow impact, cost, equity, or post-deployment performance. This imbalance matters because generative output is delivered through socially credible systems that may retain information, shape behaviour, and initiate or influence physical action in settings governed by clinical and organisational accountability.
HEART addresses this evaluative challenge by linking five complementary domains: communication, trustworthy deployment, adaptive embodiment, relationship continuity, and healthcare value. Its boundary rules separate primary objects of assessment; operational indicators and qualitative labels expose missing evidence; and non-additive E and A gates prevent favourable results elsewhere from compensating for critical governance or safety failures.
Conceptually, HEART treats LLM-enabled SARs as integrated socio-technical interventions in care rather than conversational products. Practically, it offers a common structure for study design, system comparison, stakeholder review, and implementation decisions. Retrospective profiling shows that communication or feasibility gains can coexist with unresolved longitudinal, translational, or safety questions. Prospective validation should now test content validity, inter-rater reliability, construct validity, decision-making utility, multi-site applicability, and responsiveness to model or platform updates before any overall score is considered.

Supplementary Materials

The following supporting information can be downloaded at: https://www.mdpi.com/article/10.3390/app16167904/s1, Table S1: Source-level evidence-use crosswalk for the final substantive evidence base (n = 85); Table S2: Evidence-status and claim-function descriptors for the final substantive evidence base (n = 85); Table S3: Prospective application of HEART to a geriatric LLM-enabled SAR evaluation; Table S4: Reusable HEART application template and decision record; Table S5: Controlled-database implementation, exploratory variants, and targeted Google Scholar queries; Table S6: Source-level exclusion log for excluded candidates (n = 18) from the consolidated candidate-source set. References [40,41,42] were identified after the literature lock and are cited as contextual sources outside the corpus counts.

Funding

This research received no external funding.

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

The original contributions presented in this study are included in the article and Supplementary Materials. Further inquiries can be directed to the corresponding author.

Conflicts of Interest

The author declares no conflicts of interest.

Abbreviations

The following abbreviations are used:
AIArtificial Intelligence
EDPBEuropean Data Protection Board
EUEuropean Union
FDAU.S. Food and Drug Administration
GPTGenerative Pre-trained Transformer
HEARTHuman-Centred Communication, Ethical and Trustworthy Deployment, Adaptive and Embodied Intelligence, Relationship Continuity, and Translational Healthcare Value
HRIHuman–Robot Interaction
IMDRFInternational Medical Device Regulators Forum
LLMLarge Language Model
PRISMAPreferred Reporting Items for Systematic Reviews and Meta-Analyses
PRISMA-ScRPRISMA Extension for Scoping Reviews
RQResearch Question
SARSocially Assistive Robot

Appendix A

Table A1. Local gap map for healthcare robotics, HRI, intelligent robotics, and the embodied AI literature.
Table A1. Local gap map for healthcare robotics, HRI, intelligent robotics, and the embodied AI literature.
Literature StreamWhat It EstablishesRemaining Limitation for This Review
Assistive robots and healthcare HRIHealthcare robots are framed as communicative, adaptive, care-oriented systems that extend beyond task performance.Does not yet address how LLM-enabled dialogue, generative responses, and memory-based interaction change evaluation.
Healthcare social robots and hospital HRIRobot types, healthcare settings, service roles, verbal and non-verbal behaviours, personalisation, and interaction outcomes are already mapped.Evaluation remains largely organised around robot type, service function, technical requirements, and short-term engagement behaviours.
Intelligent roboticsAI, machine learning, sensing, HRI, ethics, privacy, and responsible implementation are central to emerging robotics across sectors.Cross-sectoral accounts do not resolve healthcare-specific concerns around vulnerability, care continuity, social credibility, and clinical value.
Embodied AI in healthcareRobotic perception, planning, decision-making, and action are central to physically situated AI systems in clinical environments.Emphasis remains stronger on perception, action, logistics, rehabilitation, and surgical domains than on socially assistive dialogue and relational interaction.
AI-based clinical HRITrust, therapeutic alliance, bonding, emotional security, personalisation, and responsiveness are becoming central to healthcare HRI.Longitudinal evidence, ecological validity, conceptual clarity around trust/engagement, and translation into sustained clinical practice remain limited.
Table A2. Local gap map for older adults, dementia, frailty, and long-term care contexts.
Table A2. Local gap map for older adults, dementia, frailty, and long-term care contexts.
Literature StreamWhat It EstablishesRemaining Limitation for This Review
Older adult care ecologiesRobots for older adults involve patients, nurses, caregivers, institutions, home environments, and technical support systems.Evaluation must move beyond one-user/one-robot interaction toward care routines, workflows, domestic variability, and stakeholder coordination.
Community and nursing careCare robots may support ageing in place, cognitive support, daily activities, monitoring, companionship, autonomy preservation, and point-of-care assistance.Practical value depends on fit with home environments, professional routines, patient comfort, caregiver burden, and workflow integration.
Dementia and frailty careSocially assistive and care robots are often feasible and acceptable, with some potential benefits for physical frailty.Effects on cognition, neuropsychiatric symptoms, quality of life, psychological frailty, and social frailty remain limited or inconclusive.
Implementation in care settingsBarriers and facilitators include technological complexity, patient needs, organisational context, care setting, maintenance, staffing, and implementation conditions.Uptake depends on system-level fit, not only robot characteristics or user acceptance.
Methodological evaluationFeasibility, usability, efficacy, and effectiveness have been studied through varied designs, intervention durations, and outcome measures.Short-term, heterogeneous, and small-scale evidence weakens claims about sustained usefulness, effectiveness, and relationship continuity.
Table A3. Local gap map for paediatric care, clinical workflows, and nursing practice.
Table A3. Local gap map for paediatric care, clinical workflows, and nursing practice.
Literature StreamWhat It EstablishesRemaining Limitation for This Review
Robot-assisted autism therapyRobots may support social learning, communication, imitation, interaction, and therapeutic engagement for children with autism.Evidence remains distributed across heterogeneous interventions, target behaviours, robot types, and methodological designs.
Paediatric pain and emotional outcomesRobotic support may reduce some forms of pain, anxiety, distress, or negative affect during paediatric procedures.Effects vary by outcome, setting, procedure, intervention design, and measurement approach.
Paediatric emotion regulationActive child participation, guided breathing, motivational support, and physical comfort can shape emotional outcomes.Robot presence alone is insufficient; evaluation must address interaction quality, participation, and individualised support.
Healthcare professional rolesHCPs see potential for workload reduction, efficiency, and patient well-being, while raising concerns about reliability, ethics, privacy, training, and role clarity.Clinical use depends on organisational readiness and workflow fit, not only perceived usefulness.
Nursing and cobotic practiceCobots may assist with logistics, monitoring, medication delivery, and social interaction in clinical settings.Evidence remains limited on nurse-centred design, real-world effectiveness, and impact on nursing workload and patient care.
Obstetrics and neonatal careNurses prioritise safety, sterilisation, precision, data transfer, alarms, and autonomous navigation in assistant robots.Sensitive care environments require stronger attention to hygiene, empathy, accountability, and responsibility boundaries.
Table A4. Local gap map for trust, acceptance, ethics, autonomy, and evaluation.
Table A4. Local gap map for trust, acceptance, ethics, autonomy, and evaluation.
Literature StreamWhat It EstablishesRemaining Limitation for This Review
Deployment and acceptanceSARs are deployed across diverse healthcare-related settings, and acceptance is shaped by usefulness, ease of use, attitudes, individual characteristics, social factors, and context.Acceptance and deployment do not by themselves establish sustained clinical, psychosocial, relational, or organisational value.
Trust and trust repairTrust is dynamic, context-dependent, measurable, and potentially repairable after robot errors or failures.Trust measurement remains fragmented, often self-reported, laboratory-based, and weakly standardised for vulnerable healthcare populations.
Ethics of socially assistive devicesSocial presence, behavioural alignment, efficiency claims, care relationships, well-being, and autonomy raise ethical questions.Positive reactions to robots do not automatically justify deployment or support claims of efficiency, autonomy, and well-being.
Autonomy and agencyHuman autonomy and sense of agency are central to HRI ethics and human well-being.Evaluations often focus on interface and task levels, with limited evidence on longer-term effects on behaviour, life, society, and care relationships.
Long-term care ethicsKey concerns include consent, privacy, substitution of human care, inequitable access, infantilisation, dignity, dependency, deception, and data protection.Ethical value depends on context, vulnerability, implementation safeguards, and the meaning of the care task.
Psychological measurementInstruments assess attitudes, trust, emotions, perceptions, social presence, anxiety, and related HRI constructs.Psychometric evidence is uneven, with limited validation for healthcare, children, older adults, and other specific care populations.
LLM-integrated robotsGenerative conversational systems may enable more natural and flexible interaction.LLMs introduce risks of deceptive behaviour, hallucination, unsafe reliance, opacity, and responsibility gaps in healthcare contexts.
Table A5. Local gap map for LLMs, generative AI, and robotic autonomy.
Table A5. Local gap map for LLMs, generative AI, and robotic autonomy.
Literature StreamWhat It EstablishesRemaining Limitation for This Review
LLM-driven HRILLMs can support natural language understanding, multimodal input, contextual sensing, generative interaction, high-level reasoning, plan generation, and task execution.Conversational fluency and task execution do not by themselves establish safety, contextual appropriateness, or healthcare value.
LLMs in socially assistive robotsLLMs and vision-language models may enhance social interaction, multimodal perception, human-oriented modelling, and flexible assistance.Increased apparent social competence may intensify risks related to trust, bias, ethics, and misleading user expectations.
Robotic autonomyLLMs can connect language commands with navigation, manipulation, semantic planning, voice interaction, and adaptive execution.Moving from language to action increases risks involving inappropriate behaviour, responsibility ambiguity, latency, privacy, and reliability.
LLM-driven HRI evaluationCurrent studies are moving toward contextual sensing, generative interaction, and adaptive alignment.The evidence remains exploratory and heterogeneous, limiting comparison, replication, and clinical interpretation.
Generative AI in health care social robotsThe intersection of generative AI, social robots, and healthcare is emerging as a distinct research area.Evidence remains insufficiently consolidated regarding implementation, outcomes, acceptance, and risk.
LLM-integrated care robotsLLMs may enable more natural, open-ended, and responsive interaction in healthcare robots.Deceptive behaviour, hallucinated capabilities, unsafe reliance, and responsibility gaps are especially serious in vulnerable care contexts.
Table A6. Evidence base and evaluative implications across HEART dimensions.
Table A6. Evidence base and evaluative implications across HEART dimensions.
HEART DimensionMain Evaluative ConcernRepresentative Sources Discussed in HEART SectionsKey Evaluative Implications
H—Human-Centred CommunicationCommunication that is understandable, emotionally appropriate, context-sensitive, clinically bounded, and multimodally coherentIrfan et al. [50,51]; Kim et al. [52]; Blavette et al. [53,54]; Fu et al. [55]; Pinto-Bernal et al. [57]; Lima et al. [58]; van ’t Klooster et al. [59]Assess clarity, dialogue repair, turn-taking, factual boundaries, emotional appropriateness, multimodal synchronisation, and user understanding of limitations.
E—Ethical and Trustworthy DeploymentSafeguards for trust, autonomy, privacy, consent, dignity, accountability, and safety during deploymentGul et al. [7]; Ranisch and Haltaufderheide [11]; Loizaga et al. [32]; Haltaufderheide et al. [33,34]; Hung et al. [36]; Deusdad [60]; EDPB [62]; EU AI Act [64]; FDA [66]Prioritise calibrated trust, transparency, hallucination safeguards, privacy-by-design, consent/assent, human oversight, accountability, and regulatory alignment.
A—Adaptive and Embodied IntelligenceSafe coordination of language, perception, planning, memory, movement, and actionLiu et al. [39]; Galatolo and Winkle [56]; Pinto-Bernal et al. [57]; Bärmann et al. [69]; Yoshida et al. [70]; Kang et al. [71]; Sorrentino et al. [72]Test grounding, multimodal coordination, safe action, latency, error recovery, controlled adaptation, and clinical supervision.
R—Relationship ContinuityContinuity of repeated interaction, memory, personalisation, and adaptation over timeIrfan et al. [50,51]; Pinto-Bernal et al. [57]; Lima et al. [58]; Kang et al. [71]; Mauliana et al. [73]; Zafrani et al. [75]; Zhao et al. [76]; Hofstede et al. [77]Track memory accuracy, user control over stored data, sustained engagement, personalisation, novelty decline, routine integration, and long-term privacy risk.
T—Translational Healthcare ValueClinical, psychosocial, organisational, workflow, and implementation value beyond novelty or acceptabilityZhu et al. [3]; Lee, Lee, Choi and Kim [18]; Watanabe et al. [19]; Yu et al. [20]; Che et al. [21]; Hsu et al. [24]; Pan et al. [25]; Or et al. [26]; Lee, Hsu and Lien [28]; Babalola et al. [29]; Blavette et al. [53,54]; Rigaud et al. [82]Demonstrate clinical relevance, psychosocial benefit, care workflow fit, manageable staff burden, sustainability, scalability, and cost-sensitive implementation.

References

  1. Aymerich-Franch, L.; Ferrer, I. Socially Assistive Robots’ Deployment in Healthcare Settings: A Global Perspective. Int. J. Humanoid Robot. 2023, 20, 2350002. [Google Scholar] [CrossRef] [Scilit]
  2. Ragno, L.; Borboni, A.; Vannetti, F.; Amici, C.; Cusano, N. Application of Social Robots in Healthcare: Review on Characteristics, Requirements, Technical Solutions. Sensors 2023, 23, 6820. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  3. Zhu, T.; Ahn, H.S.; Broadbent, E.; MacDonald, B.A.; Gasteiger, N. A Deep Dive into Human-Robot Interaction in Hospitals: Scoping Review on the Services Provided, Engagement Behaviours and Interaction Outcomes. Int. J. Soc. Robot. 2025, 17, 1541–1561. [Google Scholar] [CrossRef] [Scilit]
  4. He, Y.; He, Q.; Liu, Q. Technology Acceptance in Socially Assistive Robots: Scoping Review of Models, Measurement, and Influencing Factors. J. Healthc. Eng. 2022, 2022, 6334732. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  5. Mahmoudi Asl, A.; Molinari Ulate, M.; Franco Martin, M.; Van Der Roest, H. Methodologies Used to Study the Feasibility, Usability, Efficacy, and Effectiveness of Social Robots for Elderly Adults: Scoping Review. J. Med. Internet Res. 2022, 24, e37434. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  6. Vagnetti, R.; Camp, N.; Story, M.; Ait-Belaid, K.; Mitra, S.; Zecca, M.; Di Nuovo, A.; Magistro, D. Instruments for Measuring Psychological Dimensions in Human-Robot Interaction: Systematic Review of Psychometric Properties. J. Med. Internet Res. 2024, 26, e55597. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  7. Gul, A.; Turner, L.; Fuentes, C. Conventions and Research Challenges in Considering Trust with Socially Assistive Robots for Older Adults. Front. Robot. AI 2025, 12, 1631206. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  8. Zhang, C.; Chen, J.; Li, J.; Peng, Y.; Mao, Z. Large Language Models for Human–Robot Interaction: A Review. Biomim. Intell. Robot. 2023, 3, 100131. [Google Scholar] [CrossRef] [Scilit]
  9. Lempe, P.N.; Guinemer, C.; Fürstenau, D.; Dressler, C.; Balzer, F.; Schaaf, T. Health Care Social Robots in the Age of Generative AI: Protocol for a Scoping Review. JMIR Res. Protoc. 2025, 14, e63017. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  10. Wang, Y.; Xu, Y.; Nikolova, A.; Wang, Y.; Wang, J.; Wang, C.; Tong, X. How Do We Research Human-Robot Interaction in the Age of Large Language Models? A Systematic Review. In Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems, Barcelona, Spain, 13–17 April 2026; pp. 1–28. [Google Scholar] [CrossRef] [Scilit]
  11. Ranisch, R.; Haltaufderheide, J. Rapid Integration of LLMs in Healthcare Raises Ethical Concerns: An Investigation into Deceptive Patterns in Social Robots. Digit. Soc. 2025, 4, 7. [Google Scholar] [CrossRef] [Scilit]
  12. Zoughalian, K.; Aitsam, M.; Marchang, J.; Di Nuovo, A. Guardians of Privacy: Leveraging LLMs in Assistive Robotic Systems for Healthcare. In Proceedings of the 2025 IEEE Conference on Communications and Network Security (CNS), Avignon, France, 8–11 September 2025; pp. 1–6. [Google Scholar] [CrossRef] [Scilit]
  13. Baharum, A.; Ismail, R.; Daruis, D.D.I.; Ismail, I.; Mat Noor, N.A.; Deris, F.D. Human-Robot Interaction-Healthcare: Systematic Review. In Frontiers in Artificial Intelligence and Applications; Ye, Y., Zhou, H., Eds.; IOS Press: Amsterdam, The Netherlands, 2025; Volume 404, pp. 537–543. [Google Scholar] [CrossRef] [Scilit]
  14. Licardo, J.T.; Domjan, M.; Orehovački, T. Intelligent Robotics—A Systematic Review of Emerging Technologies and Trends. Electronics 2024, 13, 542. [Google Scholar] [CrossRef] [Scilit]
  15. Mir, B.A.; Nishwa, D.E.; Lee, S.W. Embodied Artificial Intelligence in Healthcare: A Systematic Review of Robotic Perception, Decision-Making, and Clinical Impact. Healthcare 2026, 14, 572. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  16. Shafaie, M.M.; Salehnia, A.; Moslem, N.; Shafaie, V.; Movahedi Rad, M. Human-Robot Interaction Based on Artificial Intelligence in Clinical Healthcare Centers: A Systematic Review and Meta-Analysis. Comput. Hum. Behav. Rep. 2026, 22, 101080. [Google Scholar] [CrossRef] [Scilit]
  17. Dino, M.J.S.; Davidson, P.M.; Dion, K.W.; Szanton, S.L.; Ong, I.L. Nursing and Human-Computer Interaction in Healthcare Robots for Older People: An Integrative Review. Int. J. Nurs. Stud. Adv. 2022, 4, 100072. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  18. Lee, J.; Lee, H.; Choi, M.; Kim, J.A. Care Robots for Community-Dwelling Older Adults: An Integrative Review. Healthc. Inform. Res. 2025, 31, 347–366. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  19. Watanabe, T.; Li, J.; Grace Torii, M.; Miura, Y.; Dai, M.; Nakagami, G.; Hirai, S. Current State and Future Perspectives on Robotic Technology for Improving Efficiency of Nursing Care. Adv. Robot. 2025, 39, 1482–1505. [Google Scholar] [CrossRef] [Scilit]
  20. Yu, C.; Sommerlad, A.; Sakure, L.; Livingston, G. Socially Assistive Robots for People with Dementia: Systematic Review and Meta-Analysis of Feasibility, Acceptability and the Effect on Cognition, Neuropsychiatric Symptoms and Quality of Life. Ageing Res. Rev. 2022, 78, 101633. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  21. Che, R.-P.; Ruan, Y.-X.; Kodate, N.; Shi, Y.; Liu, X.; Donnelly, S.; Suwa, S.; Yu, W.; Kong, D.; Cheung, M.-C. Effectiveness and Usability of Care Robots in Supporting Older Adults Living with Frailty: A Systematic Review. Digit. Health 2025, 11, 20552076251370058. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  22. Koh, W.Q.; Felding, S.A.; Budak, K.B.; Toomey, E.; Casey, D. Barriers and Facilitators to the Implementation of Social Robots for Older Adults and People with Dementia: A Scoping Review. BMC Geriatr. 2021, 21, 351. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  23. Alabdulkareem, A.; Alhakbani, N.; Al-Nafjan, A. A Systematic Review of Research on Robot-Assisted Therapy for Children with Autism. Sensors 2022, 22, 944. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  24. Hsu, F.Y.; Lee, Y.H.; Tsai, J.-L.; Lien, A.S.-Y. Socially Assistive Robots for Pain Management and Emotional Responses in Pediatric Hospital Care: Systematic Review and Meta-Analysis. J. Med. Internet Res. 2025, 27, e76427. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  25. Pan, X.-Y.; Bi, X.-Y.; Nong, Y.-N.; Ye, X.-C.; Yan, Y.; Shang, J.; Zhou, Y.-M.; Yao, Y.-Z. The Efficacy of Socially Assistive Robots in Improving Children’s Pain and Negative Affectivity during Needle-Based Invasive Treatment: A Systematic Review and Meta-Analysis. BMC Pediatr. 2024, 24, 643. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  26. Or, X.Y.; Ng, Y.X.; Goh, Y.S. Effectiveness of Social Robots in Improving Psychological Well-Being of Hospitalised Children: A Systematic Review and Meta-Analysis. J. Pediatr. Nurs. 2025, 82, 11–20. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  27. Neerincx, A.; Plat, J.; De Graaf, M.M.A. Socially Assistive Robots in Child Healthcare: Evaluating Internal and External Emotion Regulation Interventions. Front. Robot. AI 2025, 12, 1628795. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  28. Lee, Y.H.; Hsu, F.Y.; Lien, A.S.-Y. Health Care Professionals’ Perspectives of Socially Assistive Robots in Health Care Settings: Systematic Review. J. Med. Internet Res. 2025, 27, e79634. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  29. Babalola, G.T.; Gaston, J.-M.; Trombetta, J.; Tulk Jesso, S. A Systematic Review of Collaborative Robots for Nurses: Where Are We Now, and Where Is the Evidence? Front. Robot. AI 2024, 11, 1398140. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  30. İnam, Ö.; Okay, S. Evaluation of Nurses’ Perspectives on the Design and Use of Assistant Nurse Robots in Obstetrics and Neonatal Care: A Mixed-Method Study. BMC Nurs. 2025, 24, 359. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  31. Ayoub, F.; Kerr, A.; Villing, R. An Exploration of Trust in Human-Robot Interaction: From Measurement to Repair Strategies and Design Principles. In Human-Friendly Robotics 2024; Paolillo, A., Giusti, A., Abbate, G., Eds.; Springer Nature: Cham, Switzerland, 2025; Volume 35, pp. 58–72. [Google Scholar] [CrossRef] [Scilit]
  32. Loizaga, E.; Bastida, L.; Sillaurren, S.; Moya, A.; Toledo, N. Modelling and Measuring Trust in Human–Robot Collaboration. Appl. Sci. 2024, 14, 1919. [Google Scholar] [CrossRef] [Scilit]
  33. Haltaufderheide, J.; Lucht, A.; Strünck, C.; Vollmann, J. Socially Assistive Devices in Healthcare—A Systematic Review of Empirical Evidence from an Ethical Perspective. Sci. Eng. Ethics 2023, 29, 5. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  34. Haltaufderheide, J.; Lucht, A.; Strünck, C.; Vollmann, J. Increasing Efficiency and Well-Being? A Systematic Review of the Empirical Claims of the Double-Benefit Argument in Socially Assistive Devices. BMC Med. Ethics 2023, 24, 106. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  35. Glawe, F.; Schmeckel, T.; Brauner, P.; Ziefle, M. Human Autonomy and Sense of Agency in Human-Robot Interaction: A Systematic Literature Review. Int. J. Soc. Robot. 2026, 18, 64. [Google Scholar] [CrossRef] [Scilit]
  36. Hung, L.; Zhao, Y.; Alfares, H.; Shafiekhani, P. Ethical Considerations in the Use of Social Robots for Supporting Mental Health and Wellbeing in Older Adults in Long-Term Care. Front. Robot. AI 2025, 12, 1560214. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  37. Leineweber, M.; Keusgen, C.V.; Bubeck, M.; Ranisch, R.; Haltaufderheide, J.; Klingler, C. Ethical Aspects of the Use of Social Robots in Caring for Older People—A Systematic Qualitative Review. Med. Health Care Philos. 2026, 29, 209–224. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  38. Atuhurra, J. Leveraging Large Language Models in Human-Robot Interaction: A Critical Analysis of Potential and Pitfalls. arXiv 2024, arXiv:2405.00693. [Google Scholar] [CrossRef] [Scilit]
  39. Liu, Y.; Sun, Q.; Kapadia, D.R. Integrating Large Language Models into Robotic Autonomy: A Review of Motion, Voice, and Training Pipelines. AI 2025, 6, 158. [Google Scholar] [CrossRef] [Scilit]
  40. Del-Valle-Soto, C.; Briseño, R.A.; Valdivia, L.J.; Nolazco-Flores, J.A. Unveiling Wearables: Exploring the Global Landscape of Biometric Applications and Vital Signs and Behavioral Impact. BioData Min. 2024, 17, 15. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  41. Jung, H.W.; Park, J.Y.; Holoubek, T.; Kim, W.J.; Park, J. Socially Assistive Robots in Mental Healthcare: Principles and Conceptual Framework for User-Centered Design. Int. J. Soc. Robot. 2025, 17, 2827–2851. [Google Scholar] [CrossRef] [Scilit]
  42. Asgari, E.; Montaña-Brown, N.; Dubois, M.; Khalil, S.; Balloch, J.; Au Yeung, J.; Pimenta, D. A Framework to Assess Clinical Safety and Hallucination Rates of LLMs for Medical Text Summarisation. npj Digit. Med. 2025, 8, 274. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  43. Arksey, H.; O’Malley, L. Scoping Studies: Towards a Methodological Framework. Int. J. Soc. Res. Methodol. 2005, 8, 19–32. [Google Scholar] [CrossRef] [Scilit]
  44. Levac, D.; Colquhoun, H.; O’Brien, K.K. Scoping Studies: Advancing the Methodology. Implement. Sci. 2010, 5, 69. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  45. Tricco, A.C.; Lillie, E.; Zarin, W.; O’Brien, K.K.; Colquhoun, H.; Levac, D.; Moher, D.; Peters, M.D.J.; Horsley, T.; Weeks, L.; et al. PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and Explanation. Ann. Intern. Med. 2018, 169, 467–473. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  46. Page, M.J.; McKenzie, J.E.; Bossuyt, P.M.; Boutron, I.; Hoffmann, T.C.; Mulrow, C.D.; Shamseer, L.; Tetzlaff, J.M.; Akl, E.A.; Brennan, S.E.; et al. The PRISMA 2020 Statement: An Updated Guideline for Reporting Systematic Reviews. BMJ 2021, 372, n71. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  47. Peters, M.D.J.; Marnie, C.; Tricco, A.C.; Pollock, D.; Munn, Z.; Alexander, L.; McInerney, P.; Godfrey, C.M.; Khalil, H. Updated Methodological Guidance for the Conduct of Scoping Reviews. JBI Evid. Synth. 2020, 18, 2119–2126. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  48. Braun, V.; Clarke, V. Using Thematic Analysis in Psychology. Qual. Res. Psychol. 2006, 3, 77–101. [Google Scholar] [CrossRef] [Scilit]
  49. Thomas, J.; Harden, A. Methods for the Thematic Synthesis of Qualitative Research in Systematic Reviews. BMC Med. Res. Methodol. 2008, 8, 45. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  50. Irfan, B.; Kuoppamäki, S.; Skantze, G. Recommendations for Designing Conversational Companion Robots with Older Adults through Foundation Models. Front. Robot. AI 2024, 11, 1363713. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  51. Irfan, B.; Kuoppamäki, S.; Hosseini, A.; Skantze, G. Between Reality and Delusion: Challenges of Applying Large Language Models to Companion Robots for Open-Domain Dialogues with Older Adults. Auton. Robot. 2025, 49, 9. [Google Scholar] [CrossRef] [Scilit]
  52. Kim, C.Y.; Lee, C.P.; Mutlu, B. Understanding Large-Language Model (LLM)-Powered Human-Robot Interaction. In Proceedings of the 2024 ACM/IEEE International Conference on Human-Robot Interaction, Boulder, CO, USA, 11–15 March 2024; pp. 371–380. [Google Scholar] [CrossRef] [Scilit]
  53. Blavette, L.; Dacunha, S.; Alameda-Pineda, X.; Cattoni, J.; Rigaud, A.-S.; Pino, M. Integrating a Large Language Model into a Socially Assistive Robot in a Hospital Geriatric Unit: Two-Wave Comparative Study on Performance, Engagement, and User Perceptions. JMIR Hum. Factors 2025, 12, e81936. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  54. Blavette, L.; Dacunha, S.; Alameda-Pineda, X.; Hernández García, D.; Gannot, S.; Gras, F.; Gunson, N.; Lemaignan, S.; Polic, M.; Tandeitnik, P.; et al. Acceptability and Usability of a Socially Assistive Robot Integrated with a Large Language Model for Enhanced Human-Robot Interaction in a Geriatric Care Institution: Mixed Methods Evaluation. JMIR Hum. Factors 2025, 12, e76496. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  55. Fu, M.; Shi, Z.; Huang, M.; Liu, S.; Kian, M.; Song, Y.; Matarić, M.J. Personalized Socially Assistive Robots with End-to-End Speech-Language Models for Well-Being Support. In Social Robotics + AI; Staffa, M., Cabibihan, J.-J., Siciliano, B., Ge, S.S., Bodenhagen, L., Tapus, A., Rossi, S., Cavallo, F., Fiorini, L., Matarese, M., et al., Eds.; Springer Nature: Singapore, 2026; Volume 16132, pp. 192–206. [Google Scholar] [CrossRef] [Scilit]
  56. Galatolo, A.; Winkle, K. Simultaneous Text and Gesture Generation for Social Robots with Small Language Models. Front. Robot. AI 2025, 12, 1581024. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  57. Pinto-Bernal, M.; Biondina, M.; Belpaeme, T. Designing Social Robots with LLMs for Engaging Human Interaction. Appl. Sci. 2025, 15, 6377. [Google Scholar] [CrossRef] [Scilit]
  58. Lima, M.R.; O’Connell, A.; Zhou, F.; Nagahara, A.; Hulyalkar, A.; Deshpande, A.; Thomason, J.; Vaidyanathan, R.; Matarić, M. Promoting Cognitive Health in Elder Care with Large Language Model-Powered Socially Assistive Robots. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, Yokohama, Japan, 26 April–1 May 2025; pp. 1–22. [Google Scholar] [CrossRef] [Scilit]
  59. Van ’T Klooster, J.-W.J.R.; Capasso, M.; Van Gorssel, D.; Vrolijk, E.; Rettagliata, G.; Gerritsen, D.; Hegeman, M.; Tauro, E.; Caiani, E.G.; Vonkeman, H.E. A GPT-Reinforced Social Robot for Patient Communication: A Pilot Study. Front. Digit. Health 2026, 7, 1653168. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  60. Deusdad, B. Ethical Implications in Using Robots among Older Adults Living with Dementia. Front. Psychiatry 2024, 15, 1436273. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  61. Ware, C.; Rigaud, A.-S.; Blavette, L.; Damnée, S.; Dacunha, S.; Lenoir, H.; Piccoli, M.; Cristancho-Lacroix, V.; Pino, M. Development of an Ethical Framework for the Use of Social Robots in the Care of Individuals with Major Neurocognitive Disorders: A Qualitative Study. BMC Geriatr. 2025, 25, 260. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  62. Barberá, I. AI Privacy Risks & Mitigations: Large Language Models (LLMs); European Data Protection Board: Brussels, Belgium, 2025; Available online: https://www.edpb.europa.eu/our-work-tools/our-documents/support-pool-experts-projects/ai-privacy-risks-mitigations-large_en (accessed on 10 June 2026).
  63. Kim, J.-W.; Choi, Y.-L.; Jeong, S.-H.; Han, J. A Care Robot with Ethical Sensing System for Older Adults at Home. Sensors 2022, 22, 7515. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  64. European Commission. Artificial Intelligence Act: Regulation (EU) 2024/1689 Laying Down Harmonised Rules on Artificial Intelligence; European Commission: Brussels, Belgium, 2024; Available online: https://digital-strategy.ec.europa.eu/en/policies/regulatory-framework-ai (accessed on 10 June 2026).
  65. Aboy, M.; Minssen, T.; Vayena, E. Navigating the EU AI Act: Implications for Regulated Digital Medical Products. npj Digit. Med. 2024, 7, 237. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  66. U.S. Food and Drug Administration. Marketing Submission Recommendations for a Predetermined Change Control Plan for Artificial Intelligence-Enabled Device Software Functions: Guidance for Industry and Food and Drug Administration Staff; U.S. Food and Drug Administration: Silver Spring, MD, USA, 2025. Available online: https://www.fda.gov/regulatory-information/search-fda-guidance-documents/marketing-submission-recommendations-predetermined-change-control-plan-artificial-intelligence (accessed on 10 June 2026).
  67. International Medical Device Regulators Forum. Good Machine Learning Practice for Medical Device Development: Guiding Principles; IMDRF/AIML WG/N88 FINAL:2025; International Medical Device Regulators Forum: Singapore, 2025. Available online: https://www.fda.gov/medical-devices/software-medical-device-samd/good-machine-learning-practice-medical-device-development-guiding-principles (accessed on 10 June 2026).
  68. U.S. Food and Drug Administration; Health Canada; United Kingdom’s Medicines and Healthcare Products Regulatory Agency. Transparency for Machine Learning-Enabled Medical Devices: Guiding Principles. 2024. Available online: https://www.fda.gov/medical-devices/software-medical-device-samd/transparency-machine-learning-enabled-medical-devices-guiding-principles (accessed on 10 June 2026).
  69. Bärmann, L.; Kartmann, R.; Peller-Konrad, F.; Niehues, J.; Waibel, A.; Asfour, T. Incremental Learning of Humanoid Robot Behavior from Natural Interaction and Large Language Models. Front. Robot. AI 2024, 11, 1455375. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  70. Yoshida, T.; Masumori, A.; Ikegami, T. From Text to Motion: Grounding GPT-4 in a Humanoid Robot “Alter3”. Front. Robot. AI 2025, 12, 1581110. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  71. Kang, H.; Ben Moussa, M.; Thalmann, N.M. Nadine: A Large Language Model-driven Intelligent Social Robot with Affective Capabilities and Human-like Memory. Comput. Animat. Virtual Worlds 2024, 35, e2290. [Google Scholar] [CrossRef] [Scilit]
  72. Sorrentino, A.; Fiorini, L.; Mancioppi, G.; Cavallo, F.; Umbrico, A.; Cesta, A.; Orlandini, A. Personalizing Care Through Robotic Assistance and Clinical Supervision. Front. Robot. AI 2022, 9, 883814. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  73. Mauliana, M.; Ashok, A.; Czernochowski, D.; Berns, K. Exploring LLM-Powered Multi-Session Human-Robot Interactions with University Students. Front. Robot. AI 2025, 12, 1585589. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  74. Bossema, M.; Ben Allouch, S.; Plaat, A.; Saunders, R. LLM-Enhanced Interactions in Human-Robot Collaborative Drawing with Older Adults. In Proceedings of the 2025 34th IEEE International Conference on Robot and Human Interactive Communication (RO-MAN), Eindhoven, The Netherlands, 25–29 August 2025; pp. 700–707. [Google Scholar] [CrossRef] [Scilit]
  75. Zafrani, O.; Nimrod, G.; Krakovski, M.; Kumar, S.; Bar-Haim, S.; Edan, Y. Assimilation of Socially Assistive Robots by Older Adults: An Interplay of Uses, Constraints and Outcomes. Front. Robot. AI 2024, 11, 1337380. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  76. Zhao, I.Y.; Leung, A.Y.M.; Huang, Y.; Liu, Y. A Social Robot in Home Care: Acceptability and Utility Among Community-Dwelling Older Adults. Innov. Aging 2025, 9, igaf019. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  77. Hofstede, B.M.; Askari, S.I.; Van Hoesel, T.R.C.; Cuijpers, R.H.; De Witte, L.P.; IJsselsteijn, W.A.; Nap, H.H. Huggable Integrated Socially Assistive Robots: Exploring the Potential and Challenges for Sustainable Use in Long-Term Care Contexts. Front. Robot. AI 2025, 12, 1646353. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  78. Ostrowski, A.K.; Zhang, J.; Breazeal, C.; Park, H.W. Promising Directions for Human-Robot Interactions Defined by Older Adults. Front. Robot. AI 2024, 11, 1289414. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  79. Olatunji, S.A.; Falcon, V.; Ramesh, A.; Rogers, W.A. Considerations for Designing Socially Assistive Robots for Older Adults. Front. Robot. AI 2025, 12, 1622206. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  80. Rincon Arango, J.A.; Marco-Detchart, C.; Julian Inglada, V.J. Personalized Cognitive Support via Social Robots. Sensors 2025, 25, 888. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  81. Zhao, G.; Song, E. Privacy-Preserving Large Language Models: Mechanisms, Applications, and Future Directions. arXiv 2024, arXiv:2412.06113. [Google Scholar] [CrossRef] [Scilit]
  82. Rigaud, A.-S.; Dacunha, S.; Harzo, C.; Lenoir, H.; Sfeir, I.; Piccoli, M.; Pino, M. Implementation of Socially Assistive Robots in Geriatric Care Institutions: Healthcare Professionals’ Perspectives and Identification of Facilitating Factors and Barriers. J. Rehabil. Assist. Technol. Eng. 2024, 11, 20556683241284765. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  83. Amabili, G.; Maranesi, E.; Margaritini, A.; Bonfigli, A.R.; Felici, E.; Barbarossa, F.; Benadduci, M.; Gosetto, L.; Guebey, J.; Grimstad, T.; et al. Managing Cognitive Decline Through a Social Robot–Based Intervention: Protocol for the engAGE Proof of Concept and Randomized Controlled Trial. JMIR Res. Protoc. 2025, 14, e67601. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  84. Van Dam, K.N.; Gielissen, M.F.M.; Siebelink, N.M.; Van Mastrigt, G.A.P.G.; Den Hollander, W.; Boon, B. Effectiveness and Cost-Effectiveness of Using a Social Robot in Residential Care for Individuals with Challenges in Daily Structure and Planning: Protocol for a Multiple-Baseline Single Case Trial and Health Economic Evaluation. JMIR Res. Protoc. 2025, 14, e67841. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  85. Ruan, Y.; Cheung, M. Effectiveness of Group Interventions with Socially-Assistive Robots for Older Adults: A Systematic Review. J. Nurs. Scholarsh. 2025, 57, 919–940. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  86. Shen, J.; Yu, J.; Zhang, H.; Lindsey, M.A.; An, R. Artificial Intelligence-Powered Social Robots for Promoting Physical Activity in Older Adults: A Systematic Review. J. Sport Health Sci. 2025, 14, 101045. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  87. Leoste, J.; Lubi, K.; Marmor, K.; Kangur, K. Evaluating Social Assistive Robots in Clinical Nursing Care: Mixed Method Pilot Study on Health Care Workers’ Perceptions and Adoption. JMIR Nurs. 2025, 8, e70305. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  88. Mlakar, I.; Smrke, U.; Šafran, V.; Roj, I.R.; Ilijevec, B.; Horvat, S.; Flis, V.; Plohl, N. A Randomized Pilot Study Evaluating Socially Assistive Robot Effects on Patient Engagement and Care Quality. npj Digit. Med. 2025, 8, 738. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  89. Aymerich-Franch, L.; Gómez, E. Public Perception of Socially Assistive Robots for Healthcare in the EU: A Large-Scale Survey. Comput. Hum. Behav. Rep. 2024, 15, 100465. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  90. Tobis, S.; Piasek, J.; Cylkowska-Nowak, M.; Suwalska, A. Robots in Eldercare: How Does a Real-World Interaction with the Machine Influence the Perceptions of Older People? Sensors 2022, 22, 1717. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  91. Cavallaro, A.; Perillo, F.; Romano, M.; Sebillo, M.; Vitiello, G. Social Robot in Service of the Cognitive Therapy of Elderly People: Exploring Robot Acceptance in a Real-World Scenario. Image Vis. Comput. 2024, 147, 105072. [Google Scholar] [CrossRef] [Scilit]
  92. Elsheikh, A.; Al-Thani, D.; Othman, A. Exploring Fear in Human-Robot Interaction: A Scoping Review of Older Adults’ Experiences with Social Robots. Front. Robot. AI 2025, 12, 1626471. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  93. Müller, P.; Jahn, P. Cocreative Development of Robotic Interaction Systems for Health Care: Scoping Review. JMIR Hum. Factors 2024, 11, e58046. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  94. Dinesen, B.; Hansen, H.K.; Grønborg, G.B.; Dyrvig, A.-K.; Leisted, S.D.; Stenstrup, H.; Skov Schacksen, C.; Oestergaard, C. Use of a Social Robot (LOVOT) for Persons with Dementia: Exploratory Study. JMIR Rehabil. Assist. Technol. 2022, 9, e36505. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  95. Preston, R.C.; Shippy, M.R.; Aldwin, C.M.; Fitter, N.T. How Can Robots Facilitate Physical, Cognitive, and Social Engagement in Skilled Nursing Facilities? Front. Aging 2024, 5, 1463460. [Google Scholar] [CrossRef] [Scilit] [PubMed]
Figure 1. PRISMA-informed source-selection flow and transparency boundary.
Figure 1. PRISMA-informed source-selection flow and transparency boundary.
Applsci 16 07904 g001
Figure 2. HEART dimensions and non-additive deployment gates.
Figure 2. HEART dimensions and non-additive deployment gates.
Applsci 16 07904 g002
Figure 3. Coverage of the five HEART dimensions by AI/LLM status in Corpus B.
Figure 3. Coverage of the five HEART dimensions by AI/LLM status in Corpus B.
Applsci 16 07904 g003
Figure 4. Staged HEART evaluation and translation pathway.
Figure 4. Staged HEART evaluation and translation pathway.
Applsci 16 07904 g004
Table 1. Mapping reviewed evidence streams to the HEART framework.
Table 1. Mapping reviewed evidence streams to the HEART framework.
Evidence StreamMain ContributionUnresolved Evaluative ProblemHEART Dimension
Healthcare robotics, HRI, and embodied AIRobots are treated as communicative, adaptive, embodied, and clinically situated systems.Existing work does not fully integrate language, embodiment, social presence, and care-specific evaluation.Human-Centred Communication; Adaptive and Embodied Intelligence
Older adult, dementia, frailty, and long-term care contextsSARs are relevant for companionship, support, daily living, social interaction, and care continuity.Feasibility and acceptability do not establish sustained benefit, contextual fit, or continuity of care.Relationship Continuity; Translational Healthcare Value
Paediatric care, clinical workflows, and nursing practiceRobots may support emotional coping, procedural care, workload, and clinical routines.Patient-facing support remains insufficiently connected to professional responsibility and workflow integration.Human-Centred Communication; Translational Healthcare Value
Trust, acceptance, ethics, autonomy, and evaluationTrust, acceptance, autonomy, social presence, and measurement are central but difficult to standardise.Positive reactions and trust may become unsafe if not calibrated, transparent, and ethically governed.Ethical and Trustworthy Deployment
LLMs, generative AI, and robotic autonomyLLMs connect open-ended dialogue, reasoning, planning, adaptive interaction, and embodied action.Evaluation must account for the combined effects of generative language, autonomy, embodiment, and care-context vulnerability.Ethical and Trustworthy Deployment; Adaptive and Embodied Intelligence
Integrative gap across the reviewed literatureExisting evidence is rich but fragmented across technical, clinical, relational, ethical, and implementation domains.LLM-enabled SARs require an evaluative structure linking communication, trust, embodiment, continuity, and healthcare value.HEART framework
Table 2. Search, supplementary-identification, retrieval, and contextual-update routes used in the PRISMA-informed structured review.
Table 2. Search, supplementary-identification, retrieval, and contextual-update routes used in the PRISMA-informed structured review.
Search or Update RouteFunction in the ReviewRole in Numerical Flow
Web of Science Core Collection—Advanced Search, all fieldsControlled database identification (n = 47).Exact database branch
Scopus—title, abstract, and keywordsControlled database identification (n = 81).Exact database branch
Web of Science Core Collection—Smart SearchExploratory calibration search (n = 136 displayed results).Not included in n = 128
Scopus—All FieldsExploratory calibration search (n = 5167 displayed results).Not included in n = 128
Google ScholarTargeted supplementary identification of methodological references, arXiv sources, governance or regulatory documents, and selected non-LLM SAR/healthcare-robot comparative baseline sources.Additions/overlap not retained
arXiv and governance/regulatory websitesRetrieval of documents identified through Google Scholar; not independently searched as separate systematic routes.Retrieval only
Targeted contextual update (n = 3)Framework comparison and metric operationalisation after the literature lock.Outside corpus counts
Note: The exact controlled-database branch is 47 + 81 = 128 records and 18 duplicate removals, leaving 110 unique database records. Google Scholar contributed methodological, arXiv, governance or regulatory, and selected non-LLM comparative baseline sources to the consolidated working set, but source-level additions, overlap, and attrition by route were not retained. The matching n = 110 totals refer to different source memberships after consolidation. The three post-lock contextual additions [40,41,42] are cited but excluded from the n = 85 corpus.
Table 3. Summary profile of Corpus B (n = 35).
Table 3. Summary profile of Corpus B (n = 35).
DimensionCategoryNumber of Sources
Publication year20224
20230
202411
202518
20262
Primary source orientationEmpirical, field, pilot, survey, or user evaluation study21
Technical, system design, or prototype demonstration9
Design, co-design, or user-informed design study3
Study protocol2
AI/LLM statusExplicit LLM, foundation-model, small-language-model, GPT-based, or generative AI robotic source15
Non-LLM socially assistive, healthcare, AI-enabled, or HRI baseline source20
Primary contextOlder adult, geriatric, dementia, home care, long-term care, or residential care20
Clinical, nursing, paediatric, or patient communication context6
General HRI, technical robotics, or design context with healthcare relevance9
Note: Within each descriptive variable, categories are mutually exclusive and sum to n = 35. Sources could inform multiple HEART dimensions during thematic synthesis; those overlapping tags are analysed separately in Section 4.7 and Supplementary Table S1.
Table 4. HEART dimension boundaries, LLM-specific failure modes, and deployment implications.
Table 4. HEART dimension boundaries, LLM-specific failure modes, and deployment implications.
DimensionPrimary Object and Unit of EvaluationBoundary RuleLLM-Specific Failure ModesDeployment Implication
H—Human-Centred CommunicationGenerated communicative act and user understanding.Use H when the outcome concerns clarity, appropriateness, repair, emotional sensitivity, or understood limitations.Unsupported health claims; sycophancy; misleading reassurance; fluent but incoherent dialogue.Poor H requires redesign and retesting; unsafe clinical communication may trigger an E gate.
E—Ethical and Trustworthy DeploymentJustified reliance, privacy, consent, accountability, and oversight.Use E for governance of trust, data, responsibility, escalation, contestability, and human control.Fluency-driven overtrust; privacy leakage through memory; prompt injection; absent escalation; misleading capability claims.Critical E failures block deployment regardless of performance elsewhere.
A—Adaptive and Embodied IntelligenceTechnical coupling of language, sensing, planning, movement, and action.Use A when the outcome concerns grounding, multimodal synchronisation, latency, robustness, recovery, or action safety.Language-to-action error; unsafe plan execution; grounding failure; nondeterministic embodied behaviour.Critical A failures block user-facing deployment until verified and controlled.
R—Relationship ContinuitySystem state, memory, personalisation, and user relationship across encounters.Use R for longitudinal accuracy, adaptation, user control, novelty decline, dependency, and changing needs.False or stale memory; uncontrolled personalisation; emotional dependency; inconsistent persona across sessions.R requires repeated or longitudinal evidence; single-session data are insufficient.
T—Translational Healthcare ValuePatient, caregiver, professional, workflow, organisational, and system outcomes.Use T when evaluating benefit beyond novelty, feasibility, usability, or intention to use.Persuasive novelty without benefit; hidden workload; inequitable access; poor sustainability or update burden.T requires meaningful comparative or real-world evidence before broad implementation.
Table 5. Operationalising HEART for future evaluation studies.
Table 5. Operationalising HEART for future evaluation studies.
HEART DimensionWhat to AssessPossible MethodsStakeholdersTemporal LevelIllustrative Outcomes
HClarity, factual boundaries, empathy, dialogue repair, explanation quality, multimodal coherence.Conversation analysis; expert review; adversarial and boundary scenarios; comprehension checks.Patients, caregivers, clinicians, communication specialists.Single-session and repeated interaction.Unsupported-claim rate; safe-response rate; repair success; user understanding of limitations.
ECalibrated trust, hallucination and omission risk, privacy, consent, oversight, accountability, failure handling.Severity-sensitive clinical review; trust-calibration tasks; privacy impact assessment; audit logs; escalation testing.Users, caregivers, clinicians, data protection officers, ethics committees, institutions.Pre-deployment, pilot, update, and post-deployment monitoring.Hallucination and omission rates; major or critical error rates; appropriate abstention and escalation; trust-calibration gap; privacy incidents.
ALanguage–perception–action grounding, movement safety, multimodal coordination, robustness, latency, and recovery.Simulation; task testing; hazard analysis; red-team scenarios; technical logging; supervised pilots.Engineers, clinicians, safety officers, patients, caregivers.Technical validation and controlled pilot evaluation.Grounding accuracy; unsafe-action rate; recovery success; latency thresholds; human-intervention rate.
RMemory accuracy, personalisation, sustained usefulness, user control, dependency, and novelty decline.Longitudinal logs; memory audits; repeated interviews; preference tracking; caregiver feedback.Users, family members, caregivers, care staff, developers.Multi-session and longitudinal use.False-memory rate; stable personalisation; sustained engagement; deletion success; dependency indicators.
TClinical, psychosocial, workflow, organisational, equity, cost, and implementation value beyond novelty.Comparative pilots; pragmatic studies; implementation outcomes; workflow observation; cost and workload analysis.Patients, professionals, managers, commissioners, caregivers.Pilot, real-world, and post-implementation phases.Quality-of-life or care outcomes; workload; workflow fit; equity; cost; sustainability; staff burden.
Note: Indicators are illustrative and should be adapted to robot type, user group, intended care function, deployment risk, and regulatory status. HEART does not prescribe a single composite score.
Table 6. Retrospective HEART comparison of two empirical LLM-SAR evidence strands.
Table 6. Retrospective HEART comparison of two empirical LLM-SAR evidence strands.
HEART DimensionBlavette et al.: Geriatric-Unit SAR [53,54]Irfan et al.: Companion Robot with Older Adults [50,51]
HPartial. LLM integration was associated with more error-free interactions, fewer comprehension failures, and higher interaction success; comprehensive clinical factuality and boundary testing were not primary outcomes.Concern. Participants identified value in natural and personalised dialogue, but also hallucinations, outdated or shallow responses, interruptions, latency, confusion, and frustration.
EPartial. Acceptability, usability, and user perceptions were examined, but severity-sensitive hallucination testing, formal trust calibration, privacy impact, and institutional escalation pathways were not comprehensively reported.Concern. Privacy, control over learned data, and transparency were user priorities; hallucination and worry were observed, while clinical governance and accountable deployment were outside the study scope.
APartial. System and audiovisual-tracking improvements supported interaction, but language-to-physical-action safety was not a central endpoint.Not assessed. Embodied cues and turn-taking were relevant, but systematic action-grounding, hazard, and recovery evaluation was not reported.
RPartial. Two evaluation waves demonstrated system development, but did not establish persistent individual memory, dependency risk, or longitudinal continuity for the same users.Partial. Memory and personalisation were central design requirements, but stable long-term in-the-wild performance and dependency effects remain unestablished.
TPartial. The studies provide real care-setting feasibility and interaction evidence, but do not establish clinical effectiveness, cost, workflow sustainability, or comparative benefit.Not assessed. The evidence informs design and user-facing risks rather than clinical, organisational, or economic effectiveness.
Table 7. Comparison of HEART with adjacent evaluation approaches.
Table 7. Comparison of HEART with adjacent evaluation approaches.
Existing ApproachWhat It Handles WellResidual Gap for LLM-Enabled SARs in HealthcareHEART Contribution and Boundary
Technology acceptance models [4]Perceived usefulness, ease of use, attitudes, and intention to use.Do not assess embodied risk, hallucination, dependency, privacy exposure, accountability, or healthcare value.Positions acceptance within trust calibration, continuity, and translational value, without replacing behavioural modelling or psychometric validation.
HRI usability and experience evaluation [6,7]Interaction quality, engagement, social presence, usability, and user experience.Often short-term and weakly connected to clinical boundaries, professional responsibility, or longitudinal care value.Adds clinical appropriateness, embodied risk, repeated interaction, and care-context fit, without replacing task-specific HRI protocols.
Trustworthy AI and responsible AI [64,65,66,67,68]Transparency, accountability, fairness, privacy, oversight, and risk governance.Often not specific to embodied social robots, vulnerable care relationships, or language-to-action coupling.Connects responsible-AI principles to socially credible dialogue, embodied autonomy, memory, and care reliance, without replacing legal compliance or institutional governance.
Implementation science [22,28,82]Workflow fit, adoption, sustainability, organisational readiness, and service integration.Does not directly address generative dialogue, hallucination, embodied social credibility, or LLM memory.Connects implementation outcomes with LLM-specific safety, relationship continuity, and translational care value, without replacing implementation frameworks or economic evaluation.
Medical AI and device regulation [64,65,66,67,68]Safety, performance, change control, documentation, risk management, and clinical governance.May focus on software/device performance rather than social interaction, relational effects, and embodied care communication.Adds communication, trust, personalisation, and care-relationship evaluation to regulatory thinking, without replacing medical-device classification, formal risk management, or clinical validation.
Biometric-wearable mapping and evaluation approaches [40]Continuous physiological sensing, behavioural effects, and real-world data capture.Do not evaluate socially credible generative dialogue, persistent conversational memory, or language-to-action execution.Uses sensing and privacy insights within E, A, and R while adding communication and care-reliance risks.
User-centred SAR frameworks for mental healthcare [41]Autonomy, competence, emotional experience, participatory design, contextual adaptation, personalisation, and long-term trust.Do not operationalise LLM hallucination, prompt-related failure, embodied action safety, or deployment gates.Retains user-centred design while specifying LLM-specific failure modes and healthcare evidence requirements.
Clinical LLM-safety frameworks [42]Error taxonomy, hallucination and omission rates, and severity-sensitive clinical review.Focus on generated clinical text rather than socially embodied interaction, physical action, memory, or care relationships.Adapts frequency-by-severity logic to patient-facing dialogue and connects it to trust, embodiment, continuity, and translation.
Table 8. Discussion implications and prioritised research directions for LLM-enabled SARs.
Table 8. Discussion implications and prioritised research directions for LLM-enabled SARs.
Discussion ThemeMain ImplicationFuture Research Priority
Beyond usability and acceptanceUsability, acceptability, and satisfaction are necessary but insufficient for healthcare deployment.Combine user-experience measures with trust calibration, autonomy, safety, clinical value, and ethical-risk indicators.
Implementation and healthcare translationRobots must fit clinical workflows, professional roles, institutional readiness, and care routines.Conduct pragmatic implementation studies including staff burden, training, maintenance, cost, workflow fit, and governance.
Longitudinal and real-world gapsShort-term pilots cannot determine sustained value, assimilation, dependency, privacy risk, or durability of benefit.Use longitudinal real-world designs with repeated measures, usage logs, caregiver/staff perspectives, and follow-up after novelty effects decline.
LLM-specific safety and governanceGenerative dialogue introduces risks of hallucination, overtrust, misleading capability claims, privacy exposure, and unclear accountability.Develop safety testing protocols, privacy-by-design architectures, escalation pathways, auditing procedures, and post-deployment monitoring.
Comparative and interdisciplinary evaluationLLM-enabled SARs should be evaluated against meaningful alternatives and across technical, clinical, ethical, and organisational domains.Compare LLM-enabled SARs with scripted robots, chatbots, telehealth, usual care, and human-administered interventions using interdisciplinary evaluation frameworks such as HEART.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Orehovački, T. The HEART Framework for LLM-Enabled Socially Assistive Robots in Healthcare: A PRISMA-Informed Structured Review. Appl. Sci. 2026, 16, 7904. https://doi.org/10.3390/app16167904

AMA Style

Orehovački T. The HEART Framework for LLM-Enabled Socially Assistive Robots in Healthcare: A PRISMA-Informed Structured Review. Applied Sciences. 2026; 16(16):7904. https://doi.org/10.3390/app16167904

Chicago/Turabian Style

Orehovački, Tihomir. 2026. "The HEART Framework for LLM-Enabled Socially Assistive Robots in Healthcare: A PRISMA-Informed Structured Review" Applied Sciences 16, no. 16: 7904. https://doi.org/10.3390/app16167904

APA Style

Orehovački, T. (2026). The HEART Framework for LLM-Enabled Socially Assistive Robots in Healthcare: A PRISMA-Informed Structured Review. Applied Sciences, 16(16), 7904. https://doi.org/10.3390/app16167904

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop