1. Introduction
Socially assistive robots (SARs) are embodied systems designed to support users through social interaction, guidance, companionship, coaching, monitoring, or motivation. They are studied in institutional, community, rehabilitation, and home-based care settings, where their value depends on technical reliability, verbal and non-verbal communication, personalisation, engagement, and fit with care relationships [
1,
2,
3].
Despite this breadth, evidence of healthcare value remains uneven. The literature commonly emphasises usability, acceptability, engagement, perceived usefulness, comfort, or short-term interaction quality, whereas sustained clinical, psychosocial, relational, organisational, and implementation outcomes are less well established [
4,
5,
6]. Trust is also underexamined when assessment relies on subjective measures and brief or controlled encounters rather than calibration in everyday care [
7]. These gaps are especially consequential for users whose age, health status, or cognitive, sensory, or functional impairments may increase vulnerability, the risk of dependency or privacy loss, and asymmetries of knowledge and authority.
LLM integration broadens the capabilities of socially assistive robots while making their evaluation more complex. An LLM-enabled SAR incorporates an LLM or another generative language model that performs a comparable role in one or more functions, including dialogue generation, language understanding, personalisation, retention of interaction context, multimodal interpretation, planning, or action selection. Integration may be limited to conversational functions or extend to system-level memory, multimodal reasoning, planning, and action-linked autonomy.
These functions may support open-ended conversation, contextual interpretation, multimodal interaction, and more flexible adaptation than earlier scripted or rule-based designs [
8,
9,
10]. At the same time, fluent output may make a robot appear more knowledgeable, empathic, socially competent, or clinically capable than its demonstrated performance warrants. Risks of hallucination, misleading capability claims, overtrust, privacy exposure, unclear accountability, and unsafe reliance therefore become central evaluation concerns in healthcare [
11,
12]. When generated language is coupled with social presence and possible physical action, failures can affect behaviour, care relationships, and safety rather than communication alone.
Prior reviews cover healthcare social robot characteristics, technical requirements, deployment settings, LLM-driven HRI capabilities, trust, acceptance, ethics, and implementation [
2,
8,
9]. Their contributions remain dispersed across the technical, clinical, relational, and governance literatures. Taken together, these reviews do not yet provide an integrated approach for judging systems in which generative communication, memory, adaptation, social credibility, physical agency, and healthcare accountability interact.
In response, this study proposes HEART as a healthcare-specific interpretive architecture rather than a psychometric scale or composite score. It organises evaluation across five distinct but interacting dimensions: Human-Centred Communication, Ethical and Trustworthy Deployment, Adaptive and Embodied Intelligence, Relationship Continuity, and Translational Healthcare Value. Explicit boundary rules assign primary evaluative objects, operational indicators translate the dimensions into assessable outcomes, and non-additive gates prevent strengths in one domain from masking critical safety or governance deficits in another.
Three research questions guide the analysis:
RQ1. What evaluative challenges emerge across the literature on healthcare robotics, socially assistive robots, human–robot interaction, and LLM-enabled robotic systems?
RQ2. How do LLMs and generative AI change the evaluation of socially assistive robots in healthcare, particularly in relation to communication, trust, embodiment, autonomy, continuity, and healthcare value?
RQ3. How can these evaluative challenges be synthesised into an integrative framework for assessing LLM-enabled socially assistive robots in healthcare?
To address these questions, the study combines PRISMA-informed evidence identification and selection with framework-oriented analysis of retained sources. The evidence base is organised into three functional corpora that distinguish background mapping, primary analytical studies, and governance or regulatory interpretation. This structure supports the development of an evaluative framework rather than pooled estimation of intervention effects. The resulting HEART architecture is intended to inform study design, comparative assessment, stakeholder involvement, implementation decisions, and future empirical validation.
3. Methods
3.1. Review Design
This review used a PRISMA-informed structured design that combined transparent source identification with framework-oriented analysis of LLM-enabled SARs. The aim was to organise a rapidly developing and heterogeneous literature rather than estimate pooled intervention effects. Established guidance on scoping, structured mapping, and review reporting informed the design [
43,
44,
45,
46]. The source base included reviews, empirical and technical studies, pilots, protocols, design work, preprints, and normative material across healthcare robotics, HRI, generative AI, ethics, implementation, and clinical evaluation.
PRISMA 2020 and PRISMA-ScR informed reporting of identification, deduplication, source-level assessment, inclusion, and flow documentation [
45,
46]. The qualifier PRISMA-informed reflects adaptation of these principles to a structured, framework-oriented review rather than a registered systematic review or meta-analysis. Accordingly, no pooled effect estimates or formal risk-of-bias meta-assessment were undertaken.
3.2. Information Sources
Web of Science Core Collection and Scopus supplied the controlled database coverage because both index the literature across healthcare, robotics, HRI, computer science, engineering, and social science. Google Scholar supported targeted supplementary identification of methodological references, arXiv material, governance or regulatory documents, and selected non-LLM SAR and healthcare-robot studies that were insufficiently represented in the controlled results. The purposive comparative baseline was not intended as an exhaustive search of the pre-LLM SAR literature. Publisher, arXiv, institutional, and governance websites served only as retrieval platforms after discovery. Citation tracking was not undertaken.
Sources identified through the supplementary route were interpreted according to evidentiary function: preprints informed emerging claims, whereas governance documents supported normative and regulatory interpretation rather than empirical effectiveness claims (
Section 3.9).
Exploratory queries conducted between 18 and 24 March 2026 calibrated terminology across healthcare robotics, HRI, LLMs, generative AI, ethics, implementation, and clinical evaluation. The main Web of Science and Scopus searches followed on 25–27 March. Duplicate removal and initial relevance checking took place from 28 March to 12 April, and consolidated source-level assessment and structured charting continued from 13 April to 5 May. Thematic analysis and framework development were conducted from 6 to 28 May, followed by manuscript drafting and revision between 29 May and 9 June. A targeted Google Scholar update on 10 June captured recent LLM-HRI, methodological, arXiv, governance, and regulatory sources. The literature was locked on 14 June 2026.
Coverage focused primarily on publications from 2021 to 2026, reflecting the recent emergence of generative-language and embodied-AI robotic systems. Earlier material was retained only when it supplied methodological, ethical, conceptual, or review-design foundations, including scoping-review methods, PRISMA reporting, and thematic synthesis.
3.3. Search Strategy
Four concept blocks structured the controlled database search: (1) social, socially assistive, and healthcare robots; (2) large language models, generative AI, and multimodal AI; (3) healthcare contexts; and (4) evaluation, ethics, trust, autonomy, implementation, safety, and clinical translation. Equivalent Boolean logic was implemented in Web of Science Core Collection Advanced Search and Scopus, with field syntax adapted to each platform. Broader WoS Smart Search and Scopus All Fields variants were tested during calibration but excluded from the numerical flow. Shorter Google Scholar queries targeted the supplementary categories described in
Section 3.2 and selected non-LLM comparators. This iterative formulation followed scoping and evidence-mapping approaches for heterogeneous fields [
43,
44,
47].
(“socially assistive robot*” OR “social robot*” OR “healthcare robot*” OR “care robot*” OR “human-robot interaction”) AND (“large language model*” OR LLM OR “generative AI” OR ChatGPT OR “multimodal AI”) AND (healthcare OR clinical OR hospital OR “older adult*” OR dementia OR pediatric OR nursing OR rehabilitation) AND (trust OR ethics OR acceptance OR autonomy OR evaluation OR implementation OR safety).
Database interfaces, field configurations, search dates, displayed record yields, and the supplementary Google Scholar queries are reported in
Supplementary Table S5. Audited counts derive only from the two controlled database configurations. Displayed Google Scholar totals were treated as approximate platform estimates; screening was targeted rather than exhaustive, and exact route-level additions and overlap were not retained.
Platform-specific syntax and wildcard conventions were applied to this expression in the audited searches. Materials located through Google Scholar were retrieved from the relevant publisher, arXiv, institutional, or regulatory website. Zotero Web Library was used for reference management and duplicate checking.
3.4. Study Selection and Evidence Flow
The selection flow combines an audited database-search branch with targeted Google Scholar retrieval and a preserved consolidated candidate-source log.
Figure 1 distinguishes these pathways and marks the transparency boundary created by the absence of exact counts for supplementary additions, overlap, and attrition.
The two database searches yielded 128 records: 47 from Web of Science Core Collection via Advanced Search without a field restriction and 81 from Scopus through title, abstract, and keyword queries. After 18 duplicates were removed, 110 unique items remained. The supplementary route added sources from the categories described in
Section 3.2, including selected non-LLM comparators used as a purposive baseline. Consolidation across routes produced a preserved coding log of 110 candidate sources. The matching totals are coincidental and do not indicate identical membership because additions and removals occurred without a retained route-level transition log. Source-level assessment excluded 18 candidates and retained 92 references. Of the exclusions, nine were peripheral to the healthcare SAR/LLM scope, three were overly broad or insufficiently specific to physical healthcare systems, four had unsuitable publication types or limited evidentiary contributions, and two were superseded or redundant.
Supplementary Table S6 provides the source-level reasons.
Seven of the 92 retained references were methodological sources used to justify review design, PRISMA-informed reporting, scoping logic, data charting, and thematic synthesis. They were cited but not included in substantive coding. This left 85 sources distributed across Corpus A (
n = 35), Corpus B (
n = 35), and Corpus C (
n = 15). Three contextual references added after the literature lock for framework comparison and metric operationalisation [
40,
41,
42] were also excluded from corpus counts.
Table 2 differentiates the controlled searches that generated the audited counts from exploratory interface variants, targeted Google Scholar identification, source-platform retrieval, and post-lock contextual updating.
3.5. Corpus Assembly and Classification
Analytical role determined the assignment of the 85 substantive sources to three complementary corpora. Corpus A contained reviews, surveys, and background or gap-mapping studies; Corpus B contained primary empirical, technical, pilot, protocol, and design-oriented work; and Corpus C contained governance, ethics, privacy, accountability, and regulatory material. The three groups supported landscape mapping, system-level analysis, and interpretation of trustworthy deployment, oversight, safety, and healthcare implementation, respectively.
Each source received one primary corpus assignment based on the function used most directly in the analysis. Mixed contributions retained secondary relevance through HEART tagging, allowing a paper to inform several dimensions without being counted in more than one corpus.
3.6. Eligibility Criteria
A source was eligible when it made a substantive contribution to healthcare robotics, socially assistive or generative AI-enabled robotic interaction, HRI, embodied autonomy, implementation, evaluation, safety, ethics, trust, or regulatory governance. Both topical relevance and analytical function were considered because the review combined heterogeneous source types.
Reviews and broad mappings of relevant research were assigned to Corpus A. Physical-robot studies involving care, generative AI, adaptation, multimodality, autonomy, or evaluation were assigned to Corpus B. Normative or regulatory material on trustworthy healthcare deployment, data protection, safety, or accountability was assigned to Corpus C.
Exclusions covered sources outside the healthcare, HRI, or robotics scope; disembodied chatbot papers without relevance to physical systems; industrial automation without human-facing care implications; and publications that did not contribute to the analytical objectives. Purely technical LLM work was retained only when it informed robotic autonomy, multimodal grounding, privacy, safety, or responsible deployment.
3.7. Data Extraction
A structured template informed by scoping and evidence-synthesis methods [
43,
47] was used to chart the retained sources. Recorded fields covered bibliographic details, study design, robot configuration, care or interaction context, target population, AI or LLM component, interaction modality, duration, evaluation methods, outcomes, ethical concerns, implementation issues, limitations, and implications for real-world healthcare use.
Empirical-study extraction emphasised characteristics that shaped interpretation: laboratory, simulated, clinical, educational, residential, or home setting; single-session, multi-session, or longitudinal interaction; participant group; and measurement through self-report, observation, behaviour, system performance, clinical indicators, or qualitative feedback.
Technical studies were charted according to the role of the LLM or generative component, including dialogue, planning, perception, action selection, memory, personalisation, multimodal grounding, and autonomous execution. Review-level extraction recorded the literature stream, principal conclusions, methodological limitations, gaps, and relevance to healthcare robots. Governance and ethics sources were examined for principles, risks, responsibilities, oversight mechanisms, and implications for trustworthy deployment.
The author performed duplicate checking, source-level assessment, corpus assignment, data extraction, HEART tagging, retrospective profiling, and thematic synthesis. No independent duplicate screening or coding was conducted. Explicit decision rules and the source-level crosswalks in
Supplementary Tables S1 and S2 were used to document these judgements.
3.8. Descriptive Mapping of the Evidence Base
Rather than applying formal bibliometric methods, the review used variables tailored to corpus function. Review type and gap-mapping contribution described Corpus A; publication year, primary source orientation, LLM status, context, and HEART relevance described Corpus B; and governance, ethical, privacy, accountability, and regulatory role described Corpus C. HEART relevance was an overlapping thematic tag, whereas publication year, source orientation, AI/LLM status, and primary context were mutually exclusive within Corpus B.
Corpus B formed the primary analytical set of 35 sources. Thirty-one (88.6%) were published between 2024 and 2026. Fifteen (42.9%) explicitly involved an LLM, foundation model, small language model, GPT-based component, or generative AI, while 20 (57.1%) provided non-LLM SAR, healthcare-robotics, AI-enabled, or HRI baselines. These comparators were selected purposively from the controlled results or targeted Google Scholar searching. Their role was to supply established knowledge on care relationships, workflow, longitudinal use, trust, and implementation rather than represent an exhaustive review of pre-LLM SAR research.
Table 3 reports mutually exclusive descriptive variables for Corpus B. Overlapping HEART tags are reported separately in
Supplementary Table S1 and compared by AI/LLM status in
Section 4.7. Across the 35 sources, T was represented in 22, R in 17, E in 15, H in 12, and A in 10; these values describe the coded analytical corpus rather than the prevalence of topics in the wider field.
3.9. Evidence Status and Claim Function
Because the source base was heterogeneous, a pooled risk-of-bias assessment was unsuitable. Interpretation instead used two descriptors. Evidence status distinguishes among completed peer-reviewed empirical studies, peer-reviewed technical or design work, reviews and syntheses, protocols or preprints, and governance or regulatory sources. Claim function differentiates observed user or clinical outcomes, observed technical performance, proposed designs or methods, planned evaluations, and normative requirements.
Section 4 and
Section 5 identify protocols, preprints, technical demonstrations, and governance material by function whenever maturity affects interpretation;
Supplementary Table S2 provides source-level descriptors.
This distinction preserves evidentiary differences without forcing heterogeneous materials onto a single quality hierarchy. Completed empirical studies support observed outcomes, technical work supports capability and failure-mode claims, protocols describe planned evaluation, preprints provide preliminary evidence, and governance documents establish normative requirements rather than effectiveness.
3.10. Thematic Synthesis
Recurring evaluative concerns were identified through an iterative process informed by qualitative evidence synthesis and thematic analysis [
48,
49]. Descriptive codes were expanded into broader analytical themes while preserving corpus function. Corpus A contributed landscape and gap codes; Corpus B contributed study-, system-, technical-, and interaction-level material; and Corpus C informed ethical, privacy, governance, and regulatory interpretation. This procedure enabled comparison without treating review-level, empirical, technical, and normative sources as equivalent.
Sources were initially coded by contribution, application area, evaluation focus, and limitations. Related codes were then grouped into themes such as communication quality, trustworthiness, physical interaction, adaptive behaviour, personalisation, autonomy, continuity, implementation, and healthcare value. Cross-corpus comparison was used to identify fragmentation and gaps relevant to LLM-enabled SAR evaluation.
Analysis centred on the changes introduced when an LLM is incorporated into a socially assistive robot. Acceptance, usability, and engagement were considered alongside social credibility, trust calibration, autonomy, responsibility, safety, and integration into care. Particular attention was paid to tensions between apparent capability and demonstrated reliability, short-term engagement and sustained value, personalisation and user control, conversational fluency and clinical appropriateness, and technical autonomy and professional accountability.
3.11. Framework Refinement and Methodological Boundaries
Initial HEART categories were compared with the literature streams in
Table 1 and the cross-corpus themes derived from the analysis. This comparison assessed whether the proposed dimensions captured recurring concerns in both the related-work mapping and the source-based review.
A dimension was retained when it addressed a recurring limitation, appeared across more than one corpus function, and captured an issue made more complex by generative language in a socially assistive healthcare robot. Applying these criteria refined HEART from an initial literature map into an integrative interpretive framework. The resulting framework is an evaluative structure rather than a validated measurement instrument.
The method prioritises transparent framework development over effect estimation or exhaustive coverage. Documentation preserves the exact controlled-database counts, a consolidated source-level exclusion log, retained methodological references, substantive corpus construction, and an explicit boundary around supplementary routes. The non-LLM subset was assembled purposively as a comparative baseline, not as a comprehensive review of the earlier SAR literature.
Section 6 addresses the remaining limitations.
4. The HEART Framework
HEART provides a healthcare-specific evaluative architecture for LLM-enabled socially assistive robots. Its five dimensions—abbreviated H, E, A, R, and T—are examined below through their primary objects, evidentiary requirements, and deployment implications.
The framework treats generated language, social credibility, memory, physical agency, governance, and care outcomes as interacting features of a healthcare intervention. HEART complements technical validation, clinical evaluation, ethical governance, implementation science, and regulatory oversight, but remains interpretive rather than a psychometrically validated scale or additive score.
4.1. Framework Logic, Boundaries, and Deployment Gates
The five dimensions are complementary but not interchangeable. Their primary objects are generated communication and user understanding (H), justified reliance and accountable deployment (E), coordination of language with perception and action (A), persistence and adaptation across encounters (R), and patient-, workflow-, organisational-, or system-level outcomes (T).
Table 4 specifies the boundaries, LLM-specific failure modes, and deployment implications.
Potential overlap is resolved through a primary-object rule. Multimodal coherence belongs to H when the outcome is comprehension or communicative consistency and to A when the outcome is technical synchronisation of speech, sensing, gesture, or movement. Personalisation belongs to R when continuity and adaptation are assessed and to E when consent, data minimisation, access, deletion, or privacy are assessed. Secondary cross-links may represent causal or governance relationships, but the same observation is not counted twice in a HEART profile.
Hallucinated or unsupported content is assigned to H when the evaluated outcome is factual content, comprehensibility, communicative repair, or user understanding. It is assigned to E when the evaluated outcome is potential clinical harm, justified reliance, oversight, escalation, or deployment responsibility. The same event may create a secondary cross-link, but its primary coding follows the outcome under evaluation.
Figure 2 visualises this logic by linking the system under evaluation to the five primary objects specified in
Table 4. The GATE markers show that E and A function as non-additive constraints: a critical failure in either dimension can block deployment even when other dimensions are favourable. The R and T rows additionally indicate evidence thresholds, requiring repeated interaction and outcomes beyond usability or acceptance, respectively.
4.2. H—Human-Centred Communication
Communication quality in healthcare cannot be inferred from fluency or conversational naturalness. H evaluates whether generated content is understandable, emotionally appropriate, context-sensitive, clinically bounded, and coherent with non-verbal behaviour. The central LLM-specific risk is that persuasive language can increase perceived competence despite weak factual support, absent clinical authority, or unreliable repair.
User and field studies document both potential and limitations. Studies with older participants indicate that users value continuity of relevant information, privacy controls, reminders, social connection, and empathic interaction. They also document interruptions, latency, repetition, superficiality, hallucinations, outdated information, confusion, frustration, and worry [
50,
51]. Comparative LLM-powered HRI research shows that robots can support connection-building and deliberative interaction, while participants also expect coherent non-verbal cues and identify weaknesses in logical communication [
52].
Geriatric evaluations report fewer comprehension failures and improved interaction success, acceptability, and usability after LLM integration [
53,
54]. Technical and user-evaluation studies additionally identify turn-taking, back-channelling, speech quality, gesture synchronisation, persona conditioning, multilingual adaptation, and source validation as relevant design features [
55,
56,
57,
58,
59]. Appropriate measures therefore include unsupported-claim rates, safe-response rates, repair success, explanation quality, and user understanding of limitations rather than fluency alone.
4.3. E—Ethical and Trustworthy Deployment
Safe and accountable deployment requires patients, caregivers, clinicians, and institutions to understand a robot’s capabilities and limits and to have defined means of governance, contestation, escalation, and oversight. E examines whether reliance is justified by demonstrated performance and whether privacy, consent, accountability, and human control are established before use. Trust should be calibrated to actual capability rather than maximised as an acceptance outcome [
7,
32].
LLM-enabled robots intensify this problem because fluent and confident outputs may be incorrect, incomplete, outdated, unsupported, or outside the robot’s intended role. A published ethical case study documents misleading capability claims in LLM-based care robots [
11]. For operational evaluation, the clinical-safety framework of Asgari et al. distinguishes hallucination and omission frequency from potential clinical severity and classifies errors such as fabrication, negation, contextual distortion, and causal error [
42]. Adapted to patient-facing SARs, E should therefore report rates of unsupported clinical claims and omissions, rates of major or critical errors, appropriate abstention and referral, and the gap between demonstrated reliability and user reliance.
Peer-reviewed ethical reviews and qualitative studies treat autonomy, consent, dignity, social presence, substitution of human care, dependency, equitable access, and stakeholder involvement as deployment requirements rather than optional design values [
33,
34,
36,
37,
60,
61]. These sources support normative and interpretive claims; they should not be read as evidence that a particular LLM-SAR has achieved safe deployment.
Privacy and governance are equally central because LLM-enabled robots may process speech, images, interaction histories, preferences, behavioural data, and health-related information. Technical and governance sources support data minimisation, privacy-aware memory, lifecycle monitoring, human oversight, change control, transparency, and accountable update procedures [
12,
62,
63,
64,
65,
66,
67,
68]. Privacy-sensitive personalisation belongs to E when the question concerns consent, collection, retention, access, deletion, or inference risk, and to R when the question concerns the accuracy and usefulness of personalisation across time.
4.4. A—Adaptive and Embodied Intelligence
When generated language can influence sensing, planning, movement, proximity, reminders, or action, the risk profile extends beyond dialogue quality. A evaluates the technical correctness and recoverability of that coupling. Fluent dialogue may mask an incorrectly grounded plan that produces inappropriate physical behaviour, delayed support, or unclear responsibility.
Studies of robotic autonomy show that LLMs can translate high-level language commands into lower-level control processes, support semantic planning, enable voice-based interaction, and integrate vision, speech, or proprioception. These capabilities also introduce problems of grounding, latency, reliability, error recovery, and safety [
39]. Humanoid platforms illustrate how instructions, environmental observations, execution feedback, and human input may be combined for adaptive behaviour [
69], while text-to-motion research demonstrates both the promise and limits of grounding GPT-based systems in movement [
70]. Coordinated text and gesture generation adds a further requirement for verbal and non-verbal coherence [
56].
User research adds an interactional requirement: language should align with gaze, gesture, timing, and physical presence because mismatches can reduce perceived competence or create confusion [
52]. The Nadine system illustrates how affective behaviour and human-like recall can be incorporated into a social robot while raising governance questions about memory and user expectations [
71]. Other LLM-integrated systems combine speech processing, summarisation of prior exchanges, persona conditioning, and multilingual adaptation, linking adaptive communication more closely to physical interaction and perception [
57].
Clinical settings impose a higher threshold. A GPT-reinforced patient communication robot illustrates the need for validated knowledge sources, dependability, and interaction safeguards when a physical system addresses health-related topics [
59]. Approaches based on clinical supervision likewise support personalisation without removing professional oversight or accountability [
72]. Adaptive and Embodied Intelligence therefore requires evidence that language, perception, planning, movement, and action are grounded, reliable, recoverable after errors, and appropriate to the care context.
4.5. R—Relationship Continuity
Appropriateness over time cannot be inferred from a single encounter. R examines whether stored information and adaptation remain accurate, controllable, responsive to changing needs, and useful after novelty declines. Persistent context may support reminders and shared history, but it can also preserve false information, expose sensitive data, or imply a stable understanding that the system cannot sustain.
Older participants expect companion robots to recall previous exchanges, tailor interaction, protect privacy, provide control over learned data, and adapt to social context [
50]. Open-domain systems may nevertheless fail to preserve coherent context or stable adaptation when interruptions, outdated responses, or hallucinations undermine confidence across sessions [
51]. Multi-session studies suggest that repeated interaction can support engagement but still requires assessment of accuracy, adaptation, and stability over time [
73]. Features such as persona conditioning and human-like recall may foster continuity while complicating authenticity, boundaries, and dependence [
57,
71].
Continuity also concerns cognitive support, well-being, and integration into daily routines. LLM-powered SARs are being explored for cognitive-health promotion, companionship, and adaptive engagement in elder care [
55,
58]. Collaborative drawing suggests that personalised creative activity may support social participation, although its value must persist beyond novelty [
74]. Home and real-world studies report varied trajectories of adoption, scepticism, use, constraints, and outcomes rather than a uniform pattern [
75,
76,
77].
Co-design work indicates that continuity should be shaped with users rather than imposed by designers. Older adults identify priorities involving companionship, autonomy, support, social engagement, and control over interaction [
78,
79], while personalised cognitive-support research emphasises alignment between robot functions, individual needs, and care goals [
80]. A preprint review of privacy-preserving LLM methods further supports data minimisation, transparency, user control, and protection against inference risks in systems that retain or reuse personal information [
81].
4.6. T—Translational Healthcare Value
Positive interaction outcomes do not by themselves establish healthcare value. T asks whether a robot addresses a defined care need without shifting hidden burdens to patients, caregivers, professionals, or institutions; usability, acceptance, engagement, and technical novelty remain necessary but insufficient indicators.
Evidence from geriatric and clinical settings illustrates this distinction. LLM-integrated robots may reduce comprehension failures, increase successful interactions, improve perceived usefulness, and support companionship, but short-term acceptability does not establish durable benefit, workflow fit, reliability, or safety [
53,
54]. Review-level work in age-related care reinforces the caution: feasibility is common, but benefits remain inconsistent across clinical and psychosocial outcomes [
20,
21]. Paediatric results likewise vary across procedures, measures, and intervention designs [
24,
25,
26].
Organisational fit is equally important. Professional perspectives indicate that adoption depends on patient benefit, role clarity, workload, safety, accountability, training, privacy, and integration into clinical routines [
28]. Nursing and care studies add staff capacity, maintenance, sterilisation, alarms, precision, navigation, communication, and responsibility boundaries [
19,
29,
30]. Work in geriatric and hospital settings further highlights institutional readiness, technical reliability, contextual fit, and sustained support [
3,
82].
Two peer-reviewed protocols describe planned effectiveness and cost-effectiveness evaluations rather than observed outcomes [
83,
84]. Additional intervention and implementation studies suggest possible contributions to cognitive support, residential care, group activities, physical activity, clinical nursing, and patient engagement [
85,
86,
87,
88]. Claims of translational value therefore require stronger comparative, longitudinal, economic, workload, and implementation evidence.
Appendix A,
Table A6 summarises the detailed evidence themes, representative sources, and unresolved evaluative gaps underpinning each HEART dimension.
4.7. Evidence Landscape Across the HEART Dimensions
Once the dimensions and their evaluative scope have been established, Corpus B reveals how analytical coverage is distributed. The 15 LLM-explicit and 20 non-LLM sources served complementary functions. LLM-explicit studies primarily informed generative communication, social credibility, stored context, multimodal coordination, and language-to-action risks, whereas baseline work contributed more strongly to longitudinal use, workflow integration, implementation, and care outcomes.
Figure 3 compares the overlapping HEART tags within the two subsets. H appeared in 10 of 15 LLM-explicit sources (66.7%) and 2 of 20 baseline sources (10.0%), whereas T appeared in five LLM-explicit sources (33.3%) and 17 baseline sources (85.0%). E and R were more evenly distributed, while A was more common in the LLM-explicit subset. These coding patterns describe the analytical corpus rather than topic prevalence across the wider field.
The comparison reveals three concentrations: recent LLM-enabled work on communication, social credibility, and embodied adaptation; age-related and geriatric studies of continuity, companionship, and repeated use; and non-LLM baseline research on implementation, workflow, and translational healthcare value. Few LLM-explicit studies combine generative-system testing with longitudinal, organisational, or outcome evaluation, leaving the connection between technical capability and real-world benefit underdeveloped.
4.8. Operationalising HEART
To make HEART usable in study design, assessment must operate at model, interaction, physical-system, longitudinal, and healthcare levels. Calibrated trust is not equivalent to a high trust score; it is the correspondence between demonstrated reliability, the user’s belief about that reliability, and willingness to act on the robot’s output. Hallucination assessment should report both frequency and potential clinical severity, distinguishing fabricated, negated, contextual, or causal errors and clinically important omissions [
42].
Table 5 translates the dimensions into assessable constructs, possible methods, relevant stakeholders, temporal levels, and illustrative outcomes.
User-facing work with older adults shows why model-level metrics must be combined with comprehension, privacy, and control over retained data [
50,
51]. Comparative LLM-powered HRI also indicates that social credibility and non-verbal cues shape perceived competence [
52].
A worked prospective example in
Supplementary Table S3 presents a 12-week evaluation of a medium-autonomy LLM-enabled SAR in residential care. The scenario limits the robot to bounded non-clinical dialogue, routine reminders, and companionship; clinical recommendations, care-plan modification, medication-related action, and unsupervised physical assistance are prohibited. The comparison is between an LLM-enabled SAR plus usual care and a scripted SAR plus usual care. The study design links dialogue safety, trust and privacy safeguards, controlled memory, hazard and override testing, repeated use, workflow, and healthcare-value outcomes to explicit decision gates. As a prospective template, it does not provide evidence of effectiveness.
Supplementary Table S4 offers a reusable record for study planning, evidence profiling, and deployment review.
4.9. Retrospective Application to Two Empirical LLM-SAR Evidence Strands
The two selected evidence strands permit an illustrative retrospective application of the operational scheme.
Table 6 compares companion-robot studies with older adults by Irfan et al. [
50,
51] and two-wave geriatric-unit evaluations by Blavette et al. [
53,
54]. The labels—supported, partial, concern, and not assessed—summarise the reported results and scope rather than numerical HEART scores.
The resulting profiles differ in ways relevant to evaluation design. Geriatric-unit studies provide stronger support for communication and early feasibility in a care setting, whereas companion-robot research reveals more explicit concerns involving hallucination, privacy, memory, and user expectations. Neither strand establishes translational effectiveness or comprehensive language-to-action safety; “not assessed” denotes an evidence gap rather than poor system performance.
4.10. Research Question Synthesis
The reviewed literature answers RQ1 by identifying a multidimensional evaluation problem spanning communication, justified reliance, language-to-action coordination, continuity, and healthcare value. With respect to RQ2, LLM integration connects generative language and retained context with social credibility, adaptation, and possible physical action under healthcare accountability. HEART addresses RQ3 by providing a structured basis for determining whether evidence is sufficient, incomplete, or blocked by a critical safety or governance concern. Across the three questions, responsible use depends on the combined evidence profile rather than performance on any single measure.
5. Discussion
5.1. Implications for Evaluation, Implementation, and Longitudinal Evidence
The central evaluative implication is that early interactional promise does not constitute evidence sufficient for deployment decisions. Positive attitudes, perceived usefulness, comfort, and engagement indicate feasibility but do not establish safety, trustworthiness, clinical value, or sustainability [
4,
5,
6,
7,
35,
89,
90,
91,
92]. For LLM-enabled SARs, socially responsive dialogue may inflate perceived competence, so acceptance results require complementary assessment of factual reliability, calibrated reliance, physical safety, continuity, and care outcomes.
Responsible implementation depends on more than technical performance. Role clarity, staff workload, training, maintenance, escalation, privacy governance, institutional readiness, equitable access, cost, and sustainable service models shape whether a system can be incorporated into care [
3,
19,
28,
29,
82,
84,
87,
88,
93]. Because generative dialogue can blur the boundaries between information, emotional support, clinical advice, and autonomous action, evaluation must also capture supervision, updates, workflow effects, and burdens transferred to professionals or caregivers.
Longitudinal assessment is therefore indispensable. Single-session or tightly controlled studies cannot reveal whether inaccurate stored context, hallucination, misleading capability claims, privacy exposure, overreliance, or emotional attachment accumulate through repeated use.
Evidence from real-world use indicates that outcomes vary with routine integration, individual trajectories, caregiver burden, staff needs, frailty, sensory or cognitive impairment, novelty decline, personal goals, and service sustainability [
75,
76,
77,
83,
84,
91,
94,
95]. Longitudinal designs should combine repeated measures, interaction logs, qualitative follow-up, professional and caregiver input, adverse-event monitoring, implementation outcomes, and durable care-value measures.
5.2. HEART and Adjacent Evaluation Approaches
No single adjacent approach addresses the full combination of generative communication, physical agency, memory, and healthcare accountability. HEART therefore complements rather than replaces HRI evaluation, technology-acceptance models, trustworthy-AI governance, implementation science, medical-device assessment, wearable-sensing approaches, user-centred SAR design, or clinical LLM-safety frameworks.
Table 7 compares their established contributions with the residual problems that arise when these concerns converge.
Across these approaches, HEART contributes a common boundary logic for tracing how failures in generated communication may affect reliance, physical action, longitudinal personalisation, and care delivery. Its healthcare grounding reflects vulnerable users, sensitive data, professional accountability, and the possibility of clinical harm. The novelty lies not in each criterion individually, but in assigning criteria to non-redundant primary objects, connecting them through causal and governance pathways, and constraining deployment through E and A gates.
HEART is presented as a healthcare-specific framework. Similar concerns may occur in other high-stakes applications, but transfer is not assumed. Any adaptation would require domain-specific definitions of value, stakeholders, accountability, and gate criteria, followed by separate empirical validation. Portability remains a research question rather than a claimed contribution.
5.3. Research Priorities for LLM-Enabled SARs
Five research priorities follow from the review. First, studies should specify the robot platform, model configuration and deployment mode, dialogue architecture, stored context, autonomy, intended users, care function, and prohibited claims [
9]. Second, LLM-enabled systems require comparison with scripted robots, disembodied chatbots, telehealth tools, human-administered interventions, and usual care [
8,
10]. Third, physical-system testing should assess grounding, latency, language-to-action execution, recovery, escalation, and regression after updates [
39].
Fourth, personalisation and memory-supported functions should be examined longitudinally through accuracy audits, user control over data, dependency indicators, repeated measures, and observation after novelty declines [
57,
73,
74]. Fifth, safety, privacy, governance, and healthcare value should remain core outcomes throughout the lifecycle [
11,
12,
42]. Recommended designs include adversarial prompts, vulnerable-user and clinical-boundary scenarios, severity-sensitive assessment of hallucinations and omissions, appropriate abstention and escalation, data minimisation, deletion testing, post-update validation, and monitoring of harmful or misleading outputs.
Table 8 consolidates these implications and research priorities.
5.4. Staged Evaluation and Translation Roadmap
To support practical decisions, the roadmap translates HEART into a staged process that extends from use-case definition to post-deployment monitoring.
Figure 4 arranges this process into seven sequential but revisitable stages, with arrows indicating the intended progression between them and advancement contingent on the evidence and risk criteria specified at each stage. Coloured H, E, A, R, and T markers indicate the primary HEART emphases at each stage, whereas grey markers denote dimensions that remain relevant but are not the primary focus. Stage 1 is marked H · E · A · R · T because intended use and claims define the communication role, governance boundaries, physical autonomy, memory and continuity assumptions, and healthcare-value criteria for the entire evaluation. It specifies target users, care setting, model and robot configuration, level of autonomy, memory functions, comparator, prohibited claims, and measurable success criteria. Stage 2 evaluates dialogue and model safety before user exposure through scenario-based testing of unsupported claims, clinically consequential omissions, severity of errors, appropriate abstention and escalation, adversarial prompts, privacy risks, and consistency with the robot’s stated role. At Stage 3, the physical system is verified through tests of language-to-action grounding, sensing and movement coordination, latency, interruption handling, recovery, action constraints, logging, and clinician or caregiver override. Unresolved critical E or A failures at these stages prevent progression.
Stage 4 introduces controlled user evaluation only after the pre-user safety gates are met. It examines comprehension of system limits, trust calibration, communication quality, task success, usability, emotional appropriateness, adverse responses, and the adequacy of repair and escalation, using representative users and relevant care professionals. Stage 5 extends evaluation across repeated encounters in a longitudinal care pilot. The focus shifts to memory accuracy and controllability, personalisation, changing needs, novelty decline, dependency, privacy-sensitive inference, caregiver and staff burden, and the stability of benefits and risks over time. Predefined progression criteria should specify which findings require redesign, restricted use, additional evidence, or termination.
Stage 6 evaluates implementation and healthcare value under pragmatic conditions and against a meaningful comparator. Outcomes should include patient or psychosocial benefit, workflow fit, staff and caregiver effects, cost, maintenance, equity, sustainability, residual risk, and organisational accountability. Although T is the primary translational decision focus, evidence from H, E, A, and R is considered jointly, as shown by H · E · A · R · T in the figure. Stage 7 introduces post-deployment monitoring of incidents, complaints, harmful or misleading outputs, model or data drift, memory errors, changes in user reliance, software and model updates, and emerging workflow effects. Each stage ends with an explicit decision to proceed, redesign, constrain, collect additional evidence, or stop. Major updates or adverse findings return the system to the relevant earlier stage rather than being treated as automatic progression along a simple maturity ladder.
5.5. Empirical Validation of HEART
Empirical testing should proceed in stages rather than begin with a composite score. First, a multidisciplinary content-validity and feasibility study should assess whether the dimension definitions, boundary rules, indicators, status labels, and deployment gates are relevant, comprehensive, understandable, and workable. Participants should include patients, caregivers, clinicians, HRI researchers, LLM-safety specialists, privacy and ethics experts, healthcare organisations, and regulators. A modified Delphi process combined with cognitive interviews and pilot use of
Supplementary Table S4 would support item refinement and identify missing or redundant criteria.
Second, independent raters should apply HEART to standardised evidence packages for a heterogeneous sample of LLM-SAR systems and scenarios. Inter-rater agreement should be evaluated separately for dimension assignment, evidence-status labels, risk identification, and deployment decisions. Construct validity can then be examined through preregistered hypotheses. For example, systems with documented memory leakage should receive less favourable E profiles, systems with language-to-action failures less favourable A profiles, and longitudinal real-world studies more complete R and T evidence than single-session laboratory studies.
Third, prospective multi-site studies should test utility and decision validity by comparing HEART-guided evaluation with usual study planning or governance review. Relevant outcomes include additional risks detected, changes to outcome selection, agreement among stakeholders, time and burden required, safety incidents identified before deployment, and correspondence between pre-deployment HEART profiles and subsequent real-world events. Reassessment after model, memory, or robot-platform updates would test responsiveness to known changes. Numerical weighting or an overall HEART score should be considered only after the reliability, validity, and practical consequences of the individual dimensions and decision rules have been established.
6. Limitations
The scope and design impose several limits on interpretation. This PRISMA-informed structured review is neither a systematic review nor a meta-analysis, and it does not estimate pooled effects or establish clinical effectiveness. Its conclusions concern the organisation and interpretation of a heterogeneous, rapidly developing source base.
The audited database branch comprises 128 records, 18 duplicate removals, and 110 unique items. Targeted Google Scholar searching added methodological, arXiv, governance, regulatory, and selected non-LLM comparative sources. After consolidation, a separate coding log contained 110 candidates, of which 18 were excluded and 92 retained. Seven methodological references were then separated from substantive coding, leaving 85 sources. Because route-level additions, overlap, and attrition for the supplementary pathway were not preserved, the consolidated set cannot be reconstructed as a database-only chain.
Figure 1 separates the branches and marks this transparency boundary; the design should therefore be interpreted as PRISMA-informed rather than as a systematic review with a fully reproducible search chain.
Study design, population, setting, robot type, technological maturity, and outcome measures varied substantially.
Supplementary Table S2 reports evidence status and claim function, but empirical studies, technical demonstrations, protocols, reviews, and normative documents remain methodologically non-equivalent.
Direct evidence on LLM-enabled healthcare SARs remains limited. Fifteen of the 35 primary analytical sources explicitly involved an LLM or generative AI component, and few assessed longitudinal or translational value. The purposively selected non-LLM subset established requirements for care, workflow, continuity, and implementation; it was not treated as evidence of LLM safety or effectiveness, or as an exhaustive account of earlier SAR research.
Corpus assignment, HEART tagging, retrospective profiling, and thematic synthesis involved interpretive judgement. One author performed all screening and coding, so inter-rater reliability could not be assessed. Boundary rules and source-level crosswalks strengthen traceability but cannot eliminate subjectivity. HEART has not yet been prospectively tested for reliability, validity, or predictive utility.
Publication, language, database, and grey-literature constraints also apply, and machine-exported interface histories were not retained.
Supplementary Tables S5 and S6 document the preserved Boolean expression, selected interfaces and fields, displayed yields, exploratory variants, Google Scholar queries, and 18 source-level exclusions, but they cannot reconstruct unavailable route-level transitions or imply exhaustive screening. The three post-lock contextual sources [
40,
41,
42] remained outside corpus counts. HEART does not replace formal risk management, clinical validation, data-protection assessment, medical-device regulation, or institutional governance. Its applicability outside healthcare was not examined.
7. Conclusions
Research on LLM-enabled socially assistive robots is advancing faster than the evidence for durable healthcare benefit. Current studies provide stronger support for feasibility, interaction quality, and early acceptability than for longitudinal safety, governance of retained user context, comparative effectiveness, workflow impact, cost, equity, or post-deployment performance. This imbalance matters because generative output is delivered through socially credible systems that may retain information, shape behaviour, and initiate or influence physical action in settings governed by clinical and organisational accountability.
HEART addresses this evaluative challenge by linking five complementary domains: communication, trustworthy deployment, adaptive embodiment, relationship continuity, and healthcare value. Its boundary rules separate primary objects of assessment; operational indicators and qualitative labels expose missing evidence; and non-additive E and A gates prevent favourable results elsewhere from compensating for critical governance or safety failures.
Conceptually, HEART treats LLM-enabled SARs as integrated socio-technical interventions in care rather than conversational products. Practically, it offers a common structure for study design, system comparison, stakeholder review, and implementation decisions. Retrospective profiling shows that communication or feasibility gains can coexist with unresolved longitudinal, translational, or safety questions. Prospective validation should now test content validity, inter-rater reliability, construct validity, decision-making utility, multi-site applicability, and responsiveness to model or platform updates before any overall score is considered.