Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

Article Types

Countries / Regions

Search Results (36)

Search Parameters:
Keywords = task-oriented dialogue

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
27 pages, 483 KB  
Article
A Conversational Agent with Hybrid NLU and Dual-Corpus RAG for KPI Alert Management in Tourism Business Intelligence
by Alberto Jiménez-Sánchez, Clara Rodríguez-Marcos, Silvia Domínguez-Castro, Albano Carrera and Ricardo S. Alonso
Electronics 2026, 15(16), 3704; https://doi.org/10.3390/electronics15163704 - 19 Aug 2026
Viewed by 224
Abstract
Monitoring operational Key Performance Indicators (KPIs) in Business-to-Business (B2B) tourism platforms demands continuous reconfiguration of alert systems, a task that conventional interfaces render inaccessible to non-technical stakeholders confronted with multi-screen forms and proprietary identifiers. This paper presents a conversational agent that lets such [...] Read more.
Monitoring operational Key Performance Indicators (KPIs) in Business-to-Business (B2B) tourism platforms demands continuous reconfiguration of alert systems, a task that conventional interfaces render inaccessible to non-technical stakeholders confronted with multi-screen forms and proprietary identifiers. This paper presents a conversational agent that lets such users create, modify, list, and explain KPI alerts through natural language, while guaranteeing the structural correctness of every configuration. Its core contribution is a schema-derived slot-completeness model of nine slot groups constraining a Large Language Model (LLM) tool-calling agent, paired with normalisation patterns derived at runtime from live database metadata and a bounded validation loop returning field-level errors to the model. A dual-corpus Retrieval-Augmented Generation module grounds the agent’s knowledge branch in schema documentation and a JSON-LD ontology, while configuration is grounded in live metadata; a human-in-the-loop checkpoint precedes every commit. The system is deployed as a prototype and evaluated in an automated pilot over a 50-utterance corpus, where it reaches 81.2% exact configuration match against 38.5% for the strongest unconstrained baseline (+42.7 percentage points, McNemar p<0.001) and emits no invalid schema identifier, against 15.8% for that baseline. A single-layer ablation locates the effect in the schema-aware tool layer. Corpus, annotations, prompts and evaluation scripts are released for replication. Full article
(This article belongs to the Special Issue AI-Driven Frameworks for Human–Computer Interaction)
Show Figures

Figure 1

35 pages, 10090 KB  
Article
Comparing LLM-Driven and Script-Based Non-Player Characters Under Controlled Information Boundaries: Effects on Task Performance, Immersion, and Satisfaction
by Yanzhen Li, Jing Deng, Jiaxiang Zhao and Jinho Yim
Appl. Sci. 2026, 16(14), 7254; https://doi.org/10.3390/app16147254 - 20 Jul 2026
Viewed by 455
Abstract
This study examines NPC dialogue mechanisms as an independent design factor in task-oriented digital games. Although generative NPCs are often discussed in relation to narrative richness, conversational naturalness, and social presence, limited controlled evidence exists on whether dialogue mechanisms themselves affect player outcomes [...] Read more.
This study examines NPC dialogue mechanisms as an independent design factor in task-oriented digital games. Although generative NPCs are often discussed in relation to narrative richness, conversational naturalness, and social presence, limited controlled evidence exists on whether dialogue mechanisms themselves affect player outcomes when task information is held constant. To address this gap, the study combined a formative user study with a single-factor between-subject experiment using a controlled game prototype. The LLM-NPC and Script NPC conditions shared identical task goals, scene structure, completion criteria, predefined fact set, and task-stage prompting boundaries. The experimental contrast was therefore framed as a comparison between a script-based branching dialogue interface and a boundary-controlled LLM-based adaptive dialogue interface, rather than as a comparison involving unequal task information. The results showed that, under controlled information boundaries, LLM-NPCs significantly improved task efficiency, as reflected in shorter completion time, fewer NPC inquiries, and fewer incorrect searches. For immersion, the overall multivariate effect was not statistically significant, and no confirmatory evidence of an LLM-NPC advantage in immersion was found. For satisfaction, the LLM-NPC condition showed significantly higher total satisfaction, which was specified as the primary satisfaction endpoint. Secondary affective and overall evaluative satisfaction outcomes showed consistent uncorrected differences in the same direction. These findings suggest that, under the implemented factual and procedural controls, the value of LLM-NPCs in this task-oriented setting is less likely to lie in providing more information than in reorganizing predefined task facts in a more context-sensitive and player-adaptive manner. Boundary-controlled adaptive dialogue may therefore provide a useful design perspective for future task-oriented NPC systems. Full article
(This article belongs to the Special Issue Advances in Games and Immersive Technologies)
Show Figures

Figure 1

21 pages, 524 KB  
Review
Explainable Conversational Agents for Mobile Health Coaching Systems: Trust Factors, Progress and Opportunities
by Luminous Ogochukwu Akazua, Jianlong Zhou, Fang Chen, Niusha Shafiabady, George Tian, Andreas Holzinger and Heimo Müller
Mach. Learn. Knowl. Extr. 2026, 8(6), 144; https://doi.org/10.3390/make8060144 - 25 May 2026
Viewed by 706
Abstract
Background: Artificial Intelligence (AI) and Machine Learning (ML) technologies, such as conversational agents, are becoming increasingly essential tools across multiple industries, particularly in healthcare. This paper presents a scoping review (PRISMA-ScR) of conversational agents (CAs) in mobile health coaching systems (MHCS). It [...] Read more.
Background: Artificial Intelligence (AI) and Machine Learning (ML) technologies, such as conversational agents, are becoming increasingly essential tools across multiple industries, particularly in healthcare. This paper presents a scoping review (PRISMA-ScR) of conversational agents (CAs) in mobile health coaching systems (MHCS). It examines existing applications of MHCS, focusing on development strategies, usage contexts, impacts on users, benefits, and research gaps, emphasizing the ability of explainable artificial intelligence (XAI) in making health guidance and decision-support recommendations transparent, trustworthy, and interpretable, if properly integrated. This scoping review identifies opportunities to maximize the use of conversational agents, explainable AI, and mobile technologies to make mobile health coaching systems more accessible and trustworthy, as well as further research gaps worth exploring. Objective: This scoping review maps the evidence on CAs and XAI-enabled technologies in MHCS, identifies trust-related design criteria, categorizes reported outcomes, and highlights opportunities for explainable conversational agents (XCA) in a mobile health context, especially in tackling general medical conditions pertinent in underserved settings. Eligibility criteria: Reported eligible resources evaluated, designed, or conceptually analyzed existing CAs, XAI techniques, and MHCS, AI-supported medical dialogue systems, e-coaching systems, and mobile health applications. We considered sources only relevant to healthcare, health coaching, trust, explainability, or patient engagement that were published between 2006 and 2025. Sources of Evidence: Searches were conducted in IEEE Xplore, Google Scholar, Springer, ScienceDirect/Elsevier, ProQuest, and ACM Digital Library, supplemented by targeted web searches and backward citation checks. Charting methods: Data were charted by system type, communication mode, health context, operational mode, technology used, XAI/trust features, degree of automation, study designs and outcome classification. We applied a revised outcome classification: generated desired outcome (GDO) and partially generated desired outcome (P-GDO), and did not generate desired outcome (DN-GDO). Results: A total of 201 resources were collected. Charted studies clustered around CAs in health, MHCS for chronic diseases and stress management, XAI methods such as LIME, SHAP, Prospector, and counterfactual explanations, and trust-related elements such as voice quality, communication style, appearance, social intelligence, privacy, and performance quality. Most health CAs and MHCS addressed chronic diseases, mental health, or behavior change; fewer addressed general medical diagnosis or autonomous mobile-based primary care support. Conclusions: Existing evidence suggests that CAs and MHCSs can support engagement, coaching, education, and selected decision-support tasks, but evidence for safe, autonomous, explainable general practice functionality remains limited. Future work should prioritize clinically supervised XCA designs, core safety assessment, interfaces with transparent explanation, data protection, culturally and linguistically responsive implementation, and future-oriented review in underserved mobile health settings. Full article
(This article belongs to the Section Thematic Reviews)
Show Figures

Figure 1

17 pages, 2168 KB  
Article
Benchmarking Sparse and Dense Models for Deception Detection in Negotiation: A Context-Aware and Imbalance-Sensitive Approach
by Jae-Uk Kim, Beom Jun Go, Hwan Soo Yu and Soo Young Cho
Appl. Sci. 2026, 16(9), 4301; https://doi.org/10.3390/app16094301 - 28 Apr 2026
Viewed by 486
Abstract
Automatic detection of deceptive intent in negotiation dialogue remains difficult because deceptive utterances are rare, context-dependent, and pragmatically subtle. This study develops a deployment-oriented evaluation pipeline for negotiation analytics using the Diplomacy corpus and compares sparse, dense, and imbalance-aware neural models under a [...] Read more.
Automatic detection of deceptive intent in negotiation dialogue remains difficult because deceptive utterances are rare, context-dependent, and pragmatically subtle. This study develops a deployment-oriented evaluation pipeline for negotiation analytics using the Diplomacy corpus and compares sparse, dense, and imbalance-aware neural models under a unified protocol. The pipeline integrates context-window benchmarking, validation-based threshold selection, 10-seed robustness analysis, model-agnostic explanation case studies, and controlled perturbation stress tests. Across binary speaker-intention and receiver-perception tasks, contextualized inputs consistently outperform isolated utterances, confirming that deception-related interpretation is inherently sequential. The sparse term frequency–inverse document frequency (TF-IDF) model remains the strongest and most efficient overall benchmark, whereas stronger imbalance-aware neural baselines can improve minority deceptive-instance sensitivity at substantially higher computational cost. Error analysis further shows that socially mediated deception-quadrant prediction is markedly harder than direct binary in-tent prediction, with many failures collapsing toward majority straightforward cases. Controlled perturbation tests show that sparse modeling is especially stable under light-weight surface corruption, while neural robustness remains architecture-dependent. The main contribution is therefore a practical decision framework for selecting among efficient sparse monitoring, dense baselines, and minority-sensitive neural detection under operational constraints. Full article
Show Figures

Figure 1

30 pages, 5007 KB  
Article
Developing a Protocol-Based Expressive Therapies Continuum Assessment Profile (ETC-AP): Current Achievements and Future Perspectives
by Elza Strazdiņa, Viktorija Perepjolkina, Anda Upmale-Puķīte, Elīna Akmane, Jana Duhovska and Kristīne Mārtinsone
Behav. Sci. 2026, 16(5), 640; https://doi.org/10.3390/bs16050640 - 24 Apr 2026
Viewed by 1230
Abstract
Art therapy assessment benefits from analytical clarity while preserving non-directive, process-sensitive practice. Although the Expressive Therapies Continuum (ETC) is widely used to conceptualize sensory, affective, cognitive, and symbolic processes in art-making, ETC-informed assessment often relies on implicit clinical reasoning, limiting transparency and interdisciplinary [...] Read more.
Art therapy assessment benefits from analytical clarity while preserving non-directive, process-sensitive practice. Although the Expressive Therapies Continuum (ETC) is widely used to conceptualize sensory, affective, cognitive, and symbolic processes in art-making, ETC-informed assessment often relies on implicit clinical reasoning, limiting transparency and interdisciplinary communication. This article presents the developmental stage of a protocol-based Expressive Therapies Continuum Assessment Profile (ETC-AP) developed at Rīga Stradiņš University. The ETC-AP differentiates activation and inhibition patterns around integration midpoints and organizes observation in a defined five-step interpretive sequence without positioning the method as a psychometrically validated test. It combines (i) a uniform three-task, non-directive administration with a brief post-task inquiry; (ii) criteria-guided coding of observable features across three artworks and process notes; and (iii) 0–100 descriptive profile indicators to support within-case pattern description and professional dialogue. An illustrative case vignette shows how the ETC-AP can generate trauma-informed, regulation-oriented hypotheses about channel accessibility and cautious regulation-oriented sequencing, while remaining subordinate to clinical judgment and context. Key boundaries include incomplete operational coverage in some inhibition ranges, limits of static documentation for process-dependent markers, and the need for structured training materials and programmatic studies of reliability, feasibility, and sensitivity to change. Full article
Show Figures

Figure 1

25 pages, 4998 KB  
Article
Pareto-Aware Dual-Preference Optimization for Task-Oriented Dialogue
by Shenghui Bao and Mideth Abisado
Symmetry 2026, 18(2), 372; https://doi.org/10.3390/sym18020372 - 17 Feb 2026
Viewed by 1376
Abstract
Task-oriented dialogue systems face a tension between comprehensive constraint elicitation (task adequacy) and conversational efficiency (minimizing turns). Current preference learning frameworks treat preferences as static, unable to capture the dynamic evolution of interaction states that evolve across dialogue progression. We present Dual-DPO, a [...] Read more.
Task-oriented dialogue systems face a tension between comprehensive constraint elicitation (task adequacy) and conversational efficiency (minimizing turns). Current preference learning frameworks treat preferences as static, unable to capture the dynamic evolution of interaction states that evolve across dialogue progression. We present Dual-DPO, a framework that embeds multi-objective preferences into data construction via turn-aware scoring. Our approach decouples objective balancing from policy updates through offline preference scalarization, addressing the optimization instability challenges in online multi-objective reinforcement learning. Experiments on MultiWOZ 2.4 demonstrate 28–35% dialogue turn reduction while maintaining Joint Goal Accuracy > 89% (p<0.001). Pareto frontier analysis shows 94% coverage with hypervolume HV=0.847. Independent expert evaluation by 10 PhD-level researchers (n=300 assessments, inter-rater agreement α=0.78) confirms 32% user satisfaction improvement (p<0.001). Theoretical analysis demonstrates that offline scalarization, which correlates with improved optimization stability, achieves 3.2× lower gradient variance than online multi-reward optimization by eliminating sampling stochasticity. Our approach enables balanced treatment of competing objectives through Pareto-optimal trade-offs. These results highlight a symmetric and balanced treatment of competing objectives within a Pareto-optimal optimization framework. Full article
(This article belongs to the Section A: Computer Science)
Show Figures

Figure 1

12 pages, 241 KB  
Article
Being Church Together? Exploring the Logic of Intercultural Theology and Ministry
by Daniel John Pratt Morris-Chapman
Religions 2025, 16(12), 1554; https://doi.org/10.3390/rel16121554 - 9 Dec 2025
Viewed by 921
Abstract
The shift away from mission studies to intercultural theology within a number of universities coincides with the emergence of postmodernism. This article explores the extent to which a postmodern outlook pervades intercultural theology and explores whether or not an alternative epistemological orientation, particularism, [...] Read more.
The shift away from mission studies to intercultural theology within a number of universities coincides with the emergence of postmodernism. This article explores the extent to which a postmodern outlook pervades intercultural theology and explores whether or not an alternative epistemological orientation, particularism, might be better suited to the task of bringing diverse cultures and languages into dialogue and, moreover, uniting Christian congregations. Full article
17 pages, 1327 KB  
Article
MA-HRL: Multi-Agent Hierarchical Reinforcement Learning for Medical Diagnostic Dialogue Systems
by Xingchuang Liao, Yuchen Qin, Zhimin Fan, Xiaoming Yu, Jingbo Yang, Rongye Shi and Wenjun Wu
Electronics 2025, 14(15), 3001; https://doi.org/10.3390/electronics14153001 - 28 Jul 2025
Viewed by 2989
Abstract
Task-oriented medical dialogue systems face two fundamental challenges: the explosion of state-action space caused by numerous diseases and symptoms and the sparsity of informative signals during interactive diagnosis. These issues significantly hinder the accuracy and efficiency of automated clinical reasoning. To address these [...] Read more.
Task-oriented medical dialogue systems face two fundamental challenges: the explosion of state-action space caused by numerous diseases and symptoms and the sparsity of informative signals during interactive diagnosis. These issues significantly hinder the accuracy and efficiency of automated clinical reasoning. To address these problems, we propose MA-HRL, a multi-agent hierarchical reinforcement learning framework that decomposes the diagnostic task into specialized agents. A high-level controller coordinates symptom inquiry via multiple worker agents, each targeting a specific disease group, while a two-tier disease classifier refines diagnostic decisions through hierarchical probability reasoning. To combat sparse rewards, we design an information entropy-based reward function that encourages agents to acquire maximally informative symptoms. Additionally, medical knowledge graphs are integrated to guide decision-making and improve dialogue coherence. Experiments on the SymCat-derived SD dataset demonstrate that MA-HRL achieves substantial improvements over state-of-the-art baselines, including +7.2% diagnosis accuracy, +0.91% symptom hit rate, and +15.94% symptom recognition rate. Ablation studies further verify the effectiveness of each module. This work highlights the potential of hierarchical, knowledge-aware multi-agent systems for interpretable and scalable medical diagnosis. Full article
(This article belongs to the Special Issue Advanced Techniques for Multi-Agent Systems)
Show Figures

Figure 1

17 pages, 6837 KB  
Article
Mitigating LLM Hallucinations Using a Multi-Agent Framework
by Ahmed M. Darwish, Essam A. Rashed and Ghada Khoriba
Information 2025, 16(7), 517; https://doi.org/10.3390/info16070517 - 21 Jun 2025
Cited by 8 | Viewed by 20443
Abstract
The rapid advancement of Large Language Models (LLMs) has led to substantial investment in enhancing their capabilities and expanding their feature sets. Despite these developments, a critical gap remains between model sophistication and their dependable deployment in real-world applications. A key concern is [...] Read more.
The rapid advancement of Large Language Models (LLMs) has led to substantial investment in enhancing their capabilities and expanding their feature sets. Despite these developments, a critical gap remains between model sophistication and their dependable deployment in real-world applications. A key concern is the inconsistency of LLM-generated outputs in production environments, which hinders scalability and reliability. In response to these challenges, we propose a novel framework that integrates custom-defined, rule-based logic to constrain and guide LLM behavior effectively. This framework enforces deterministic response boundaries while considering the model’s reasoning capabilities. Furthermore, we introduce a quantitative performance scoring mechanism that achieves an 85.5% improvement in response consistency, facilitating more predictable and accountable model outputs. The proposed system is industry-agnostic and can be generalized to any domain with a well-defined validation schema. This work contributes to the growing research on aligning LLMs with structured, operational constraints to ensure safe, robust, and scalable deployment. Full article
(This article belongs to the Special Issue Intelligent Agent and Multi-Agent System)
Show Figures

Figure 1

14 pages, 1659 KB  
Article
Multi-HM: A Chinese Multimodal Dataset and Fusion Framework for Emotion Recognition in Human–Machine Dialogue Systems
by Yao Fu, Qiong Liu, Qing Song, Pengzhou Zhang and Gongdong Liao
Appl. Sci. 2025, 15(8), 4509; https://doi.org/10.3390/app15084509 - 19 Apr 2025
Cited by 3 | Viewed by 4349
Abstract
Sentiment analysis is pivotal in advancing human–computer interaction (HCI) systems as it enables emotionally intelligent responses. While existing models show potential for HCI applications, current conversational datasets exhibit critical limitations in real-world deployment, particularly in capturing domain-specific emotional dynamics and context-sensitive behavioral patterns—constraints [...] Read more.
Sentiment analysis is pivotal in advancing human–computer interaction (HCI) systems as it enables emotionally intelligent responses. While existing models show potential for HCI applications, current conversational datasets exhibit critical limitations in real-world deployment, particularly in capturing domain-specific emotional dynamics and context-sensitive behavioral patterns—constraints that hinder semantic comprehension and adaptive capabilities in task-driven HCI scenarios. To address these gaps, we present Multi-HM, the first multimodal emotion recognition dataset explicitly designed for human–machine consultation systems. It contains 2000 professionally annotated dialogues across 10 major HCI domains. Our dataset employs a five-dimensional annotation framework that systematically integrates textual, vocal, and visual modalities while simulating authentic HCI workflows to encode pragmatic behavioral cues and mission-critical emotional trajectories. Experiments demonstrate that Multi-HM-trained models achieve state-of-the-art performance in recognizing task-oriented affective states. This resource establishes a crucial foundation for developing human-centric AI systems that dynamically adapt to users’ evolving emotional needs. Full article
Show Figures

Figure 1

19 pages, 979 KB  
Article
A Conversational Agent for Empowering People with Parkinson’s Disease in Exercising Through Motivation and Support
by Patricia Macedo, Rui Neves Madeira, Pedro Albuquerque Santos, Pedro Mota, Beatriz Alves and Carla Mendes Pereira
Appl. Sci. 2025, 15(1), 223; https://doi.org/10.3390/app15010223 - 30 Dec 2024
Cited by 4 | Viewed by 3200
Abstract
Parkinson’s disease (PD) is a neurodegenerative disorder characterized by motor and non-motor symptoms. The MoveONParkinson project aims to enhance exercise engagement among people with Parkinson’s Disease (PwPD) in the Portuguese context through the ONParkinson digital platform, which provides mobile and web interfaces. While [...] Read more.
Parkinson’s disease (PD) is a neurodegenerative disorder characterized by motor and non-motor symptoms. The MoveONParkinson project aims to enhance exercise engagement among people with Parkinson’s Disease (PwPD) in the Portuguese context through the ONParkinson digital platform, which provides mobile and web interfaces. While the broader MoveONParkinson project has been previously described from a health-focused perspective, this study specifically focuses on the development and integration of an AI-driven conversational agent (CA) for the Portuguese language, called PANDORA, within the mobile interface of the solution to assist and motivate PwPD in their exercise routines. PANDORA (Parkinson Assistant in Natural Dialogue and Oriented by Rules and Assessments), designed based on Self-Determination Theory (SDT), addresses the psychological needs of autonomy, competence, and relatedness. A preliminary study involving 20 PwPD, 10 caregivers, and 5 healthcare professionals informed the design requirements for PANDORA. The development process involved four main phases: (1) Design of the Chatbot’s Motivation Model, (2) Design and implementation of the conversational agent, (3) Technical Performance Evaluation, and (4) User Experience Evaluation. Technical Performance Evaluation, conducted with three physiotherapists, assessed domain coverage, coherence response capacity, and dialog management capacity, achieving 100% accuracy in domain coverage and coherence response capacity and 89% in dialog management capacity. The User Experience Study involved eight PwPD users recruited from Portuguese healthcare units performing predefined tasks, with user satisfaction scores ranging from 4.2 to 4.9 on a five-point Likert scale. The findings indicate that integrating a conversational agent with motivational cues tends to increase patient engagement. However, further studies are required to determine PANDORA’s impact on exercise engagement in PwPD. Full article
(This article belongs to the Special Issue Artificial Intelligence in Digital Health)
Show Figures

Figure 1

22 pages, 6160 KB  
Article
WaterGPT: Training a Large Language Model to Become a Hydrology Expert
by Yi Ren, Tianyi Zhang, Xurong Dong, Weibin Li, Zhiyang Wang, Jie He, Hanzhi Zhang and Licheng Jiao
Water 2024, 16(21), 3075; https://doi.org/10.3390/w16213075 - 27 Oct 2024
Cited by 41 | Viewed by 9990
Abstract
This paper introduces WaterGPT, a language model designed for complex multimodal tasks in hydrology. WaterGPT is applied in three main areas: (1) processing and analyzing data such as images and text in water resources, (2) supporting intelligent decision-making for hydrological tasks, and (3) [...] Read more.
This paper introduces WaterGPT, a language model designed for complex multimodal tasks in hydrology. WaterGPT is applied in three main areas: (1) processing and analyzing data such as images and text in water resources, (2) supporting intelligent decision-making for hydrological tasks, and (3) enabling interdisciplinary information integration and knowledge-based Q&A. The model has achieved promising results. One core aspect of WaterGPT involves the meticulous segmentation of training data for the supervised fine-tuning phase, sourced from real-world data and annotated with high quality using both manual methods and GPT-series model annotations. These data are carefully categorized into four types: knowledge-based, task-oriented, negative samples, and multi-turn dialogues. Additionally, another key component is the development of a multi-agent framework called Water_Agent, which enables WaterGPT to intelligently invoke various tools to solve complex tasks in the field of water resources. This framework handles multimodal data, including text and images, allowing for deep understanding and analysis of complex hydrological environments. Based on this framework, WaterGPT has achieved over a 90% success rate in tasks such as object detection and waterbody extraction. For the waterbody extraction task, using Dice and mIoU metrics, WaterGPT’s performance on high-resolution images from 2013 to 2022 has remained stable, with accuracy exceeding 90%. Moreover, we have constructed a high-quality water resources evaluation dataset, EvalWater, which covers 21 categories and approximately 10,000 questions. Using this dataset, WaterGPT achieved the highest accuracy to date in the field of water resources, reaching 83.09%, which is about 17.83 points higher than GPT-4. Full article
Show Figures

Figure 1

35 pages, 836 KB  
Article
Enhancing Task-Oriented Dialogue Systems through Synchronous Multi-Party Interaction and Multi-Group Virtual Simulation
by Ellie S. Paek, Talyn Fan, James D. Finch and Jinho D. Choi
Information 2024, 15(9), 580; https://doi.org/10.3390/info15090580 - 19 Sep 2024
Cited by 2 | Viewed by 5857
Abstract
This paper presents two innovative approaches: a synchronous multi-party dialogue system that engages in simultaneous interactions with multiple users, and multi-group simulations involving virtual user groups to evaluate the resilience of this system. Unlike most other chatbots that communicate with each user independently, [...] Read more.
This paper presents two innovative approaches: a synchronous multi-party dialogue system that engages in simultaneous interactions with multiple users, and multi-group simulations involving virtual user groups to evaluate the resilience of this system. Unlike most other chatbots that communicate with each user independently, our system facilitates information gathering from multiple users and executes 17 administrative tasks for group requests adeptly by leveraging a state machine-based framework for complete control over dialogue flow and a large language model (LLM) for robust context understanding. Assessing such a unique dialogue system poses challenges, as it requires many groups of users to interact with the system concurrently for an extended duration. To address this, we simulate various virtual groups using an LLM, each comprising 10–30 users who may belong to multiple groups, in order to evaluate the efficacy of our system; each user is assigned a persona and allowed to interact freely without scripts. As a result, our system shows average success rates of 87% for task completion and 89% for natural language understanding. Comparatively, our virtual simulation, which has an average success rate of 80%, is juxtaposed with a group of 15 human users, depicting similar task diversity and error trends. To our knowledge, it is the first work to show the LLM’s potential in both task execution and the simulation of a synchronous dialogue system to fully automate administrative tasks. Full article
(This article belongs to the Special Issue Feature Papers in Artificial Intelligence 2024)
Show Figures

Figure 1

12 pages, 1086 KB  
Article
Enhancing Task-Oriented Dialogue Modeling through Coreference-Enhanced Contrastive Pre-Training
by Yi Huang, Si Chen, Yaqin Chen, Junlan Feng and Chao Deng
Appl. Sci. 2024, 14(17), 7614; https://doi.org/10.3390/app14177614 - 28 Aug 2024
Cited by 1 | Viewed by 2889
Abstract
Pre-trained language models (PLMs) are proficient at understanding context in plain text but often struggle with the nuanced linguistics of task-oriented dialogues. The information exchanges in dialogues and the dynamic role-shifting of speakers contribute to complex coreference and interlinking phenomena across multi-turn interactions. [...] Read more.
Pre-trained language models (PLMs) are proficient at understanding context in plain text but often struggle with the nuanced linguistics of task-oriented dialogues. The information exchanges in dialogues and the dynamic role-shifting of speakers contribute to complex coreference and interlinking phenomena across multi-turn interactions. To address these challenges, we propose Coreference-Enhanced Contrastive Pre-training (CECPT), an innovative pre-training framework specifically designed to enhance dialogue modeling. CECPT utilizes unsupervised dialogue datasets to capture both semantic richness and structural coherence. Our experimental results demonstrate that the CECPT model significantly outperforms established baselines in three critical applications: intent recognition, dialogue act prediction, and dialogue state tracking. These findings suggest that CECPT is more adept at following the information flow within dialogues and accurately linking statuses to their respective references. Full article
(This article belongs to the Section Computing and Artificial Intelligence)
Show Figures

Figure 1

16 pages, 410 KB  
Article
STOD: Towards Scalable Task-Oriented Dialogue System on MultiWOZ-API
by Hengtong Lu, Caixia Yuan and Xiaojie Wang
Appl. Sci. 2024, 14(12), 5303; https://doi.org/10.3390/app14125303 - 19 Jun 2024
Cited by 1 | Viewed by 4974
Abstract
Task-oriented dialogue systems (TODs) enable users to complete specific goals and are widely used in practice. Although existing models have achieved delightful performance for single-domain dialogues, scalability to new domains is far from well explored. Traditional dialogue systems rely on domain-specific information like [...] Read more.
Task-oriented dialogue systems (TODs) enable users to complete specific goals and are widely used in practice. Although existing models have achieved delightful performance for single-domain dialogues, scalability to new domains is far from well explored. Traditional dialogue systems rely on domain-specific information like dialogue state and database (DB), which limits the scalability of such systems. In this paper, we propose a Scalable Task-Oriented Dialogue modeling framework (STOD). Instead of labeling multiple dialogue components, which have been adopted by previous work, we only predict structured API queries to interact with DB and generate responses based on the complete DB results. Further, we construct a new API-schema-based TOD dataset MultiWOZ-API with API query and DB result annotation based on MultiWOZ 2.1. We then propose MSTOD and CSTOD for multi-domain and cross-domain TOD systems, respectively. We perform extensive qualitative experiments to verify the effectiveness of our proposed framework. We find the following. (1) Scalability across multiple domains: MSTOD achieves 2% improvements than the previous state-of-the-art in the multi-domain TOD. (2) Scalability to new domains: our framework enables satisfying generalization capability to new domains, a significant margin of 10% to existing baselines. Full article
(This article belongs to the Special Issue Natural Language Processing (NLP) and Applications—2nd Edition)
Show Figures

Figure 1

Back to TopTop