Next Article in Journal
Computational Jurisprudence: Verifiable Law for Machine Societies
Previous Article in Journal
Deep-Learning-Based Multi-Camera Framework for Indoor Human Detection and Presence Management
Previous Article in Special Issue
HOSPIT-LLM: A Human-Centered Multimodal Dataset and Edge-Deployed LLM Pipeline for Emotion-Aware Hospitality Assistants
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

The Telephone AI Paradox: How Voice Agents Can Help Counter Unwanted Telemarketing Through Role-Based Automation, Transparency, and Governance

1
Chair of Business Informatics, Processes and Systems, University of Potsdam, Karl-Marx-Straße 67, 14482 Potsdam, Germany
2
Capgemini, Bahnhofstraße 30, 90402 Nuremberg, Germany
3
Capgemini, Potsdamer Platz 5, 10785 Berlin, Germany
4
The James & Gail Ellis School of Business Leadership, The University of New Mexico, Albuquerque, NM 87131, USA
5
Department of Electrical Engineering and Information Technology, FH Aachen University of Applied Sciences, 52066 Aachen, Germany
6
iDA Institute for Digitalization, FH Aachen University of Applied Sciences, 52066 Aachen, Germany
*
Author to whom correspondence should be addressed.
Future Internet 2026, 18(8), 436; https://doi.org/10.3390/fi18080436
Submission received: 10 July 2026 / Revised: 6 August 2026 / Accepted: 10 August 2026 / Published: 14 August 2026
(This article belongs to the Special Issue Human-Centered Artificial Intelligence—2nd Edition)

Abstract

Unwanted telemarketing calls are a persistent source of consumer frustration and a legally regulated issue in Germany. At first glance, the idea of addressing this problem with AI-based voice technology appears contradictory: why should an automated caller help restore trust in a communication channel that has been damaged by aggressive outbound practices? This design-oriented case and prototype study argues that the paradox can be resolved through a different design logic. Rather than using AI to intensify persuasion, we present a role-based voice-agent architecture that constrains conversational behavior through narrow task boundaries, explicit escalation rules, and auditable data handling. The paper reports a transfer project involving FH Aachen students, Capgemini, and Fairdient GmbH. Methodologically, the work is positioned as a design-oriented case study with a prototype artifact. The contribution is threefold: first, we describe a three-agent architecture for outbound screening, consent-aware explanation, and inbound service; second, we derive governance principles for legally and ethically sensitive telephony, including transparency, bounded knowledge, privacy-preserving deployment, and human fallback; and third, we propose an evaluation framework covering conversion, compliance, hallucination control, user trust, and cost per validated outcome. The prototype does not yet claim large-scale field effectiveness. Instead, it offers a structured and empirically testable design for trustworthy voice automation in a domain where misuse, opacity, and user distrust are especially pronounced.

1. Introduction

Telephone communication remains one of the most intrusive channels in digital customer interaction because it competes directly for attention and interrupts private routines. This is particularly problematic in the context of unwanted marketing calls. In Germany, telemarketing towards consumers without prior explicit consent is unlawful, and violations can trigger substantial fines by the Bundesnetzagentur [1,2]. At the same time, victims often struggle to document, interpret, and pursue such cases in a legally meaningful way. This creates a practical gap between formal regulation and everyday enforceability.
The project discussed in this paper emerged exactly in this gap. Fairdient GmbH supports individuals affected by unlawful telemarketing and therefore faces an unusual operational challenge: in order to help people respond to problematic calls, the company must itself initiate telephone contact. This creates what we call the “telephone AI paradox”. The same medium that is widely associated with pressure, manipulation, and nuisance becomes the medium through which support, explanation, and procedural clarity must be offered.
The central research question is therefore not whether AI should replace human calling in a generic sense, but under which design conditions voice agents can be used in a way that reduces communicative harm rather than reproducing it. We argue that the answer lies in a human-centered and governance-driven architecture: narrow roles instead of general persuasion, bounded knowledge instead of unrestricted generation, and measurable controls instead of improvised scripts. In line with the manuscript structure recommended by Future Internet, this paper presents a design-oriented artifact, explains the project setting and method, and outlines an evaluation framework for subsequent pilot studies.

2. Theoretical Foundations and Related Work

This paper draws on five strands of prior work: design science research as the methodological frame; research on conversational agents in service encounters; research on anthropomorphism, trust, and disclosure in human–AI interaction; work on hallucination control and retrieval-constrained generation in large language models (LLMs); and the regulatory and abuse context of automated telephony. We review each strand and then position the contribution of this paper relative to the identified gap.

2.1. Design Science Research as Methodological Frame

Design science research (DSR) in information systems emphasizes the construction and evaluation of purposeful artifacts that address relevant organizational problems, and it provides explicit criteria for when such artifacts constitute research contributions rather than mere engineering [3]. Peffers et al. operationalize this stance as an iterative process of problem identification, objective definition, design and development, demonstration, evaluation, and communication [4]. Gregor and Hevner further distinguish contribution types by maturity, ranging from situated instantiations (level 1) over nascent design theories such as principles and architectures (level 2) to well-developed design theories (level 3) [5].
Within this framework, the present paper is positioned as a level 1/level 2 contribution: a situated prototype instantiation together with an explicit set of design principles (role decomposition, bounded knowledge, orchestration-enforced governance, fail-closed escalation) and an evaluation specification that renders the embedded design hypotheses empirically testable. This positioning matters because it defines the standard against which the paper should be judged: not measured field effectiveness, which is explicitly not claimed, but the rigor, novelty, and evaluability of the design knowledge.

2.2. Conversational Agents in Service Encounters

Conversational agents (CAs) have a long history in human–computer interaction, beginning with early demonstrations that even simple pattern-matching dialogue systems elicit strong social responses from users [6]. Contemporary research on CAs in customer service—where they are also called chatbots or assistants—has produced design knowledge on cooperative and socially adequate agent behavior [7], systematic taxonomies of the social cues through which agents shape user perception [8], and organizing reviews of the design space and its open questions [9]. A consistent finding across this literature is that interaction quality is not a function of linguistic capability alone; it depends on expectation management, task fit, and the handling of failure and handover situations.
Most of the existing CA research focuses on user-initiated or already-accepted service encounters—where the user is already aware of their service need, knows what organization needs to be contacted, and already has a willingness to interact with that organization. Outbound voice agent calls differ in an important way—they begin before the service encounter has been accepted by the user. The callee has not necessarily requested the interaction, may not recognize the caller, and has to decide whether the call is legitimate, useful, risky, or intrusive. Thus, CA calls can be understood as screened encounters, where the callee must decide whether to answer, continue, resist, or terminate the contact. This screening problem is amplified by the potential for abuse through caller-ID spoofing, robocall campaigns, voicemail spam, and scam calls—giving callees strong reasons to treat unknown automated calls as suspicious [10,11]. Synthetic or recorded human-like voices in CA can create an illusion of trust early in the conversation [12]. However, caleees decide to answer or ignore calls under conditions of incomplete information [13], and they frequently resist attempts to convert the initial contact into a sales encounter [14].
Thus, in addition to service quality, an important problem in outbound AI calls is interactional legitimacy at the point of contact.
The literature on AI-based conversational agents for outbound calls is still in its infancy, with many LLM-focused research still in pre-print stage. AI shopping assistants can increase sales and reduce returns by improving information provision [15]. AI agents seem to approach human performance on routing sales calls, but underperform on dimensions such as persuasion and objection handling [16]. And in some contexts, such as financial telemarketing, they may outperform human agents [17]. Performance improvements are also found when the agents are used in a hybrid mode, supporting human employees. Customer support assistants can significantly increase productivity for human support agents, with larger gains for less-experienced and lower-skilled workers [18]. Using LLMs as sounding boards seems to improve the work of non-experts, but LLM use also produces anchoring effects that could hurt experts [19]. And while AI outbound conversational agents look promising for structured, high-volume, information-heavy sales calls, their performance depends heavily on the voice technology stack, the task complexity, and the emotional context of the interaction [20,21,22,23,24].
The agent-oriented decomposition used in this paper also connects to the classical multi-agent systems literature, in which agents are defined by bounded competences, explicit interaction protocols, and organizational roles rather than by general intelligence [25]. Recent surveys of LLM-based multi-agent systems adopt a similar view: decomposing a task across specialized agents with restricted action spaces can improve controllability and observability compared to a single unconstrained model [26]. Our three-role architecture applies this principle to a legally sensitive telephony setting.

2.3. Anthropomorphism, Trust, and Disclosure

The “computers are social actors” paradigm demonstrates that users mindlessly apply social heuristics to machines, including politeness norms and personality attributions [27]. For voice interfaces this effect is amplified: prosody, turn-taking, and conversational fluency are strong anthropomorphic cues, which creates a specific risk of misplaced trust. The trust literature distinguishes trust as a willingness to be vulnerable based on perceived ability, benevolence, and integrity [28], and analyzes how trust in automation should be calibrated to actual system capability rather than maximized [29]. Empirical reviews of human trust in AI further show that trust antecedents differ between embodied, voice-based, and text-based systems, with voice occupying a middle ground of high anthropomorphism and low inspectability [30].
Business research on conversational AI in service encounters shows that customers respond not only to task performance but also to the agent’s perceived social presence, conversational skill, anthropomorphism, and capacity for repair [31,32,33]. Customers often penalize chatbots, even when their service is identical to human-provided service, partly because they associate chatbots with cost-cutting at the customer’s expense [34]. The design of the conversational agents also affects customer reactions: anthropomorphic chatbots can backfire with angry customers [22] but they can also strengthen consumer–brand relationships when mutual understanding matters [35]. In addition, sales effects vary by sales funnel stage [36].
Disclosure further complicates customer reactions: revealing that an agent is an AI bot can reduce trust and retention in high-criticality service settings, but it can also help customers interpret service failures more appropriately [37]. Directly relevant to our setting is the field-experimental finding by Luo et al. that undisclosed AI voice agents can match experienced human sales agents in outbound performance, whereas disclosing the agent’s machine identity before the conversation substantially reduces purchase rates [23]. This result frames the central tension of the present paper: disclosure carries a measurable short-term conversion cost, yet it is legally mandated and normatively constitutive for procedural predictability. Our design therefore treats disclosure not as an optimization variable but as a fixed constraint, and treats its interaction with trust and long-term acceptance as an explicit evaluation hypothesis. For measuring the callee-side perception of the resulting procedure, we draw on established constructs from organizational justice research [38].

2.4. Hallucination Control and Retrieval-Constrained Generation

LLM-based dialogue systems are subject to hallucination, i.e., the generation of fluent but unsupported or false content; surveys document that this risk is systemic rather than incidental and is particularly consequential in high-stakes domains [39]. Retrieval-augmented generation grounds model output in an external, controllable knowledge source [40], and empirical work shows that retrieval augmentation reduces hallucination in knowledge-grounded conversation [41]. Complementary to grounding, programmable guardrail frameworks constrain dialogue flow, topic scope, and tool access at the orchestration level rather than inside the model [42].
The architecture presented in this paper combines both mechanisms: a bounded, approved knowledge base as the only permitted answer source for the inbound role, and orchestration-level constraints (state machines, end-call tools, structured output schemas, escalation rules) that make deviations technically observable. The paper thereby follows the position that conversational safety in deployed systems is primarily an architectural property, with model-level quality as a necessary but not sufficient condition.

2.5. Regulatory and Abuse Context of Automated Telephony

Automated and unwanted telephony is a well-documented abuse ecosystem. Security research has systematized the techniques of telephone spam and the defensive countermeasures available to carriers and users [43], and large-scale measurement studies characterize robocall operations through audio and metadata analysis [11]. This literature is predominantly defensive: it asks how unwanted calls can be detected and blocked, not how legitimate automated calling could be designed to be trustworthy.
On the regulatory side, three layers apply to the German setting of this study. First, §7 of the German Act Against Unfair Competition (UWG) prohibits telemarketing towards consumers without prior express consent, enforced by the Bundesnetzagentur with substantial fines [1,44]. Second, the General Data Protection Regulation imposes lawfulness, purpose limitation, storage limitation, and data protection by design and by default on all processing of call-related personal data [45]; the latter principle continues the earlier privacy-by-design tradition [46]. Third, the EU Artificial Intelligence Act introduces a horizontal transparency obligation: providers must ensure that natural persons are informed when they interact with an AI system, unless this is obvious from the circumstances [47]. Human-centered AI research argues that such obligations should not be treated as external compliance burdens but as design inputs for reliable, safe, and trustworthy systems [48].

2.6. Research Gap and Positioning

Taken together, the reviewed strands reveal a gap. The CA literature provides rich design knowledge for text-based service chatbots but comparatively little for regulated outbound voice telephony, where the interaction is initiated by the system, the legal exposure is high, and user trust in the channel itself is damaged. The literature further suggests that LLM-based assistants need to be designed for a specific task and context—not as generic human replacements. They are most fitted for high-volume, structured, information-heavy interactions such as FAQs, order status, lead qualification, product explanation, appointment setting, and routine follow-up. In contrast, they are not well suited for angry, emotionally charged, ambiguous, high-stakes, or relationship-sensitive interactions. The trust and disclosure literature quantifies the cost of transparency but offers little architectural guidance on how to operate under mandatory disclosure. The hallucination literature provides model- and retrieval-level mechanisms but rarely connects them to workflow orchestration, consent handling, and auditability requirements. The telephony abuse literature, finally, treats automated calling almost exclusively as a threat to be blocked.
The contribution of this paper is located at this intersection: a design-science artifact that translates regulatory constraints (UWG, GDPR, AI Act) and trust-calibration insights into a role-based, retrieval-constrained, orchestration-governed voice-agent architecture, together with an operationalized evaluation framework that makes the resulting design hypotheses falsifiable in subsequent pilot studies.

3. Materials and Methods

This study uses a design-oriented case-study approach. The project constitutes a transfer setting between academia and practice: master’s students in business information systems at FH Aachen contributed conceptual design, prototype implementation, and evaluation logic; Capgemini contributed an industrial perspective on governance, long-term operability, and architectural robustness; and Fairdient GmbH contributed the real-world use case and operational constraints.
The resulting artifact is a prototype multi-agent system for legally sensitive telephone interaction. The aim was not to maximize conversational breadth, but to identify the minimum viable set of roles and controls required for trustworthy operation. The design process followed an iterative logic: problem framing, role decomposition, boundary definition, orchestration design, privacy review, and metric design. In this sense, the artifact is the central research result, while the proposed pilot protocol defines the basis for later empirical validation.
Methodologically, the paper should be read as an early-stage artifact paper rather than as a finished field experiment. No controlled field-study data and no quantitative benchmark measurements are claimed at this stage. Instead, the scientific value lies in the explicit articulation of design decisions, constraints, and evaluable hypotheses, meaning hypotheses that have not yet been empirically validated but have been specified in a way that enables subsequent measurement. This is important because many industrial AI prototypes remain anecdotal; they demonstrate what can be built, but not why the design should be considered trustworthy or transferable.

4. Prototype Architecture and Agent Role Design

The system consists of three specialized voice-agent roles coordinated by a central orchestration layer. The architecture deliberately rejects the idea of a single, highly flexible conversational agent. Instead, each agent is assigned a narrow function, a limited dialogue scope, and explicit stopping conditions. This role separation reduces behavioral ambiguity, supports policy enforcement, and improves auditability. Table 1 summarizes this role-based decomposition.
The first role is the Outbound Survey Agent. Its function is limited to respectful first contact and early qualification. The agent asks whether the called person would like support in relation to an earlier marketing call or a possible legal claim. If the person declines, the interaction ends immediately and the number is marked for removal or suppression in accordance with policy. The design intention is minimal intrusion: no persuasion loop, no repeated reframing, and no escalation into hidden sales behavior.
The second role is the Sales and Process Explanation Agent. This agent is only activated after explicit interest has been expressed. Its purpose is not aggressive selling but transparent explanation. It describes the service model, outlines the procedural steps, explains potential compensation pathways, and clarifies what information is required for further processing. Importantly, this role separates explanation from qualification. The conversation should remain intelligible, documented, and reversible. Consent, clarification, and data capture are treated as distinct interaction states rather than merged into one persuasive script.
The third role is the Inbound Service Agent. It answers incoming questions strictly on the basis of a limited knowledge base. In order to reduce hallucination risk, the agent is not allowed to improvise beyond that bounded source. If the available knowledge does not support an answer, the system explicitly communicates uncertainty and offers a handover path instead of fabricating legal or procedural advice. This is a crucial design decision: in sensitive service scenarios, reliability often improves when the model is allowed to know less.
The orchestration layer manages routing, state transitions, consent checks, logging, and escalation. It also enforces hard constraints such as maximum call duration, prohibited topics, and transfer to a human operator. In architectural terms, the orchestration layer is as important as the language model itself. Without orchestration, even a technically strong model can drift into non-compliant behavior; with orchestration, the system becomes governable.

5. Technical Architecture and Implementation Design

The prototype was implemented as a layered voice-agent system rather than as a single monolithic chatbot. The architecture separates telephony transport, real-time voice processing, conversational intelligence, workflow automation, data persistence, and governance controls. This separation is central to the design claim of this paper: trustworthy telephone automation does not emerge from the language model alone, but from the boundaries and handover points between technical components. Figure 1 depicts this layered architecture.
At the communication layer, customer calls are routed through a telephony provider and connected to the voice-agent platform by Session Initiation Protocol (SIP) trunking. The project documentation distinguishes this production-oriented configuration from a simpler API-based import path. The API-based path was associated with choppy or robotic audio and higher latency because media handling depends on repeated request-response cycles. The recommended SIP trunk creates a persistent voice session and transmits audio as a real-time media stream, which reduces jitter and improves audio fidelity for conversational turn-taking.
Vapi acts as the real-time voice orchestration layer. It coordinates the call session, the assigned assistant, the speech-to-text service, the language model, the text-to-speech service, function tools, and structured output extraction. In the implemented configuration, Deepgram provides German speech recognition, a Gemini Flash model provides the low-latency reasoning layer, and ElevenLabs provides German text-to-speech. This STT-LLM-TTS pipeline creates the observable voice interaction, while Vapi enforces assistant assignment, call initiation, end-call tools, handoff tools, and JSON-based output generation.
The workflow automation layer is implemented with n8n. It starts outbound campaigns, retrieves lead records, filters eligible contacts, triggers Vapi calls through authenticated HTTP requests, waits for call-completion callbacks, and writes post-call results back to the operational data store. The workflow uses Google Sheets as a lightweight CRM-like state store in the prototype, with fields such as phone number, name, call status, number of call attempts, availability time, survey participation, sales opt-in, rejection reason, and extracted contract data. In a production deployment, the same logic can be transferred to a dedicated CRM or case-management backend.
The three agent roles are therefore not only conversational personas but also different workflow states. The Outbound Survey Agent qualifies the contact, records survey answers, and schedules a follow-up only after explicit interest. The Outbound Sales Agent calls only prequalified contacts at the previously indicated time slot and collects contract-relevant data through structured output. The Inbound FAQ Agent answers incoming questions from a bounded knowledge base and transfers to sales only when the caller explicitly expresses purchase or subscription intent. This mapping between role, routing condition, and data schema supports auditability because every transition is technically observable.
Structured outputs are a key implementation mechanism. Instead of relying only on transcripts, the agents extract predefined fields into machine-readable JSON objects. For the survey flow, these include summary, frequency, annoyance score, participation, sales opt-in, availability time, call outcome, and rejection reason. For the sales flow, the structured output includes personal details, address, birthdate, occupation, nationality, email, payment method, IBAN where applicable, legal consent confirmations, contract start, and follow-up timing. These schemas reduce manual interpretation effort and make downstream processing testable.
The hosting strategy follows a hybrid pattern. Vapi remains a cloud-first component because it provides managed real-time voice infrastructure and model-provider integration. n8n and data storage can be hosted on German infrastructure, for example on a dedicated server, in order to retain stronger control over workflow execution, lead data, call metadata, and contractual records. This hybrid architecture balances latency, managed telephony capabilities, operational cost, and data protection control. It also creates a migration path: sensitive persistence and orchestration can remain under the operator’s control, while specialized real-time voice services are consumed where self-hosting would add disproportionate complexity.
Several technical safeguards translate governance requirements into executable system behavior. First, assistant assignment is explicit for inbound and outbound calls, preventing unbound numbers from initiating uncontrolled interactions. Second, end-call tools terminate calls when a wrong number, refusal, voicemail, or completed interaction is detected. Third, low temperature settings and short maximum response lengths constrain the language model and reduce long, improvised monologues. Fourth, knowledge-base tools restrict FAQ answers and objection handling to approved content. Fifth, webhook callbacks create an auditable post-call event, including call result and structured extraction. Sixth, call-attempt filters and status fields prevent repeated contact beyond the configured campaign rules. Table 2 lists the main technical components and their architectural responsibilities.

6. Empirical Case Setting and Operational Requirements

This section specifies the empirical operating frame of the Fairdient case and translates the business objective into measurable system requirements. The section does not report a completed large-scale field experiment. Instead, it documents the observed operational baseline, the target operating model, and the measurable pilot criteria that a production-ready voice-agent system would need to satisfy.
The current acquisition process is based on outsourced human calling. According to the project documentation, the status quo consists of three human call-center agents, each conducting approximately 60 calls per day (an average provided by Fairdient based on recent operations measurements). This results in an operational baseline of approximately 180 calls per day. The reported conversion rate, defined as the percentage of total calls resulting in customer contracts, is approximately 8–8.3%. In absolute terms, this corresponds to roughly 14–15 successful outcomes per day. For the target AI-supported operation, business objectives were jointly defined with Fairdient during the project to reflect Fairdient’s management expectations from the system. These business objectives therefore become explicit design targets for the voice-agent system, translating managerial expectations into operational requirements for scale, conversion performance, and contract completion.The target state is substantially more ambitious: Fairdient expects that the voice-agent system support at least 1000 calls per day and reach a conversion rate of at least 10%. At this target level, the expected number of successful outcomes (completed customer contracts) increases to approximately 100 per day. Table 3 contrasts the current manual baseline with the target operating model.
The target call volume also defines a capacity planning problem. If 1000 calls are distributed across an eight-hour calling window, the system must handle an average of 125 calls per hour, or approximately 2.1 initiated calls per minute. This average rate is not sufficient for infrastructure sizing because outbound campaigns rarely distribute perfectly evenly across the day. A conservative pilot design should therefore include peak-load assumptions. With a peak factor of 2, the system must support approximately 4.2 call starts per minute. With a peak factor of 3, it must support approximately 6.3 call starts per minute. If the average call lasts three minutes, this corresponds to approximately 6.3 concurrent calls under average load, 12.5 concurrent calls under a 2× peak, and 18.8 concurrent calls under a 3× peak. Table 4 summarizes these capacity-planning scenarios.
These numbers should be interpreted as system design requirements rather than measured production values. They clarify the operational scale that the architecture must be able to handle during a realistic pilot and later rollout. The relevant engineering question is therefore not only whether one call works, but whether the system remains stable when multiple calls run concurrently, when callbacks arrive asynchronously, and when post-call workflow steps write structured results back to the operational data store.
The empirical target frame also affects the latency requirement. Telephone interaction is highly sensitive to delay because even short pauses can make the agent appear robotic or inattentive. The project documentation defines a maximum response delay below 500 ms as an important target for natural turn-taking. In the implemented architecture, the end-to-end voice pipeline consists of speech recognition, language-model response generation, text-to-speech synthesis, telephony routing, and orchestration overhead. For this reason, latency should be measured at several levels: speech-to-text latency, time to first token of the language model, text-to-speech latency, and user-perceived end-to-end response time. A pilot should report at least median, 95th percentile, and maximum latency, because rare latency spikes can harm trust and conversion even when average latency appears acceptable.
The prototype already creates a suitable data basis for such measurement. The workflow automation layer retrieves lead records, filters eligible contacts, triggers calls, waits for completion callbacks, and writes post-call results back to the data store. The recorded fields include call status, number of call attempts, call summaries, availability windows, survey participation, sales opt-in, annoyance score, rejection reasons, and call outcomes. These fields make it possible to evaluate the system beyond raw conversion. In particular, they allow the pilot to distinguish between technical failure, non-contact, explicit rejection, qualified interest, scheduled follow-up, and completed conversion. Table 5 specifies the resulting pilot measurement framework.
A successful pilot should therefore not be defined solely by reaching 1000 calls per day. The more defensible success criterion combines scale, quality, and legitimacy. At minimum, the pilot should demonstrate that the system can operate at the required daily volume, maintain stable latency under peak load, avoid repeated or unwanted contact, preserve at least the current conversion baseline of approximately 8%, and trend toward the target conversion rate of 10% without increasing perceived pressure or procedural opacity. In this sense, the empirical requirement is not only “more calls”, but controlled, measurable, and auditable scaling.
For the present paper, the most conservative interpretation is therefore as follows: the Fairdient prototype is empirically grounded in a real operational baseline and a concrete target operating model, but its effectiveness remains a matter for future controlled validation. The contribution of the current artifact lies in making this validation possible. Because calls, callbacks, transcripts, structured outputs, routing decisions, refusal handling, and post-call status updates are technically observable, later field studies can test whether the architecture delivers economic scalability without sacrificing transparency, user autonomy, and compliance.

7. Operationalization of Bounded Legitimacy and Evaluation Readiness

This section addresses how the prototype operationalizes the claim that the agents should not improvise beyond their assigned roles and approved knowledge sources. The current implementation does not rely on fine-tuning, repeated self-critique by the model, or a large reasoning model as the primary safety mechanism. Instead, bounded behavior is achieved through a layered control design that combines role-specific prompting, deterministic dialogue structure, tool restrictions, structured outputs, workflow orchestration, and human fallback.
The first control layer is the role prompt. Each agent receives a narrow identity, task, and exit policy. The outbound survey agent is implemented as a deterministic “verbal form filler” with a linear-state-machine logic: identity check, permission, data collection, transition, and scheduling. The prompt also contains explicit technical directives such as immediate termination after voicemail detection and natural call termination through an end-call function rather than spoken meta-statements. Low temperature settings and short maximum responses further reduce creative deviation.
The second control layer is tool and knowledge-base access. The inbound FAQ agent answers only from an approved Fairdient knowledge base and must state that information is unavailable when the knowledge base does not support an answer. The sales agent uses a separate knowledge-base tool for approved Fairdient information and an objection-handling tool for predefined responses. This is closer to retrieval-constrained generation than to unrestricted conversation: the model may formulate the answer conversationally, but the permitted source and action space are deliberately narrow. The design follows the general logic of retrieval-augmented generation and hallucination mitigation research, where external knowledge access and source constraints can reduce unsupported generation [39,40].
The third control layer is orchestration. n8n controls call initiation, waiting for Vapi callbacks, status updates, and post-call routing. Vapi Structured Output schemas convert conversation results into machine-readable fields such as call_outcome, sales_opt_in, annoyance_score, availability_time, consent flags, payment details, and follow-up timing. These schemas do not guarantee truth by themselves, but they make the expected outputs explicit and testable. They also support transcript audits because every structured field can be compared with the corresponding spoken evidence.
The fourth control layer is fail-closed behavior. If a topic is outside the role, outside the knowledge base, legally sensitive, emotionally charged, or unresolved, the system should not continue as if it were competent. It should either communicate uncertainty, end the call, schedule a human follow-up, or transfer the caller to a human or specialized agent. In the present prototype, these mechanisms are implemented through prompt rules and Vapi tools, not through model fine-tuning. Future versions should add automated post-call transcript checks and adversarial test suites before any larger pilot.
For evaluation purposes, we distinguish between fluency and bounded legitimacy. Fluency describes whether the interaction works as a voice interaction: users can understand the agent, the agent responds with acceptable latency, interruptions are handled adequately, and the call does not feel technically broken. Bounded legitimacy describes whether the interaction stays within the approved role, knowledge, consent, and escalation boundaries. Prior work has largely treated AI voice agents as tools for conversion, productivity, routing, information provision, or employee augmentation. In contrast, we examine whether an outbound AI voice interaction can be structured as a bounded and procedurally legitimate encounter. In this setting, higher conversational fluency or higher conversion is not necessarily evidence of better system performance. A CA may be fluent but illegitimate if it obscures its identity, exceeds its approved role, makes unsupported claims, applies inappropriate pressure, or fails to escalate when required. Conversely, a CA may be procedurally legitimate but commercially ineffective if it complies with disclosure and escalation rules yet suffers from delays, misunderstands the callee’s responses, or fails to hold the callee’s attention. This distinction separates conversational competence from bounded legitimacy.
In this paper, the safer primary construct is procedural predictability: the callee can understand who is calling, why the call occurs, what the agent may do, how refusal is handled, and when a human can take over. Perceived procedural fairness can then be treated as a testable downstream hypothesis rather than an assumed property. If future work uses fairness language, it should explicitly measure whether callees perceive the procedure as consistent, unbiased, correctable, respectful, and sufficiently explained, drawing on established justice and trust measurement traditions [28,38]. The current artifact has therefore been implemented and architecturally specified, but it has not yet been validated through a controlled field study or benchmark. The contribution is not a claim of measured superiority. It is a design artifact plus an evaluation specification. The hypotheses are evaluable because the system produces transcripts, structured outputs, routing decisions, and workflow logs that can be checked against predefined metrics. Table 6 translates the most important constructs into operational definitions and measurement candidates.

8. Legal, Compliance, and Governance Requirements as Design Constraints

Legal, compliance and governance requirements for AI conversational agents are highly dependent on country, jurisdiction, sector, use case, data-processing context, and organizational policy. For that reason, we treat the applicable governance requirements as context-specific constraints that must be determined by the organization operating the agent, in consultation with appropriate legal and compliance expertise. The system architecture should then technically enforce them by constraining what the agent can say, what information it can use, when it must stop, when it must escalate, and how the interaction is logged or reviewed. In the case examined here, the system was implemented in Germany, and the design necessarily reflects Fairdient’s interpretation of the relevant legal, compliance and governance requirements in their context. In this paper, we do not seek to provide a comprehensive legal analysis of automated telephony in Germany or to generalize Fairdient’s specific compliance setup to other jurisdictions. Rather, we treat the applicable legal, compliance and governance requirements as design inputs that must be translated into the system architecture. The contribution of our project is to show how such requirements can be implemented through role specialization and governance-by-design. In many conversational AI projects, the dominant optimization target is fluency. In this project, fluency is secondary to bounded legitimacy. The architecture is therefore closer to a socio-technical control system than to a generic AI assistant.
Several system design hypotheses follow from this view. First, narrower roles should reduce the probability of conversational drift because the agent has fewer opportunities to invent new objectives. Second, bounded knowledge should reduce hallucination rates in comparison to unrestricted generation. Third, explicit disclosure that the caller is an AI system may initially lower acceptance in some interactions, but may improve procedural predictability and support future measurements of perceived procedural fairness and long-term trust. Fourth, self-hosting or sovereign hosting on German infrastructure may increase organizational feasibility in privacy-sensitive contexts by reducing data transfer risks and strengthening auditability. This leads to a broader conceptual point. The paradox of AI in telephony is not solved by claiming that machines are friendlier than humans. It is addressed by redesigning the channel itself. The quality of the interaction emerges from constraints: standardized scripts, transparent role boundaries, explicit off-ramps, limited knowledge access, and measurable compliance. In this sense, the project contributes to current debates on trustworthy AI by showing that conversational safety is not only a model question; it is an architectural question.
The project team identified privacy and compliance as primary rather than secondary requirements. This was especially important because many commercially attractive voice-agent stacks rely on U.S.-based large language model services and distributed data processing patterns that may be difficult to reconcile with strict organizational interpretations of data protection obligations. To address this issue, the prototype design explored a self-hosted deployment model on German servers. The main rationale was not technological nationalism but control: control over storage, over data flows, over retention, and over access to sensitive records. In addition, the architecture separates conversational logic, orchestration metadata, and knowledge-base access, which supports least-privilege principles.
Firm-specific governance requirements were translated into practical system rules. Examples include mandatory identification of the system as AI at the beginning of the call, prohibition of legal advice beyond the approved information base, retention limits for call-related data, mandatory escalation for unresolved or emotionally charged cases, and traceable reasons for every termination or transfer. Such governance rules are important because conversational quality alone cannot guarantee lawful or ethically acceptable behavior. A polite but unauthorized answer may still be harmful.

9. Discussion and Future Work

Because the current state of the project is a prototype, the most important next step is controlled empirical validation. We propose a multi-dimensional evaluation framework that combines operational, quality-and-safety, human-centered, and compliance indicators. Table 7 summarizes the proposed pilot metrics by evaluation dimension.
At the operational level, key metrics include contact rate, conversation completion rate, qualified-interest rate, transfer rate, average handling time, and cost per validated outcome. At the quality level, relevant measures include hallucination frequency, unsupported-answer rate, script deviation, and successful completion of mandatory disclosure steps. At the human-centered level, the pilot should capture perceived transparency, perceived pressure, perceived helpfulness, and willingness to continue the interaction. At the compliance level, the system should be evaluated for consent handling, deletion execution, retention adherence, and traceability of escalation decisions. A/B testing can be used cautiously to compare variants of disclosure wording, turn-taking style, escalation thresholds, or knowledge-base phrasing. However, experimental design in this domain must remain ethically restrained: aggressive behavioral nudges or ambiguity in system identity would undermine the very rationale of the project. The purpose of experimentation should therefore be to identify the most understandable and least manipulative version of the interaction, not the most exploitative one.
The project offers a useful counterpoint to the widespread assumption that more generative flexibility automatically leads to better conversational systems. In settings shaped by mistrust, legal sensitivity, and reputational risk, the opposite may be true. The most valuable system may be the one that is less fluent, less omniscient, and more strictly governed.
At the same time, the current work has limitations. First, the artifact has not yet been validated through a controlled real-world field study or benchmark against alternative voice-agent designs; validation so far is limited to prototype-level implementation and workflow plausibility. Second, the legal and organizational environment may vary across jurisdictions and sectors, which limits immediate generalization. Third, user reactions to AI callers are likely to depend on disclosure wording, voice design, prior experiences, and socio-demographic context. Fourth, self-hosting increases control but may also increase operational complexity and cost, especially with respect to model maintenance, observability, and security hardening. These limitations do not weaken the main contribution. Rather, they clarify its scope. This paper does not claim that voice agents have already solved unwanted telemarketing. It claims that there is a scientifically and operationally defensible way to prototype such systems without reproducing the harms of traditional outbound calling.

10. Conclusions

This paper reframes a hands-on German project as a design science contribution for the intelligent, human-centered web. Its key conclusion is that AI improves telephone communication only when it reduces imbalances of power rather than reinforcing them. Approaches such as role-based structuring, limited knowledge scopes, transparent disclosure, privacy-focused deployment, and verifiable governance turn voice automation from a tool of persuasion into a controlled and accountable service mechanism.
For future research, the priority lies in empirical testing under realistic pilot conditions. This involves examining how trust develops, how robust compliance remains, how effectively hallucinations are prevented, and whether the system proves economically sustainable over time. In practice, the takeaway is that organizations should not begin by asking whether a voice system sounds convincingly human, but whether the entire system is sufficiently controllable to be trustworthy.

Author Contributions

Conceptualization, C.C. and E.S.; methodology, E.S.; software, J.A., T.B., E.B., Y.H.F., S.I., E.R. and S.U.; validation, A.L. and C.C.; formal analysis, J.A., T.B., E.B., Y.H.F., S.I., E.R. and S.U.; investigation, E.S. and C.C.; resources, data curation, J.A., T.B., E.B., Y.H.F., S.I., E.R. and S.U.; writing—original draft preparation, E.S.; writing—review and editing, A.C., A.L. and C.C.; visualization, E.S.; supervision, C.C.; project administration, E.S.; funding acquisition, not applicable. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Data Availability Statement

No new data were created or analyzed in this study. Data sharing is not applicable to this article.

Acknowledgments

The authors gratefully acknowledge Alexander Vierlein from AV Solutions and Alp San from Fairdient for contributing the practical case context that informed this study. Their insights into market needs, operational requirements, and implementation-oriented design considerations provided an important foundation for translating the proposed architecture into a realistic application scenario.

Conflicts of Interest

Eldar Sultanow and Alexander Loosley are employed by Capgemini. Fairdient GmbH provided the real-world use case, operational baseline data, and implementation-oriented requirements for the prototype; AV Solutions contributed practical case context. These industry partners had no role in the analysis or interpretation of the reported operational data, in the writing of the manuscript, or in the decision to publish the results. Beyond the stated employment, the authors declare no conflict of interest.

References

  1. Bundesnetzagentur. Nachweis von Telefon-Werbeeinwilligungen [Documentation of Consent for Telephone Advertising]; Bundesnetzagentur: Bonn, Germany, 2022; Available online: https://www.bundesnetzagentur.de/DE/Fachthemen/Telekommunikation/Unternehmenspflichten/Telefonwerbung/start.html (accessed on 21 July 2026).
  2. European Commission. AI Act. Shaping Europe’s Digital Future, Directorate-General for Communications Networks, Content and Technology, Last Updated 11 May 2026. Available online: https://digital-strategy.ec.europa.eu/en/policies/regulatory-framework-ai (accessed on 21 July 2026).
  3. Hevner, A.R.; March, S.T.; Park, J.; Ram, S. Design science in information systems research. MIS Q. 2004, 28, 75–105. [Google Scholar] [CrossRef] [Scilit]
  4. Peffers, K.; Tuunanen, T.; Rothenberger, M.A.; Chatterjee, S. A Design Science Research Methodology for Information Systems Research. J. Manag. Inf. Syst. 2007, 24, 45–77. [Google Scholar] [CrossRef] [Scilit]
  5. Gregor, S.; Hevner, A.R. Positioning and Presenting Design Science Research for Maximum Impact. MIS Q. 2013, 37, 337–355. [Google Scholar] [CrossRef] [Scilit]
  6. Weizenbaum, J. ELIZA—A Computer Program for the Study of Natural Language Communication Between Man and Machine. Commun. ACM 1966, 9, 36–45. [Google Scholar] [CrossRef] [Scilit]
  7. Gnewuch, U.; Morana, S.; Maedche, A. Towards Designing Cooperative and Social Conversational Agents for Customer Service. In Proceedings of the 38th International Conference on Information Systems (ICIS), Seoul, Republic of Korea, 10–13 December 2017. [Google Scholar]
  8. Feine, J.; Gnewuch, U.; Morana, S.; Maedche, A. A Taxonomy of Social Cues for Conversational Agents. Int. J. Hum.-Comput. Stud. 2019, 132, 138–161. [Google Scholar] [CrossRef] [Scilit]
  9. Diederich, S.; Brendel, A.B.; Morana, S.; Kolbe, L.M. On the Design of and Interaction with Conversational Agents: An Organizing and Assessing Review of Human–Computer Interaction Research. J. Assoc. Inf. Syst. 2022, 23, 96–138. [Google Scholar] [CrossRef] [Scilit]
  10. Gupta, P.; Srinivasan, B.; Balasubramaniyan, V.; Ahamad, M. Phoneypot: Data-driven understanding of telephony threats. Proc. NDSS 2015, 107, 108. [Google Scholar]
  11. Prasad, S.; Bouma-Sims, E.; Mylappan, A.K.; Reaves, B. Who’s Calling? Characterizing Robocalls through Audio and Metadata Analysis. In Proceedings of the 29th USENIX Security Symposium; USENIX Association: Berkeley, CA, USA, 2020; pp. 397–414. [Google Scholar]
  12. Márquez Reiter, R.; Iveson, M. The establishment and breakdown of trust in human-bot marketing calls. Discourse Commun. 2025, 19, 522–545. [Google Scholar] [CrossRef] [Scilit]
  13. Grandhi, S.A.; Schuler, R.P.; Jones, Q. To answer or not to answer: That is the question for cell phone users. In Proceedings of the CHI’09 Extended Abstracts on Human Factors in Computing Systems; Association for Computing Machinery (ACM): New York, NY, USA, 2009; pp. 4621–4626. [Google Scholar]
  14. Humă, B.; Stokoe, E. Resistance in business-to-business “cold” sales calls. J. Lang. Soc. Psychol. 2023, 42, 630–652. [Google Scholar] [CrossRef] [Scilit]
  15. Wang, L.; Huang, N.; He, Y.; Liu, D.; Guo, X.; Sun, Y.; Chen, G. Artificial intelligence (AI) assistant in online shopping: A randomized field experiment on a livestream selling platform. Inf. Syst. Res. 2025, 36, 2358–2374. [Google Scholar] [CrossRef] [Scilit]
  16. Kaewtawee, K.; Modecrua, W.; Pachtrachai, K.; Kraisingkorn, T. Cloning a Conversational Voice 624 AI Agent from Call, Recording Datasets for Telesales. arXiv 2025, arXiv:2509.04871. [Google Scholar]
  17. Kan, Y.R.; Zhang, M.; Qiu, W.; Chen, F.; Tan, Y. Agentic Agent, Better Agent: Evidence from Agentic AI in Telemarketing. SSRN 2026. Available online: https://ssrn.com/abstract=6502379 (accessed on 31 March 2026).
  18. Brynjolfsson, E.; Li, D.; Raymond, L. Generative AI at work. Q. J. Econ. 2025, 140, 889–942. [Google Scholar] [CrossRef] [Scilit]
  19. Chen, Z.; Chan, J. Large language model in creative work: The role of collaboration modality and user expertise. Manag. Sci. 2024, 70, 9101–9117. [Google Scholar] [CrossRef] [Scilit]
  20. Wang, L.; Huang, N.; Hong, Y.; Liu, L.; Guo, X.; Chen, G. Voice-based AI in call center customer service: A natural field experiment. Prod. Oper. Manag. 2023, 32, 1002–1018. [Google Scholar] [CrossRef] [Scilit]
  21. Huang, M.H.; Rust, R.T. The caring machine: Feeling AI for customer care. J. Mark. 2024, 88, 1–23. [Google Scholar] [CrossRef] [Scilit]
  22. Crolic, C.; Thomaz, F.; Hadi, R.; Stephen, A.T. Blame the bot: Anthropomorphism and anger in customer–chatbot interactions. J. Mark. 2022, 86, 132–148. [Google Scholar] [CrossRef] [Scilit]
  23. Luo, X.; Tong, S.; Fang, Z.; Qu, Z. Frontiers: Machines vs. Humans: The Impact of Artificial Intelligence Chatbot Disclosure on Customer Purchases. Mark. Sci. 2019, 38, 937–947. [Google Scholar] [CrossRef] [Scilit]
  24. Jing, Z.; Xu, X.; Jin, Y.; Shen, J. Emotion vs. information: Understanding the effect of AI-powered call systems on potential customer decision from a field experiment. Decis. Support Syst. 2025, 201, 114579. [Google Scholar] [CrossRef] [Scilit]
  25. Wooldridge, M.; Jennings, N.R. Intelligent Agents: Theory and Practice. Knowl. Eng. Rev. 1995, 10, 115–152. [Google Scholar] [CrossRef] [Scilit]
  26. Guo, T.; Chen, X.; Wang, Y.; Chang, R.; Pei, S.; Chawla, N.V.; Wiest, O.; Zhang, X. Large Language Model Based Multi-Agents: A Survey of Progress and Challenges. In Proceedings of the 33rd International Joint Conference on Artificial Intelligence (IJCAI), Jeju, Republic of Korea, 3–9 August 2024. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  27. Nass, C.; Moon, Y. Machines and Mindlessness: Social Responses to Computers. J. Soc. Issues 2000, 56, 81–103. [Google Scholar] [CrossRef] [Scilit]
  28. Mayer, R.C.; Davis, J.H.; Schoorman, F.D. An Integrative Model of Organizational Trust. Acad. Manag. Rev. 1995, 20, 709–734. [Google Scholar] [CrossRef] [Scilit]
  29. Lee, J.D.; See, K.A. Trust in Automation: Designing for Appropriate Reliance. Hum. Factors 2004, 46, 50–80. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  30. Glikson, E.; Woolley, A.W. Human Trust in Artificial Intelligence: Review of Empirical Research. Acad. Manag. Ann. 2020, 14, 627–660. [Google Scholar] [CrossRef] [Scilit]
  31. Van Doorn, J.; Mende, M.; Noble, S.M.; Hulland, J.; Ostrom, A.L.; Grewal, D.; Petersen, J.A. Domo arigato Mr. Roboto: Emergence of automated social presence in organizational frontlines and customers’ service experiences. J. Serv. Res. 2017, 20, 43–58. [Google Scholar]
  32. Schuetzler, R.M.; Grimes, G.M.; Scott Giboney, J. The impact of chatbot conversational skill on engagement and perceived humanness. J. Manag. Inf. Syst. 2020, 37, 875–900. [Google Scholar] [CrossRef] [Scilit]
  33. Sheehan, B.; Jin, H.S.; Gottlieb, U. Customer service chatbots: Anthropomorphism and adoption. J. Bus. Res. 2020, 115, 14–24. [Google Scholar] [CrossRef] [Scilit]
  34. Castelo, N.; Bos, M.W.; Lehmann, D.R. Task-dependent algorithm aversion. J. Mark. Res. 2019, 56, 809–825. [Google Scholar] [CrossRef] [Scilit]
  35. Bergner, A.S.; Hildebrand, C.; Häubl, G. Machine talk: How verbal embodiment in conversational AI shapes consumer–brand relationships. J. Consum. Res. 2023, 50, 742–764. [Google Scholar] [CrossRef] [Scilit]
  36. Adam, M.; Wessel, M.; Benlian, A. AI-based chatbots in customer service and their effects on user compliance: M. Adam et al. Electron. Mark. 2021, 31, 427–445. [Google Scholar]
  37. Mozafari, N.; Weiger, W.H.; Hammerschmidt, M. Trust me, I’m a bot–repercussions of chatbot disclosure in different service frontline settings. J. Serv. Manag. 2022, 33, 221–245. [Google Scholar] [CrossRef] [Scilit]
  38. Colquitt, J.A. On the Dimensionality of Organizational Justice: A Construct Validation of a Measure. J. Appl. Psychol. 2001, 86, 386–400. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  39. Ji, Z.; Lee, N.; Frieske, R.; Yu, T.; Su, D.; Xu, Y.; Ishii, E.; Bang, Y.; Madotto, A.; Fung, P. Survey of Hallucination in Natural Language Generation. ACM Comput. Surv. 2023, 55, 248. [Google Scholar] [CrossRef] [Scilit]
  40. Lewis, P.; Perez, E.; Piktus, A.; Petroni, F.; Karpukhin, V.; Goyal, N.; Küttler, H.; Lewis, M.; Yih, W.; Rocktäschel, T.; et al. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. Proc. Adv. Neural Inf. Process. Syst. 2020, 33, 9459–9474. [Google Scholar]
  41. Shuster, K.; Poff, S.; Chen, M.; Kiela, D.; Weston, J. Retrieval Augmentation Reduces Hallucination in Conversation. Proc. Find. Assoc. Comput. Linguist. EMNLP 2021, 2021, 3784–3803. [Google Scholar] [CrossRef] [Scilit]
  42. Rebedea, T.; Dinu, R.; Sreedhar, M.N.; Parisien, C.; Cohen, J. NeMo Guardrails: A Toolkit for Controllable and Safe LLM Applications with Programmable Rails. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing: System Demonstrations; Association for Computational Linguistics (ACL): Stroudsburg, PA, USA, 2023; pp. 431–445. [Google Scholar] [CrossRef] [Scilit]
  43. Tu, H.; Doupé, A.; Zhao, Z.; Ahn, G.J. SoK: Everyone Hates Robocalls: A Survey of Techniques Against Telephone Spam. In Proceedings of the 2016 IEEE Symposium on Security and Privacy (SP); IEEE Computer Society (Conference Publishing Services/CPS): Los Alamitos, CA, USA, 2016; pp. 320–338. [Google Scholar] [CrossRef] [Scilit]
  44. Bundesrepublik Deutschland. Gesetz Gegen den Unlauteren Wettbewerb (UWG), §7: Unzumutbare Belästigungen. Bundesgesetzblatt. As Amended. 2004. Available online: https://www.gesetze-im-internet.de/uwg_2004/ (accessed on 9 August 2026).
  45. European Parliament and Council of the European Union. Regulation (EU) 2016/679 on the Protection of Natural Persons with Regard to the Processing of Personal Data (General Data Protection Regulation). Off. J. Eur. Union 2016, L 119, 1–88. Available online: http://data.europa.eu/eli/reg/2016/679/oj (accessed on 9 August 2026).
  46. Cavoukian, A. Privacy by Design: The 7 Foundational Principles; Information and Privacy Commissioner of Ontario: Toronto, ON, Canada, 2009. [Google Scholar]
  47. European Parliament and Council of the European Union. Regulation (EU) 2024/1689 Laying Down Harmonised Rules on Artificial Intelligence (Artificial Intelligence Act) See in particular Article 50 on transparency obligations. Off. J. Eur. Union 2024, L 1689, 1–144. Available online: http://data.europa.eu/eli/reg/2024/1689/oj (accessed on 9 August 2026).
  48. Shneiderman, B. Human-Centered Artificial Intelligence: Reliable, Safe & Trustworthy. Int. J. Hum.-Interact. 2020, 36, 495–504. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Layered technical architecture of the Fairdient voice-agent prototype.
Figure 1. Layered technical architecture of the Fairdient voice-agent prototype.
Futureinternet 18 00436 g001
Table 1. Role-based decomposition of the prototype system.
Table 1. Role-based decomposition of the prototype system.
Agent RolePrimary TaskHard BoundaryExpected Benefit
Outbound Survey AgentRespectful first contact and early qualificationEnds interaction immediately after refusal; no persuasion loop; suppression/removal workflowLower intrusiveness and clearer consent handling
Sales and Process Explanation AgentTransparent explanation of service, process, and required informationActivated only after explicit interest; no hidden transition from survey to salesImproved procedural clarity and auditable data capture
Inbound Service AgentAnswer incoming questions based on approved knowledge baseNo external improvisation; no legal advice outside approved content; explicit uncertainty allowedReduced hallucination risk and higher answer reliability
Table 2. Technical components and architectural responsibilities.
Table 2. Technical components and architectural responsibilities.
Layer/ComponentPrototype ImplementationArchitectural Responsibility
Telephony transportTelnyx SIP trunkingLow-latency call routing, caller-ID control, persistent media session
Voice orchestrationVapiCall session management, assistant routing, tools, handoff, structured outputs
Speech-to-textDeepgram Nova modelsGerman transcription and numerical recognition for phone conversations
Conversational modelGemini Flash/Flash LiteLow-latency dialogue generation under narrow role prompts
Text-to-speechElevenLabs Flash voiceNatural German voice output with stable pronunciation
Workflow and data layern8n plus Google Sheets prototype storeCampaign scheduling, lead filtering, callback handling, status updates, post-call persistence
Table 3. Operational baseline and target operating model.
Table 3. Operational baseline and target operating model.
MetricCurrent Manual Operation (Observed Operating Baseline)Target AI-Supported Operation (Business Objectives for AI System)
Daily call volumeApprox. 180 calls/dayAt least 1000 calls/day
Calling resourcesThree human call-center agentsVoice-agent system with automated orchestration
Calls per human agentApprox. 60 calls/dayNot directly applicable; capacity depends on concurrent call handling
Conversion rate (successful outcomes/total calls)Approx. 8–8.3%At least 10%
Successful outcomes per dayApprox. 14–15Approx. 100 at 1000 calls/day and 10% conversion
Cost structureHigh labor-dependent cost per callLower and more transparent cost per call/minute
ScalabilityLimited by human staffingNear-linear scaling subject to concurrency, provider limits, and governance controls
Expansion potentialOperationally limitedDACH first, later additional European markets
Table 4. Illustrative capacity model for the target AI-supported operating state (capacity planning scenarios).
Table 4. Illustrative capacity model for the target AI-supported operating state (capacity planning scenarios).
ScenarioCall Starts per MinuteAssumed Average Call DurationEstimated Concurrent Calls
Average load, 8-h window2.13 min6.3
Peak factor 24.23 min12.5
Peak factor 36.33 min18.8
Table 5. Pilot measurement framework for empirical validation.
Table 5. Pilot measurement framework for empirical validation.
DimensionMetricInterpretation
ReachCalls initiated per day; calls initiated per minute; peak call-start rateTests whether the campaign can reach the target volume
ContactabilityAnswer rate; voicemail rate; wrong-number rate; no-answer rateSeparates market/contact-list quality from agent performance
ConversionQualified-interest rate; sales-opt-in rate; completed-conversion rateMeasures whether automation preserves or improves business outcomes
Operational efficiencyAverage handling time; cost per call; cost per validated outcomeTests unit economics at the target volume
Latency and fluencyEnd-to-end response latency; interruption handling; silence duration; drop rateTests whether the interaction is technically usable as a phone call
GovernanceDisclosure completion; refusal recognition; maximum call attempts; escalation traceabilityTests whether the system stays within approved procedural boundaries
Quality and safetyUnsupported-answer rate; hallucination rate; script-deviation rate; complaint rateTests whether the system remains bounded and auditable
Table 6. Proposed constructs for evaluating bounded voice-agent behavior.
Table 6. Proposed constructs for evaluating bounded voice-agent behavior.
ConstructOperational DefinitionExample IndicatorsData Source
FluencyThe degree to which the voice interaction is technically and conversationally usable.End-to-end latency; interruption rate; ASR confidence; average silence duration; user-rated naturalness; completion without technical failure.Call logs; telephony metrics; ASR metadata; post-call questionnaire.
Bounded legitimacyThe degree to which the agent remains within approved role, knowledge, consent, and escalation boundaries.Unsupported-answer rate; prohibited-topic violations; role-deviation count; successful fallback; escalation correctness; no-persuasion-loop adherence.Transcript audit; structured outputs; policy checklist; human reviewer labels.
Hallucination controlThe degree to which factual or procedural claims are supported by the approved knowledge base or script.Share of claims supported by source; unsupported legal/procedural advice; uncertainty statement when source coverage is absent.Transcript-to-source comparison; knowledge-base coverage audit.
Procedural predictabilityThe degree to which callees can understand the identity, purpose, choices, refusal path, and handover options of the interaction.Disclosure completion; clarity of purpose; explicit refusal recognition; handover availability; no hidden transition from survey to sales.Transcript coding; user survey; callback complaint review.
Perceived procedural fairnessThe callee-side perception that the procedure is consistent, respectful, correctable, and adequately explained.Perceived respect; perceived pressure; perceived ability to refuse; perceived explanation quality; willingness to continue.Validated questionnaire items adapted from organizational justice and trust research.
Compliance robustnessThe degree to which operational behavior can be audited and defended against policy and legal requirements.Consent capture; deletion/suppression execution; retention adherence; traceable termination and transfer reasons.Workflow logs; CRM records; audit trail; retention reports.
Table 7. Proposed pilot metrics by evaluation dimension.
Table 7. Proposed pilot metrics by evaluation dimension.
DimensionIllustrative MetricsInterpretive Purpose
OperationalContact rate; completion rate; qualified-interest rate; average handling time; cost per validated outcomeTests whether the system is practically viable and efficient
Quality and safetyHallucination frequency; unsupported-answer rate; script deviation; successful disclosure completionTests whether conversation quality is bounded and policy-compliant
Human-centeredPerceived transparency; perceived pressure; perceived helpfulness; willingness to continueTests whether users experience the interaction as understandable and respectful
ComplianceConsent handling; deletion execution; retention adherence; escalation traceabilityTests whether operational behavior can be audited and defended
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Sultanow, E.; Loosley, A.; Chircu, A.; Arnold, J.; Bayer, T.; Bauer, E.; Firdaus, Y.H.; Ivanov, S.; Rofalski, E.; Ugur, S.; et al. The Telephone AI Paradox: How Voice Agents Can Help Counter Unwanted Telemarketing Through Role-Based Automation, Transparency, and Governance. Future Internet 2026, 18, 436. https://doi.org/10.3390/fi18080436

AMA Style

Sultanow E, Loosley A, Chircu A, Arnold J, Bayer T, Bauer E, Firdaus YH, Ivanov S, Rofalski E, Ugur S, et al. The Telephone AI Paradox: How Voice Agents Can Help Counter Unwanted Telemarketing Through Role-Based Automation, Transparency, and Governance. Future Internet. 2026; 18(8):436. https://doi.org/10.3390/fi18080436

Chicago/Turabian Style

Sultanow, Eldar, Alexander Loosley, Alina Chircu, Jonas Arnold, Timon Bayer, Emilia Bauer, Yudha Hefitra Firdaus, Stoyan Ivanov, Elisa Rofalski, Serhat Ugur, and et al. 2026. "The Telephone AI Paradox: How Voice Agents Can Help Counter Unwanted Telemarketing Through Role-Based Automation, Transparency, and Governance" Future Internet 18, no. 8: 436. https://doi.org/10.3390/fi18080436

APA Style

Sultanow, E., Loosley, A., Chircu, A., Arnold, J., Bayer, T., Bauer, E., Firdaus, Y. H., Ivanov, S., Rofalski, E., Ugur, S., & Czarnecki, C. (2026). The Telephone AI Paradox: How Voice Agents Can Help Counter Unwanted Telemarketing Through Role-Based Automation, Transparency, and Governance. Future Internet, 18(8), 436. https://doi.org/10.3390/fi18080436

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop