Identifying and Evaluating Artificial Intelligence Use Cases

The future of customer service is not about replacing people with AI—it is about redesigning service so humans and AI each do what they do best, creating better experiences, greater efficiency, and responsible value.

Sanchez P.

8/25/2026119 min read

Abstract

Artificial intelligence (AI) is rapidly transforming customer service, yet organisations face a fundamental challenge: identifying where AI can create sustainable value rather than simply where automation is technically possible. This paper develops an integrated framework for identifying, evaluating, prioritising, and implementing AI use cases in customer service. It combines the practical AI Use Case Canvas proposed by Deepsearch (2026) with evidence from research on AI-enabled service, human–AI complementarity, chatbot adoption, customer experience, employee productivity, and responsible AI.

The analysis advances the original six-dimensional framework by introducing an explicit governance and risk dimension. The resulting seven-dimensional framework evaluates customer-service AI use cases according to strategic relevance, use case and channel characteristics, data and system integration, human–AI roles and stakeholder implications, customer and operational outcomes, economic value, and governance and risk. The paper further develops a weighted prioritisation model, a four-quadrant portfolio, an iterative implementation roadmap, and an evidence-based approach to vendor evaluation.

The literature indicates that AI adoption should be analysed primarily at the task rather than job level. AI is particularly suited to structured, information-intensive, and transactional activities, whereas human capabilities remain important in ambiguous, relational, and emotionally sensitive interactions. Empirical evidence also demonstrates that AI productivity effects are heterogeneous: generative-AI assistance can substantially improve customer-support productivity, but benefits vary according to employee experience and task characteristics (Brynjolfsson et al., 2025). Customer research similarly suggests that automation creates value only when customers perceive genuine service benefits and when AI provides accurate, useful, and effective problem resolution.

The paper therefore argues that customer-service AI should be understood not primarily as an automation programme but as a service-system redesign problem. Sustainable value emerges when organisations deliberately allocate tasks between AI and humans according to comparative capabilities, customer expectations, economic value, and risk. The proposed framework provides a structured basis for making these decisions and for translating AI opportunities into measurable, governable, and scalable customer-service transformation.

Keywords: artificial intelligence; customer service; AI use cases; human–AI collaboration; service automation; customer experience; generative AI; AI governance; AI adoption; digital transformation

1. Introduction

Customer service is a particularly promising domain for the application of artificial intelligence (AI). Customer interactions generate large volumes of structured and semi-structured data; many requests recur across customers and channels; and a substantial proportion of service processes follow established rules, workflows, and knowledge structures. These characteristics create favourable conditions for AI-supported classification, information retrieval, response generation, decision support, and process automation (Huang & Rust, 2018; Deepsearch, 2026). At the same time, customer service extends beyond the efficient execution of transactions. It constitutes a critical interface between an organisation and its customers and can therefore shape customer satisfaction, trust, perceived service quality, and the broader customer relationship (Ashfaq et al., 2020; Zhao et al., 2022; Xie et al., 2024).

The strategic question for organisations is consequently not simply whether AI can be deployed in customer service, but which service activities should be automated, augmented, or deliberately retained as human-led interactions. This distinction is fundamental because technological feasibility does not automatically translate into customer value, organisational value, or sustainable economic benefit. Research on service AI suggests that the suitability of automation depends on the characteristics of the underlying task and interaction. Mechanical and analytical activities are generally more amenable to AI, whereas tasks requiring contextual judgement, intuition, empathy, warmth, or relationship management remain more dependent on human capabilities (Huang & Rust, 2018; Xiao & Kumar, 2021). Recent empirical research further indicates that the effects of AI adoption are heterogeneous: AI assistance can substantially improve employee productivity, but the magnitude of these gains varies with employee experience and task characteristics (Brynjolfsson et al., 2025).

This task-oriented perspective is particularly important because customer-service work comprises a heterogeneous set of activities. A single customer interaction may involve intent recognition, information retrieval, authentication, diagnosis, policy interpretation, response generation, transaction execution, documentation, and escalation. These activities differ considerably in their degree of structure, ambiguity, risk, and dependence on human judgement. Accordingly, AI use-case identification should focus on specific tasks and service processes rather than on entire occupations or functions (Huang & Rust, 2018). This approach also opens the possibility of human–AI complementarity, in which AI does not replace employees but instead augments their capabilities, reduces information-search costs, disseminates organisational knowledge, and enables employees to devote more time to complex or relational customer needs (Xiao & Kumar, 2021; Brynjolfsson et al., 2025).

The customer perspective reinforces the need for such differentiation. Research on AI-enabled customer service shows that customer satisfaction depends substantially on functional outcomes such as information quality, service quality, usefulness, and successful problem resolution (Ashfaq et al., 2020; Xie et al., 2024). Conversely, customers may respond negatively to AI-based service when systems are perceived as ineffective, slow, or incapable of resolving problems, particularly where human interaction is expected or where the service context involves emotional or relational needs (Zhao et al., 2022). Recent research comparing AI chatbots and human agents similarly suggests that AI can perform particularly well in competency-oriented, transactional interactions, whereas human agents retain important advantages in interactions characterised by warmth, trust, and relational engagement (Zhang et al., 2026). The implication is that successful AI adoption requires more than maximising automation rates: the allocation of tasks between AI and humans must reflect the nature and value of the customer interaction.

Against this background, the Deepsearch guide KI-Anwendungsfälle im Kundenservice identifizieren & bewerten provides a useful practice-oriented starting point for systematic AI use-case identification. Its AI Use Case Canvas structures the assessment around six dimensions: (1) the decision basis, (2) use cases and service channels, (3) data and system integration, (4) stakeholders, (5) success metrics, and (6) the business case (Deepsearch, 2026). The guide recommends beginning with concrete customer-service pain points and prioritising repetitive, relatively well-structured requests, particularly in high-volume channels such as email and telephone. It further emphasises the importance of historical interaction data, knowledge resources, system interfaces, stakeholder alignment, measurable performance targets, and an explicit economic assessment before selecting or scaling an AI solution (Deepsearch, 2026).

The practical logic of this framework is broadly consistent with the academic literature, but the literature also suggests several important extensions. First, use-case assessment should explicitly distinguish between automation and augmentation. Evidence from real-world customer support demonstrates that generative AI can increase employee productivity, with Brynjolfsson et al. (2025) reporting an average increase of approximately 15% in issues resolved per hour among 5,172 customer-support agents. Importantly, the effects were not uniform: less experienced and lower-performing employees benefited substantially more than experienced employees. These findings suggest that AI can function as an organisational knowledge and capability multiplier rather than merely as a mechanism for labour substitution (Brynjolfsson et al., 2025).

Second, customer value and service quality must be treated as explicit evaluation criteria. A technically successful AI system may nevertheless damage the customer experience if it provides inaccurate information, fails to resolve the customer's problem, or makes escalation to a human unnecessarily difficult. Empirical research demonstrates that information quality, service quality, usefulness, and problem-solving capability are central to customers' evaluations of AI-mediated service (Ashfaq et al., 2020; Xie et al., 2024; Zhao et al., 2022). The appropriate objective is therefore not maximum automation, but an optimal configuration of automated and human service that improves outcomes for both customers and the organisation.

Third, governance, transparency, and regulatory requirements need to be integrated into use-case selection from the outset. This is particularly important in regulated industries, where AI systems may process personal data or contribute to decisions with material consequences for customers. Data availability and system integration therefore need to be assessed alongside privacy, security, explainability, human oversight, and regulatory requirements rather than treated as purely technical implementation issues (Deepsearch, 2026). The European regulatory environment further strengthens this requirement, making governance an integral component of responsible AI deployment.

These considerations lead to the central research question of this paper:

How can organisations systematically identify and prioritise AI use cases in customer service so that technological feasibility, customer value, employee impact, economic value, and regulatory requirements are evaluated simultaneously?

The paper makes three contributions. First, it integrates the practice-oriented Deepsearch AI Use Case Canvas with insights from the academic literature on service AI, chatbots, human–AI collaboration, and generative AI. Second, it critically examines the limitations of automation-centric approaches, highlighting the importance of customer experience, task characteristics, employee augmentation, and human–AI complementarity. Third, it extends the existing framework into a more comprehensive AI Customer Service Use Case Canvas that incorporates human–AI role allocation and governance alongside strategic fit, data and technology readiness, stakeholder requirements, performance measurement, and economic value.

Taken together, the paper advances the proposition that AI use-case selection in customer service should be treated not primarily as a technology-selection exercise, but as a systematic service-design and organisational decision problem. The relevant question is therefore not which AI technology can automate the greatest proportion of customer interactions, but which configuration of AI and human capabilities creates the greatest sustainable customer and organisational value at an acceptable level of operational, employee, and regulatory risk (Huang & Rust, 2018; Brynjolfsson et al., 2025; Zhang et al., 2026).

2. Theoretical Foundations

2.1 AI Adoption as Task Transformation Rather Than Job Replacement

A central premise for evaluating AI in customer service is that adoption should be analysed primarily at the task level rather than at the level of entire occupations or organisational functions. Huang and Rust (2018) conceptualise artificial intelligence in service according to four forms of intelligence—mechanical, analytical, intuitive, and empathetic—and argue that these forms of intelligence differ in their susceptibility to technological substitution. Their framework suggests that AI is initially most capable of performing mechanical and increasingly analytical tasks, while intuitive and empathetic capabilities remain more strongly associated with human service provision. AI adoption should therefore be understood as a process of task substitution and augmentation rather than as a binary choice between human and machine labour (Huang & Rust, 2018).

This distinction is particularly relevant in customer service because customer-service work consists of multiple interdependent activities with substantially different levels of structure, ambiguity, and judgement. A single customer interaction may require the employee to identify the customer's intent, retrieve relevant information, authenticate the customer, interpret policies or tariffs, diagnose a problem, formulate a response, execute a transaction, document the interaction, recognise exceptional or emotionally sensitive circumstances, and determine whether escalation is necessary. Treating such an interaction as one indivisible "customer-service task" obscures the fact that these component activities differ considerably in their suitability for AI.

Activities such as intent classification, information retrieval, summarisation, knowledge retrieval, transcription, and response drafting are generally characterised by structured inputs, identifiable outputs, and the availability of historical examples or explicit rules. These characteristics make them relatively suitable for AI-based automation or decision support. Other activities, including negotiation, complex complaint resolution, discretionary decision-making, exceptional-case management, and interactions involving significant emotional or relational needs, are more dependent on contextual interpretation, judgement, empathy, or trust. In such cases, human involvement may not simply compensate for technological limitations; it may constitute an important source of service value (Huang & Rust, 2018; Xiao & Kumar, 2021).

The relevant strategic question is therefore not whether customer service as a whole can be automated, but rather:

Which tasks within a customer-service process can AI perform reliably and advantageously, and where does human involvement create additional customer and organisational value?

This task-oriented perspective provides an important theoretical foundation for the Deepsearch (2026) recommendation to prioritise repetitive and relatively well-structured customer requests. Such requests are attractive not merely because they are repetitive, but because their underlying tasks tend to have predictable inputs, explicit process logic, measurable outcomes, and relatively limited requirements for discretionary judgement. The value of an AI application therefore depends on the characteristics of the task being performed rather than simply on the availability of an AI technology.

At the same time, task suitability should not be interpreted as a fixed property. The boundary between tasks that can and cannot be automated is influenced by technological capabilities, data availability, process design, error tolerance, and the consequences of failure. Advances in generative AI, for example, have expanded the range of activities that can be supported through natural-language interaction and knowledge retrieval. However, increased technical capability does not eliminate the need to assess whether a particular application is appropriate from a customer, organisational, or regulatory perspective. The task-level perspective should therefore be understood as the starting point for use-case evaluation, not as a presumption that technically automatable tasks ought necessarily to be automated.

2.2 Service Automation as Human–AI Complementarity

The task perspective also challenges a purely substitution-oriented understanding of service automation. Rather than assuming that AI creates value primarily by replacing human labour, an alternative perspective conceptualises AI as a mechanism for human–AI complementarity and employee augmentation. Xiao and Kumar (2021) argue that service automation should be understood within a broader human–technology system in which technological capabilities interact with employee characteristics, customer expectations, service contexts, and organisational processes. From this perspective, the performance of an AI application cannot be evaluated independently of the human and organisational system in which it is deployed.

This perspective is strongly supported by recent empirical evidence. Brynjolfsson, Li, and Raymond (2025) examined the introduction of a generative-AI conversational assistant among 5,172 customer-support agents and found an average increase of approximately 15% in issues resolved per hour. Importantly, the productivity effects were heterogeneous. Less experienced and lower-performing employees experienced substantially greater improvements than more experienced employees. The authors interpret these effects partly as evidence that AI can disseminate knowledge and practices associated with higher-performing workers, thereby reducing differences in employee performance and accelerating the acquisition of effective service practices (Brynjolfsson et al., 2025).

This finding has important implications for the evaluation of AI business cases. If AI increases the productivity of existing employees, its value cannot be assessed solely through potential reductions in headcount. AI may instead generate value by reducing knowledge gaps, accelerating employee onboarding, improving access to organisational knowledge, supporting employees during complex interactions, and increasing the amount of customer demand that existing teams can handle. In this sense, the relevant economic outcome may be capacity creation rather than labour elimination (Brynjolfsson et al., 2025).

The distinction is particularly important because customer service involves considerable variation in employee experience. Newer employees may spend substantial time searching for information, interpreting policies, or determining how similar cases were previously resolved. An AI assistant capable of retrieving relevant knowledge, summarising customer histories, suggesting responses, or recommending next steps can reduce these information and experience gaps. The resulting productivity improvement can therefore emerge from a more effective distribution of cognitive work between employees and AI rather than from the removal of the employee from the process.

A theoretically and managerially useful framework should consequently distinguish between three broad modes of human–AI interaction: automation, augmentation, and decision support. In an automation configuration, AI performs a defined service process with little or no routine human intervention. Examples include responding to standard information requests, providing order or case-status information, or executing predefined appointment changes. Automation is most appropriate when the process is highly structured, the desired outcome is clearly defined, and the consequences of an incorrect action are limited or can be readily contained.

In an augmentation configuration, AI supports an employee who remains responsible for the customer interaction and the final outcome. The system may retrieve relevant knowledge, summarise previous interactions, classify the customer's intent, recommend information, or generate a draft response. The employee evaluates and, where necessary, modifies the AI output before communicating with the customer. This configuration is particularly valuable where AI can substantially reduce cognitive or administrative workload while human judgement remains important (Xiao & Kumar, 2021; Brynjolfsson et al., 2025).

The third configuration is AI-based decision support, in which AI analyses information and recommends an action while the final decision remains with an authorised human or organisational process. Examples include routing and prioritisation, identifying potentially relevant knowledge, recommending a next-best action, or highlighting patterns that may warrant further investigation. This model is particularly relevant where AI can process information at a scale or speed that exceeds human capability but where the consequences of an incorrect decision make unrestricted automation inappropriate.

These configurations should not be regarded as mutually exclusive or permanent. A single customer-service process may combine all three. For example, AI may automatically classify an incoming request, provide decision support to determine the appropriate workflow, and then augment a human employee during the resolution of a complex case. The appropriate configuration depends on factors including task complexity, data quality, customer expectations, error tolerance, regulatory requirements, and the consequences of incorrect decisions.

This perspective also provides a more nuanced interpretation of the concept of "automation rate". A high automation rate is not inherently evidence of a successful AI deployment. In some processes, the greatest value may arise from high-quality augmentation with limited autonomous execution, particularly where human judgement contributes substantially to customer trust or service quality. Conversely, routine and low-risk processes may generate greater value through end-to-end automation. Use-case evaluation must therefore determine not only whether AI can perform a task, but also which degree of AI autonomy is appropriate.

The distinction between automation and augmentation is further supported by research on customer perceptions of AI-enabled service. Ashfaq et al. (2020) show that information quality, service quality, usefulness, and related perceptions influence satisfaction and continued use of AI-powered service agents. Zhao et al. (2022) similarly identify inadequate problem-solving capability and the perceived absence of human interaction as important sources of dissatisfaction with AI customer service. These findings imply that the optimal human–AI configuration is partly determined by what customers need from a particular interaction. A highly transactional request may benefit from rapid autonomous resolution, whereas a complex or emotionally sensitive interaction may benefit from AI assistance combined with direct human involvement.

The theoretical implication is therefore that AI adoption in customer service should be conceptualised as the redesign of task allocation between humans and machines. The objective is not to maximise the proportion of work performed by AI, but to allocate individual tasks to the actor—human or AI—that can perform them most effectively while maintaining acceptable levels of quality, risk, and customer value. This provides the theoretical basis for evaluating AI use cases according to both their automation potential and their potential to augment human capabilities.

3. Customer-Service AI Use Cases

3.1 Why Customer Service Is Particularly Suitable for AI

Customer service is a particularly attractive domain for AI because it combines several characteristics that facilitate the application of data-driven and language-based technologies. Service organisations typically handle large volumes of customer interactions, many of which fall into recurring intent categories and follow established processes. These interactions also generate extensive historical data in the form of emails, chat transcripts, call recordings and transcripts, customer records, knowledge-base articles, product information, and case histories. At the same time, customer-service processes often have measurable operational outcomes, such as response time, average handling time, first-contact resolution, escalation, and customer satisfaction. Together, these characteristics create an environment in which AI systems can be trained, evaluated, integrated into workflows, and monitored against clearly defined performance objectives (Deepsearch, 2026; Huang & Rust, 2018).

The attractiveness of customer service is therefore not simply a consequence of high labour costs or the availability of increasingly capable AI models. More fundamentally, customer-service activities frequently combine high interaction volume, recurring patterns, structured information, and identifiable process outcomes. These characteristics make it possible to identify specific tasks that can be automated or augmented while retaining human involvement where contextual judgement or interpersonal capabilities are required. The Deepsearch (2026) framework accordingly recommends beginning with concrete service problems and prioritising repetitive, relatively well-structured requests, particularly in high-volume channels such as email and telephone.

This channel-oriented perspective is important because AI value depends not only on the type of task but also on where and how the customer interaction occurs. Email, for example, provides a naturally digital input that can be classified, summarised, routed, answered, or escalated using AI. Telephone service presents a different set of opportunities and challenges because speech recognition, intent detection, dialogue management, and real-time response generation must operate under less controlled conditions. Nevertheless, the underlying principle is the same: use cases should be prioritised where customer demand is substantial, the underlying process is sufficiently structured, and the potential benefits can be measured reliably (Deepsearch, 2026).

Importantly, high transaction volume alone does not make a use case a good candidate for AI. The task must also possess characteristics that allow AI to perform reliably. A process with millions of interactions but highly ambiguous inputs, constantly changing rules, limited training data, or severe consequences of error may be a poorer candidate than a smaller-volume process with stable information and well-defined outcomes. Use-case prioritisation should therefore consider the interaction between volume, task structure, data availability, process stability, expected value, and risk rather than treating volume as a sufficient selection criterion.

Evidence from generative-AI deployment in customer support further complicates the assumption that the most routine tasks necessarily provide the greatest marginal value. Brynjolfsson, Li, and Raymond (2025), in their study of 5,172 customer-support agents, found that generative-AI assistance increased issues resolved per hour by approximately 15% on average. However, the productivity effects varied substantially across types of work and employee experience. The greatest benefits occurred among less experienced and lower-performing employees, suggesting that AI can transfer knowledge and practices from higher-performing workers to employees who otherwise have less experience with particular customer problems. The researchers also found that the benefits were not simply proportional to the routine nature of the task: highly familiar problems offered less room for improvement because experienced employees were already able to resolve them efficiently, while extremely rare problems could provide insufficient training information for the AI system (Brynjolfsson et al., 2025).

This finding has an important implication for use-case selection. The relationship between task repetitiveness and AI value should not be assumed to be linear. A use case can be highly repetitive but generate limited incremental value if employees already resolve it quickly and accurately. Conversely, a structured but knowledge-intensive task may offer substantial potential because AI can provide employees with rapid access to information or recommended solutions that would otherwise require considerable search time or accumulated experience. The most promising use cases may therefore be those in which AI can combine reliable process execution with access to organisational knowledge.

This distinction also reinforces the importance of evaluating AI use cases in terms of comparative advantage rather than automation potential alone. A task is a strong candidate for AI when the technology can perform it more rapidly, consistently, accurately, or economically than the current process, or when AI can materially improve human performance. In contrast, a task should not be automated simply because an AI system is technically capable of performing it. The evaluation must consider whether automation improves the overall service system and whether the resulting benefits justify the technological, organisational, and regulatory risks involved.

3.2 Customer-Facing Chatbots

Customer-facing chatbots represent one of the most visible applications of AI in customer service and have consequently received substantial attention in both academic research and industry practice. Their appeal derives from the possibility of providing immediate responses at scale, extending service availability beyond traditional operating hours, and handling a proportion of customer requests without requiring direct employee intervention. However, research increasingly shows that the effectiveness of chatbots depends less on their ability to simulate human conversation than on their ability to deliver useful, accurate, timely, and contextually appropriate service.

Ashfaq et al. (2020), based on empirical data from 370 actual chatbot users, found that information quality and service quality positively influence customer satisfaction. Their findings also demonstrate the importance of perceived usefulness, ease of use, and enjoyment for continued chatbot use, while the customer's need for human interaction moderates the relationship between chatbot characteristics and user responses. These findings suggest that chatbot effectiveness is fundamentally connected to the quality of the service outcome rather than simply to the sophistication of the conversational interface (Ashfaq et al., 2020).

This conclusion is reinforced by the meta-analysis of Xie, Wang, and Cheng (2024), which synthesised evidence from 12 studies and found that utilitarian gratification was the strongest determinant among the examined dimensions of chatbot satisfaction. The finding is significant because it challenges the assumption that customers primarily value AI systems when they appear increasingly human-like. While conversational fluency can contribute to a positive interaction, customers ultimately have a functional objective: they want their problem resolved, their question answered, or the necessary transaction completed. Consequently, utility, information quality, and successful problem resolution should generally take precedence over anthropomorphic design (Xie et al., 2024).

This does not mean that the relational dimension of customer service is irrelevant. On the contrary, research demonstrates that customers can react negatively when AI systems fail to provide adequate problem-solving capability or when the interaction deprives them of access to human support. Zhao et al. (2022), analysing 17,673 Weibo posts together with 33 interviews, identified significant dissatisfaction with AI customer service, particularly where users perceived inadequate problem-solving ability, delayed or ineffective responses, and an absence of human interaction. Their findings highlight a critical distinction between automation that removes friction and automation that creates friction. A chatbot that provides an immediate and correct answer can substantially improve the customer experience; a chatbot that prevents access to a human while failing to resolve the problem can have the opposite effect.

The evidence therefore suggests that the key question is not whether customers prefer "AI" or "humans" in the abstract. Rather, customer preferences depend on the characteristics and objectives of the interaction. Zhang et al. (2026) provide further support for this contingency perspective, finding that AI chatbots can perform particularly well in competency-driven and transactional service interactions, while human agents retain important advantages in relational interactions where warmth and trust are central. This indicates that AI and human service should be viewed as complementary modes of service delivery, with their relative value depending on the type of customer need.

A useful distinction can consequently be drawn between transactional and relational customer-service interactions. Transactional interactions are primarily concerned with obtaining information, completing a straightforward process, or resolving a clearly defined issue. Examples include checking an order or claim status, requesting required documentation, changing a customer address, obtaining tariff information, or asking about the next step in a standard process. Such interactions generally have relatively clear objectives, predictable information requirements, and measurable outcomes. They are therefore often strong candidates for chatbot automation, provided that the AI has reliable access to the relevant information and systems.

Relational or high-stakes interactions have fundamentally different characteristics. A customer who has lost their income and cannot meet a loan obligation, for example, may require discretion and sensitivity rather than simply the retrieval of a standard procedure. Similarly, a customer whose insurance claim has been rejected may require an explanation, interpretation of the decision, and potentially a meaningful discussion of alternatives or escalation options. A long-standing customer who feels unfairly treated is likewise expressing a relationship problem that cannot necessarily be reduced to information retrieval. In such cases, the value of human involvement may derive from empathy, contextual judgement, trust, negotiation, and the ability to exercise discretion.

The distinction should nevertheless not be interpreted as a rigid division in which transactional interactions are always automated and relational interactions are always handled by humans. A sophisticated customer-service architecture can combine both forms of service. AI may initially identify the customer's intent, retrieve the relevant customer and product information, summarise the interaction history, and recommend an appropriate next step. If the case is straightforward and low-risk, the AI may complete the interaction autonomously. If uncertainty, emotional sensitivity, financial significance, or another predefined escalation criterion is detected, the interaction can be transferred to a human employee together with the information already assembled by the AI. Such a model preserves the efficiency advantages of automation while reducing the risk that customers become trapped in an inappropriate automated process (Xiao & Kumar, 2021; Brynjolfsson et al., 2025).

The implications for use-case selection are therefore twofold. First, organisations should prioritise customer-service applications according to the structure and purpose of the underlying interaction, rather than simply according to channel or technology. Second, they should design explicit escalation mechanisms for cases in which AI lacks sufficient confidence, information, authority, or contextual understanding. This is particularly important because customer acceptance of AI is influenced by the perceived need for human interaction and by the system's ability to resolve the customer's actual problem (Ashfaq et al., 2020; Zhao et al., 2022).

Overall, the literature supports a contingency-based approach to customer-service AI. AI is most attractive where it can deliver reliable functional value at scale; human involvement becomes increasingly important where interactions require contextual judgement, empathy, trust, or discretion. The strategic objective is therefore not to maximise chatbot penetration or automation rates, but to design an appropriate allocation of customer-service activities between AI and employees. This principle provides an important foundation for the subsequent evaluation of data requirements, integration, stakeholder implications, performance metrics, business value, and governance.

4. The AI Use Case Canvas

The AI Use Case Canvas proposed by Deepsearch (2026) provides a practical structure for moving from broad AI ambitions to concrete, assessable customer-service applications. Its central strength is that it frames AI adoption as a business and process-design question rather than as a technology-selection exercise. The original canvas addresses six dimensions: the strategic decision basis, the use case and service channel, data and system integration, stakeholders, success metrics, and the business case. Building on this structure and the preceding theoretical discussion, this paper extends the framework by treating governance, risk, and regulatory requirements as an explicit additional dimension.

The resulting framework is intended to support three stages of decision-making. First, it establishes whether a genuine business problem exists and whether AI is an appropriate means of addressing it. Second, it determines whether the required data, technology, organisational capabilities, and human roles are available. Third, it establishes whether the proposed application can generate measurable value while remaining acceptable from a customer, employee, operational, and regulatory perspective (Deepsearch, 2026; Huang & Rust, 2018).

4.1 Dimension 1: Strategic Decision Basis

The starting point for an AI use case should be the business problem rather than the AI technology. This principle is fundamental because technology-led initiatives can easily result in solutions searching for problems. The relevant question is not whether an organisation could deploy a chatbot, voicebot, generative-AI assistant, or another AI application, but whether a specific customer-service problem is sufficiently important to justify technological intervention.

A robust assessment should therefore establish the current state of the service process before any AI solution is selected. Relevant baseline measures include customer-service volumes, the distribution of request types, average handling time, response and resolution times, first-contact resolution, escalation rates, customer satisfaction, staffing requirements, demand fluctuations and peak periods, error rates, and the current cost of service delivery (Deepsearch, 2026). These measures provide the reference point against which the effects of an AI intervention can subsequently be evaluated.

The baseline is particularly important because an apparently attractive automation rate can have little economic significance if the underlying process is inexpensive, low-volume, or already highly efficient. Conversely, a relatively modest improvement in a high-volume process can create substantial value. The strategic assessment should therefore consider the magnitude of the underlying problem, the potential value of improvement, and the extent to which AI is capable of addressing its root causes.

The distinction between a business problem and a technology solution can be expressed simply:

Do not define the use case as "implement a chatbot." Define it as a measurable customer-service problem that a chatbot, AI assistant, or another technology may be capable of solving.

This distinction is also consistent with the task-based perspective developed in the preceding sections. Huang and Rust (2018) argue that AI adoption should be understood in terms of the characteristics of individual tasks and the relative capabilities of humans and machines. The strategic decision should therefore begin by identifying where performance is constrained and then determining whether AI can improve the relevant tasks.

For example, "customer emails take too long to answer" describes a symptom, while "60% of incoming emails concern five standard request categories and require agents to search multiple knowledge sources before responding" provides a basis for evaluating a specific AI intervention. The latter formulation identifies the volume, structure, process friction, and potential mechanism through which AI could create value.

This approach also establishes an important discipline for subsequent business-case development. The organisation should be able to articulate the expected change in operational terms—for example, a reduction in response time, handling time, escalation, or manual information retrieval—before estimating the financial value of the intervention. In this way, the strategic decision basis becomes the foundation for the entire use-case assessment rather than a retrospective justification for a technology purchase.

4.2 Dimension 2: Use Case, Channel, and Request Type

Once the underlying business problem has been established, the next step is to define the specific customer interaction and task that AI is expected to perform. A use case should be sufficiently precise that its process boundaries, inputs, outputs, human responsibilities, and escalation conditions can be understood and evaluated independently.

The Deepsearch (2026) framework emphasises the importance of specifying the relevant channel and request type. This can be extended into a more detailed description encompassing the customer segment, communication channel, request or intent, process trigger, expected AI action, required data, resulting output, human role, and conditions under which the interaction must be escalated. Such specification prevents broad initiatives such as "automate customer service" from being treated as single use cases when they actually contain many distinct processes with different technical, economic, and regulatory characteristics.

A useful formulation is:

When [trigger] occurs, the AI system [performs action] using [specified data], producing [defined output], unless [exception condition], in which case [defined human role] takes over.

This formulation has several advantages. It establishes the boundaries of the AI system, makes the expected behaviour explicit, identifies dependencies on information and other systems, and defines the circumstances under which autonomous operation should end. It also provides a basis for testing and measurement because the organisation can compare actual system performance with the expected process outcome.

Consider an incoming customer email requesting a standard address change. A narrowly defined use case could specify that the AI identifies the customer's intent, verifies whether the request contains the information required by the applicable procedure, retrieves the relevant process instructions, and either initiates an approved workflow through an integrated system or prepares a response for human approval. If the request contains contradictory information, fails an authentication requirement, or falls outside the standard procedure, the case is transferred to an employee.

This is substantially more evaluable than the general objective of "using AI to automate email." The former defines a bounded process, while the latter describes a technology ambition without specifying the problem, task, decision boundaries, or expected outcome.

The distinction is also important because the same channel may contain use cases with very different automation potential. An email channel, for example, can contain straightforward requests for information as well as complex complaints, disputes, financial difficulties, or exceptional cases. Similarly, a telephone channel can encompass simple status enquiries alongside emotionally sensitive or high-risk interactions. Channel selection is therefore insufficient on its own; the relevant unit of analysis is the interaction between channel, request type, task characteristics, and risk.

This perspective is consistent with research showing that customers evaluate AI-mediated service partly according to its ability to provide useful information and resolve their actual problems (Ashfaq et al., 2020; Xie et al., 2024). It also supports the distinction between transactional and relational interactions developed in Chapter 3. AI autonomy can be appropriate where the task is clearly defined and low risk, whereas interactions requiring judgement, empathy, trust, or discretion may require augmentation or human-led resolution (Zhao et al., 2022; Zhang et al., 2026).

Consequently, a mature use-case definition should specify not only what AI is expected to do, but also what AI is explicitly not authorised to do. Defining these boundaries is essential for reliable service design and provides the basis for subsequent human oversight and governance.

4.3 Dimension 3: Data and System Integration

A technically capable AI system can only perform a customer-service task effectively if it has access to the information and systems required to perform that task. Data readiness and system integration are therefore fundamental determinants of AI use-case feasibility (Deepsearch, 2026).

Potential information sources in customer service include historical emails, chat and call transcripts, frequently asked questions, knowledge bases, product and tariff information, policy documents, customer records, CRM systems, case-management and ticketing platforms, workflow engines, telephony infrastructure, and robotic process automation interfaces. Depending on the use case, the AI system may need access to one or several of these sources in order to interpret the request, retrieve relevant information, generate an appropriate response, or execute the associated process.

The presence of data, however, should not be confused with the suitability or legal permissibility of using that data. Data used in AI systems must be assessed in terms of quality, relevance, completeness, provenance, access rights, security, retention, and legal basis. In customer service, this is particularly important because interaction data frequently contains personal information and may include commercially sensitive or otherwise protected information. Consequently, data readiness is simultaneously a technical, organisational, and governance question.

Data quality is also critical to AI performance. Historical customer interactions may contain inconsistent terminology, outdated procedures, incomplete records, biased examples, or previous employee errors. Simply providing large volumes of historical data to an AI system does not guarantee reliable performance. The organisation must establish whether the available information is sufficiently accurate, current, representative, and structured for the intended task.

System integration presents a parallel challenge. An AI system may be capable of generating an accurate response but still create limited operational value if it cannot access the relevant customer record, case history, product information, or transaction system. Likewise, an AI system that identifies the correct action but cannot execute the corresponding workflow may simply transfer additional work back to employees. The practical value of AI therefore depends on its position within the broader service architecture rather than on model performance in isolation (Deepsearch, 2026).

For this reason, use-case assessment should distinguish three related forms of feasibility.

Data feasibility concerns whether the AI system can obtain the information necessary to perform the task and whether that information is sufficiently accurate, current, complete, and legally usable.

Technical feasibility concerns whether the AI application can connect securely and reliably to the systems required to support or execute the process, including CRM, ticketing, telephony, knowledge-management, workflow, and other relevant enterprise systems.

Operational feasibility concerns whether the resulting end-to-end process can operate reliably at the required volume and speed, including appropriate exception handling, human escalation, monitoring, maintenance, and support.

These dimensions should be assessed together because weakness in any one of them can undermine the overall business case. A use case may have excellent theoretical automation potential but limited practical value if the necessary data are inaccessible, the relevant APIs do not exist, or the surrounding workflow cannot support reliable autonomous execution.

The importance of this integrated perspective is reinforced by research on human–AI complementarity. Brynjolfsson et al. (2025) demonstrate that generative AI can substantially improve customer-support productivity, but the observed benefits arise within a specific organisational and technological environment in which the AI system can provide relevant information and guidance to employees. The result is therefore not simply a property of the underlying language model; it reflects the interaction between AI capability, organisational knowledge, employee behaviour, and the service process.

Accordingly, data and integration should be evaluated before selecting a vendor or committing to large-scale implementation. A high-quality use-case assessment should identify the required data sources, determine their availability and quality, map the systems involved, establish integration dependencies, and identify technical or governance gaps that could prevent the proposed AI application from delivering its intended value.

The three dimensions developed in this section—strategic decision basis, use-case definition, and data and system integration—therefore form the foundation of the AI Use Case Canvas. They move the organisation from a general ambition to "use AI in customer service" toward a specific, measurable and technically assessable proposition: a defined customer-service problem, expressed as a bounded task and supported by the data and systems required to perform it reliably.

5. Customer Experience and Service Quality

The evaluation of AI in customer service should not define success primarily in terms of automation rates or cost reduction. Customer service is simultaneously an operational process and a central component of the customer relationship. Consequently, an AI application can generate substantial efficiency gains while nevertheless destroying value if it reduces service quality, increases customer effort, undermines trust, or prevents customers from obtaining an appropriate resolution. Customer experience and service quality should therefore be treated as core outcome dimensions of AI adoption rather than secondary consequences of operational efficiency.

Research on AI-mediated service consistently indicates that customers evaluate these interactions according to both functional and relational criteria. Ashfaq et al. (2020), for example, demonstrate that information quality and service quality positively influence customer satisfaction with AI-powered service agents. Their findings further indicate that perceived usefulness and ease of use contribute to continued use, while the customer's need for human interaction moderates responses to chatbot-based service. These findings suggest that customers do not evaluate AI primarily according to whether it appears technologically sophisticated; they evaluate whether it enables them to accomplish what they came to customer service to accomplish.

This distinction is reinforced by the meta-analysis of Xie, Wang, and Cheng (2024), which identifies utilitarian gratification as a particularly important determinant of satisfaction with AI-powered chatbots. Customers generally approach service interactions with a practical objective: obtaining information, resolving a problem, completing a transaction, or understanding what they need to do next. Consequently, the fundamental measure of a successful AI interaction should be whether it delivers an accurate and useful outcome with an appropriate level of effort.

The implication is that functional service quality should precede conversational sophistication. An AI system that communicates fluently but provides inaccurate information, fails to understand the customer's request, or cannot complete the required process does not create a high-quality service experience. Indeed, conversational naturalness may increase frustration if it creates an expectation of competence that the system cannot fulfil. For customer-service applications, therefore, language quality should be considered an enabler of service delivery rather than an objective in itself.

The importance of effective problem resolution is also evident in research examining negative customer reactions to AI service. Zhao et al. (2022), analysing 17,673 social-media posts and 33 interviews, found substantial dissatisfaction with AI customer service where users perceived inadequate problem-solving ability, ineffective responses, and a lack of human interaction. These findings illustrate that automation can become counterproductive when it introduces additional effort into the customer's journey. A customer who is required to repeat information to a chatbot, navigate multiple unsuccessful automated responses, or overcome barriers before reaching a human employee may experience greater frustration than under the original human-led process.

This creates an important distinction between automation that removes customer effort and automation that transfers effort from the organisation to the customer. The latter may improve an organisation's apparent operational efficiency while simultaneously reducing customer value. AI use-case evaluations should therefore examine not only how much employee work can be automated, but also how the proposed process changes the customer's effort, waiting time, uncertainty, and probability of successful resolution.

The appropriate balance between AI and human service also depends on the nature of the interaction. Zhang et al. (2026) find that AI chatbots can perform particularly well in competency-driven, transactional interactions, whereas human agents retain advantages in relational interactions in which warmth and trust are important. This supports a differentiated service architecture rather than a universal automation strategy. AI can be highly effective where the customer primarily requires accurate information or efficient transaction processing, while human involvement becomes increasingly valuable where the interaction involves emotional sensitivity, ambiguity, negotiation, trust, or relationship management.

This distinction is consistent with the broader human–AI complementarity perspective developed by Huang and Rust (2018) and Xiao and Kumar (2021). Human involvement should not be understood merely as a fallback mechanism for technological failure. In some service contexts, human interaction is itself part of the value proposition. The ability of an employee to exercise judgement, demonstrate empathy, interpret an unusual situation, or reassure a customer can contribute directly to perceived service quality. The strategic objective is therefore to determine where human capabilities are complementary to AI rather than simply where AI capabilities are insufficient.

The design of the interaction also matters. Shin (2023) demonstrates that chatbot interaction characteristics, including the use of humour and other human-like behaviours, can influence consumers' evaluations of service. However, the findings should not be interpreted as evidence that increasing anthropomorphism necessarily improves customer experience. Human-like characteristics can influence how customers perceive and respond to a chatbot, but they cannot compensate for fundamental deficiencies in information quality or problem resolution. The strategic lesson is therefore that anthropomorphic design should support functional service quality rather than substitute for it.

This point is particularly relevant as generative AI enables increasingly natural conversational interfaces. The ability of a system to maintain context, produce fluent language, and respond in a conversational manner can make an interaction feel more sophisticated. However, natural language generation can also obscure uncertainty or create unwarranted perceptions of competence. For this reason, customer-service AI should be designed to communicate limitations appropriately and to provide clear pathways to human assistance when the system cannot reliably resolve the customer's request.

A high-quality AI customer-service experience should consequently be evaluated across several interconnected dimensions. First, the system should provide correct and relevant information based on authoritative organisational sources. Second, it should achieve successful problem resolution, rather than merely producing plausible conversational responses. Third, it should provide an appropriate response and resolution time, particularly for high-volume requests where speed is a major component of customer value. Fourth, the interaction should provide sufficient transparency, including appropriate disclosure of AI involvement where required and clear communication of relevant limitations. Fifth, customers should have easy access to escalation, particularly when the AI lacks sufficient confidence, information, authority, or contextual understanding. Sixth, the service should be consistent, avoiding arbitrary differences in outcomes across customers and channels. Finally, the system should incorporate appropriate human involvement where the interaction requires judgement, empathy, discretion, or relationship management.

These dimensions also suggest that conventional operational metrics such as automation rate and average handling time are insufficient as standalone measures of AI success. A system that achieves a 70% containment rate but causes a substantial decline in customer satisfaction or an increase in repeat contacts may have reduced apparent workload without improving the underlying service process. Conversely, an AI assistant that resolves only a modest proportion of interactions autonomously but substantially reduces employee handling time while improving first-contact resolution may create considerable value.

Customer experience should therefore be integrated directly into the AI use-case business case. At minimum, organisations should establish a baseline and monitor changes in customer satisfaction, customer effort, first-contact resolution, repeat contacts, escalation rates, abandonment, complaint volumes, and resolution time. These measures should be analysed alongside productivity and cost metrics so that management can determine whether AI is creating net service value rather than merely shifting costs or effort between the organisation, employees, and customers.

The central principle is consequently straightforward: the purpose of AI in customer service is not to make the interaction appear more automated or more human-like; it is to make the service more effective for the customer and the organisation. A chatbot that speaks naturally but cannot solve the customer's problem creates frustration rather than value. By contrast, an AI system that provides accurate information, resolves appropriate requests quickly, communicates transparently, and transfers complex cases smoothly to capable employees can improve both operational performance and customer experience (Ashfaq et al., 2020; Xie et al., 2024; Zhao et al., 2022; Zhang et al., 2026).

Customer experience should therefore be treated as a design constraint and success criterion from the beginning of AI use-case selection, not as a metric applied only after implementation. This principle reinforces the broader argument of this paper: successful customer-service AI depends on designing an appropriate division of work between technology and people, with the allocation determined by the characteristics of the customer need and the value created by each mode of service.

6. Stakeholder Alignment and Employee Impact

AI implementation in customer service should be understood as an organisational and work-system transformation rather than a software implementation alone. Introducing AI changes how customer interactions are processed, how decisions are made, how information is accessed, how employees perform their work, and how responsibility is distributed between humans and machines. Consequently, the success of an AI initiative depends not only on model performance and technical integration but also on stakeholder alignment, employee acceptance, process redesign, and the organisation's ability to establish appropriate forms of human oversight (Huang & Rust, 2018; Xiao & Kumar, 2021; Deepsearch, 2026).

This broader perspective has direct implications for the identification and evaluation of AI use cases. Relevant stakeholders typically include customer-service leadership, frontline employees, IT and enterprise architecture teams, data and analytics functions, information security, legal and compliance departments, data-protection officers, telecommunications teams where voice applications are involved, process owners, senior management, and, where applicable, works councils or other employee representatives. Customers should also be considered a central stakeholder because the consequences of AI deployment ultimately affect the way in which they access and experience the organisation's services (Deepsearch, 2026).

The involvement of these groups should occur before rather than after the technical solution has been selected. Different stakeholders evaluate the same AI use case from fundamentally different perspectives. Customer-service leaders may focus on service levels, productivity, and scalability; frontline employees may be concerned with workload, job design, autonomy, and the reliability of AI recommendations; IT teams must assess integration, architecture, security, and operational resilience; legal and compliance functions must consider applicable regulatory requirements and accountability; data-protection officers must assess the processing of personal data; and senior management must determine whether the proposed investment generates sufficient strategic and economic value. Early alignment is therefore necessary to prevent technically successful projects from encountering organisational resistance or governance barriers during implementation (Deepsearch, 2026; Xiao & Kumar, 2021).

6.1 Employees as a Central Dimension of AI Value

Employees deserve particular attention because AI changes not only the number of tasks performed by humans but also the composition and nature of human work. Huang and Rust (2018) argue that AI adoption progressively changes the relative importance of different forms of human intelligence. As mechanical and analytical activities become increasingly automatable, employees may devote a greater proportion of their time to tasks requiring intuition, judgement, empathy, creativity, and relationship management. AI adoption can therefore be understood as a transformation in the division of cognitive and interpersonal labour rather than simply as the elimination of human work.

This perspective is particularly important in customer service, where employees often perform both routine information-processing tasks and higher-value relational activities within the same interaction. If AI takes over information retrieval, summarisation, classification, documentation, or routine response generation, employees may have more time to focus on complex cases, complaints, negotiation, customer retention, and situations requiring discretion. Whether this transition actually occurs, however, depends on how the organisation redesigns the work system. If the organisation simply introduces AI while maintaining all existing responsibilities, employees may experience additional monitoring, exception handling, and quality-control burdens rather than genuine workload relief.

The empirical findings of Brynjolfsson, Li, and Raymond (2025) provide particularly strong evidence for considering AI as an employee-augmentation mechanism. Their study of 5,172 customer-support agents found that access to a generative-AI assistant increased issues resolved per hour by approximately 15% on average. Importantly, the gains were substantially larger among less experienced and lower-performing employees. This suggests that AI can help distribute organisational knowledge and effective problem-solving practices more broadly across the workforce, reducing the performance gap between less and more experienced employees (Brynjolfsson et al., 2025).

The finding has important implications for both workforce strategy and business-case evaluation. AI may create value by raising the capability of existing employees, rather than by reducing the number of employees required. A new employee supported by an AI system may become productive more quickly; an experienced employee may spend less time searching for information; and employees handling unfamiliar cases may receive relevant guidance without requiring immediate assistance from a supervisor or specialist. In each case, the economic value arises from increased organisational capability and capacity rather than direct labour elimination.

Accordingly, an AI business case should not be constructed exclusively around headcount reduction. A broader assessment should consider the potential effects on handling time, training requirements, onboarding speed, employee productivity, service capacity, employee turnover, work quality, and the proportion of employee time available for higher-value activities. It should also consider whether AI creates new responsibilities, such as monitoring AI outputs, managing exceptions, validating recommendations, maintaining knowledge sources, or supervising automated processes.

This broader conception of value is particularly relevant where labour markets are constrained or customer-service demand is growing. If AI enables employees to resolve more cases within existing staffing levels, the resulting capacity may have significant economic value even when no immediate reduction in employment occurs. Conversely, if AI reduces handling time but generates substantial additional verification or correction work, the apparent productivity gain may be overstated. The business case must therefore evaluate the net effect on the complete work process, rather than the performance of the AI component in isolation.

6.2 Employee Acceptance and Human–AI Collaboration

Employee acceptance is another critical determinant of AI implementation success. Xiao and Kumar (2021) emphasise that service technologies operate within a broader human–technology system in which employee and customer acceptance influence outcomes. An AI system that employees do not trust, understand, or perceive as useful is unlikely to generate its intended benefits, even if its technical performance is strong.

Trust is particularly important in customer service because employees may be held accountable for AI-generated outputs even when they have limited control over how those outputs are produced. If employees are expected to review every AI recommendation in detail, the system may create additional workload. If they are discouraged from questioning AI outputs, the organisation may increase the risk of inappropriate decisions or customer interactions. Effective implementation therefore requires a clear allocation of responsibility between AI and employees, including explicit rules concerning when employees must review, override, or escalate AI-generated outputs.

Training should consequently extend beyond teaching employees how to use an AI interface. Employees need to understand the system's capabilities and limitations, recognise potentially unreliable outputs, know when human judgement is required, and understand the appropriate escalation procedures. In this respect, AI literacy becomes part of operational risk management as well as employee development.

The role of employees may also evolve from direct execution toward supervision, exception management, and relationship-oriented service. This transformation is consistent with Huang and Rust's (2018) theoretical argument that the increasing automation of mechanical and analytical tasks shifts the relative importance of capabilities that are less readily substituted by AI. In customer service, these may include empathy, negotiation, contextual interpretation, complex judgement, and relationship management.

However, such a transition should not be assumed to occur automatically. Organisations need to redesign performance measures, training, job descriptions, escalation procedures, and career pathways accordingly. If employees continue to be evaluated primarily according to the number of routine interactions processed, the organisation may unintentionally discourage the higher-value activities that AI is intended to make possible.

6.3 Stakeholder Alignment as a Governance Mechanism

Stakeholder involvement should also be understood as a form of governance. Different stakeholders provide complementary perspectives on whether a use case is appropriate, feasible, and acceptable. Frontline employees can identify process exceptions that may not appear in formal process documentation. IT and security teams can identify architectural and cybersecurity constraints. Data-protection and compliance functions can identify legal risks. Customer-service leaders can assess operational consequences, while customers can reveal whether the proposed interaction actually improves the service experience.

The Deepsearch (2026) framework therefore appropriately emphasises early stakeholder involvement and recommends preparing responses to the principal concerns of different stakeholder groups before vendor discussions. This can also improve procurement quality because vendors can be evaluated against a clearly defined organisational problem and a set of requirements agreed across the relevant functions.

A useful stakeholder process should therefore establish, for each use case, who owns the business outcome, who owns the AI system, who is accountable for decisions, who monitors performance, who manages exceptions, and who has authority to suspend or modify the system when performance or risk becomes unacceptable. Such clarity becomes increasingly important as AI systems move from advisory roles toward autonomous execution.

6.4 Reframing Employee Impact in the Business Case

The combined theoretical and empirical evidence suggests that employee impact should be incorporated directly into the economic evaluation of customer-service AI. At least four forms of value should be distinguished.

First, AI can generate productivity value by reducing the time required to complete individual tasks. The evidence from Brynjolfsson et al. (2025) demonstrates that this effect can be substantial and may be particularly pronounced among less experienced workers.

Second, AI can generate capacity value by enabling existing employees to handle greater volumes of customer demand without a proportional increase in staffing. This is especially relevant where demand is growing, seasonal, or difficult to forecast.

Third, AI can generate capability value by accelerating employee learning, improving access to organisational knowledge, and reducing differences in performance between employees with different levels of experience (Brynjolfsson et al., 2025).

Fourth, AI can generate job-quality value by reducing repetitive administrative work and enabling employees to devote more time to complex, meaningful, or relationship-oriented activities. This potential is consistent with the broader task-transformation perspective proposed by Huang and Rust (2018).

These benefits must, however, be balanced against potential employee costs. AI can create new monitoring responsibilities, increase cognitive demands associated with reviewing automated outputs, reduce employee autonomy, or generate anxiety concerning job security. It may also intensify performance expectations if productivity improvements are converted directly into higher workloads rather than improved service capacity. Consequently, employee impact should be evaluated as both a potential source of value and a potential source of implementation risk.

The central implication is that organisations should not ask only whether AI can reduce the amount of human labour required. They should ask how AI will change the work performed by employees, whether those changes improve the overall service system, and whether employees have the skills, authority, and organisational support required to work effectively with AI.

AI implementation in customer service is therefore best understood as a redesign of the human–technology work system. The strongest business cases are likely to emerge where AI removes low-value cognitive and administrative work while strengthening employees' ability to resolve complex and relational customer needs. This perspective is consistent with the empirical evidence that AI can augment rather than simply substitute for human performance (Brynjolfsson et al., 2025), with the theoretical task-transformation framework of Huang and Rust (2018), and with the broader human–technology perspective developed by Xiao and Kumar (2021).

Ultimately, stakeholder alignment and employee impact should be treated as integral components of AI use-case selection rather than implementation considerations that arise after the business case has been approved. A technically feasible use case that employees cannot effectively adopt, customers do not value, or governance functions cannot support is not a viable use case. Conversely, an application that improves employee capability, customer outcomes, and organisational capacity can generate value well beyond the direct automation of individual tasks.

7. Success Metrics

The evaluation of customer-service AI requires a measurement framework that extends beyond model accuracy or automation rates. AI systems operate within a broader service system, and their value ultimately depends on whether they improve customer outcomes, operational performance, employee effectiveness, economic results, and organisational control. The Deepsearch (2026) guide proposes indicative 12-month targets such as intent-recognition accuracy above 84%, automation rates of approximately 50–60%, email response times below five minutes, and reductions in average handling time. These figures can provide useful orientation when developing an initial business case, but they should not be interpreted as universal scientific benchmarks. Their relevance depends on the characteristics of the process, the quality of the underlying data, the degree of process standardisation, and the organisation's starting position.

This distinction is particularly important when evaluating vendor proposals. A vendor-reported accuracy rate or automation percentage may have been obtained under conditions that differ substantially from those of the intended deployment. Performance can vary according to language, customer segment, request complexity, domain knowledge, data quality, escalation rules, and the consequences of incorrect responses. Organisations should therefore establish their own baseline, define the intended use case precisely, and require vendors to explain how claimed performance levels were measured. AI performance should ideally be assessed using representative organisational data and realistic operating conditions rather than relying exclusively on generic vendor benchmarks (Deepsearch, 2026).

A robust measurement framework should consequently combine customer, operational, AI-quality, employee, financial, and risk outcomes. This multidimensional approach reflects the central argument of this paper: an AI application should be evaluated according to the value it creates for the overall customer-service system rather than according to the performance of the AI component in isolation.

7.1 Customer Outcomes

Customer outcomes represent the most important external test of whether an AI application actually improves service. Core indicators include Customer Satisfaction (CSAT), Customer Effort Score (CES), Net Promoter Score (NPS), complaint rates, abandonment rates, and escalation rates.

CSAT can provide a direct measure of how customers evaluate an interaction, while CES captures the amount of effort customers perceive as necessary to achieve their objective. The latter is particularly relevant for AI because automation can reduce employee workload while inadvertently increasing customer effort. For example, a chatbot may reduce the number of contacts reaching human agents but require customers to navigate several unsuccessful dialogue turns before obtaining a resolution. In such a situation, containment may improve while the customer experience deteriorates.

This concern is supported by research showing that customers evaluate AI-mediated service according to information and service quality, usefulness, problem-solving capability, and the need for human interaction (Ashfaq et al., 2020; Zhao et al., 2022). Consequently, customer metrics should be analysed together with the actual resolution outcome. A high level of automated interaction is not necessarily desirable if it is accompanied by increased repeat contacts, complaints, or customer abandonment.

NPS can provide a broader indicator of relationship effects, although it should generally be interpreted alongside more immediate interaction-level measures such as CSAT and CES. Importantly, customer outcomes should be segmented by use case, channel, customer group, and AI-versus-human journey wherever possible. Aggregate measures can conceal significant differences between straightforward transactional requests and complex or relational interactions, for which the relative performance of AI and human service may differ substantially (Zhang et al., 2026).

7.2 Operational Outcomes

Operational measures assess whether AI improves the efficiency and effectiveness of the underlying service process. Important indicators include Average Handling Time (AHT), First Contact Resolution (FCR), response time, resolution time, automation or containment rate, and transfer rate.

AHT is useful for identifying whether AI reduces the time required to handle an interaction, particularly where AI assists employees with information retrieval, summarisation, documentation, or response generation. However, AHT should not be optimised independently of service quality. Reducing handling time while increasing repeat contacts or unresolved cases can simply shift work to a later stage of the customer journey.

FCR is therefore an important complementary measure because it captures whether the customer's issue is resolved without requiring additional contacts. Similarly, response and resolution times measure two different dimensions of service performance: how quickly the organisation acknowledges or responds to a request and how quickly the customer's underlying problem is actually resolved.

Automation or containment rate should likewise be interpreted carefully. A high containment rate indicates that interactions are being completed without human intervention, but it does not demonstrate that customers received an appropriate or successful resolution. Transfer rates provide useful information about how frequently AI hands cases to employees, but high transfer rates are not necessarily negative if escalation is deliberately designed into the process for complex, sensitive, or high-risk cases.

The operational measurement framework should therefore focus on end-to-end process performance. The objective is not to maximise automation but to determine whether the redesigned process resolves customer needs more effectively and efficiently. This is consistent with the human–AI complementarity perspective developed by Huang and Rust (2018), Xiao and Kumar (2021), and Brynjolfsson et al. (2025).

7.3 AI Quality and Reliability

Operational and customer outcomes must be supplemented by measures that directly assess AI quality. Relevant indicators include intent-recognition accuracy, answer accuracy, hallucination rate, inappropriate-response rate, escalation precision, and retrieval accuracy.

Intent-recognition accuracy is particularly relevant for routing and classification use cases because an incorrect classification can place the customer into an inappropriate workflow. Answer accuracy is central to knowledge-based applications, while retrieval accuracy measures whether the AI system accesses the correct organisational information before generating its response. Hallucination rates are particularly important for generative-AI applications because a fluent but factually incorrect answer can create significant customer, operational, and regulatory risks.

The evaluation of AI quality should, however, extend beyond aggregate accuracy. An average accuracy figure can conceal highly consequential failures. An AI system might achieve 95% accuracy overall while producing unreliable outputs for a small but high-risk category of interactions. In a regulated environment, the remaining 5% may contain cases involving financial consequences, legal rights, sensitive personal data, or other material risks.

The relevant principle is therefore risk-weighted AI evaluation. Performance should be measured separately for different request categories and should give particular attention to cases in which errors have disproportionate consequences. This also reinforces the importance of predefined escalation mechanisms. Where confidence is insufficient or the interaction falls outside the system's validated scope, transferring the case to a human may represent successful system behaviour rather than AI failure.

This distinction is particularly important when interpreting vendor benchmarks. A high model-level accuracy rate does not establish that the AI system is appropriate for a specific business process. The relevant question is whether the system achieves sufficient performance within the intended operational and risk context.

7.4 Employee Outcomes

Because AI changes the distribution of work between humans and machines, employee outcomes should form a further component of the measurement framework. Relevant indicators include employee productivity, training and onboarding time, employee satisfaction, workload, turnover, and the proportion of employee time devoted to higher-value activities.

The importance of these measures is supported by Brynjolfsson et al. (2025), who found substantial productivity improvements following the introduction of generative-AI assistance, with particularly strong benefits among less experienced and lower-performing customer-support employees. This suggests that AI can create value through capability enhancement and knowledge transfer rather than solely through the elimination of labour.

Productivity should therefore be assessed not only through the number of interactions processed but also through changes in the time required to become proficient, the amount of supervisory support required, and the proportion of work that employees can devote to complex or relational activities. If AI reduces routine administrative work and enables employees to spend more time resolving difficult cases, the resulting value may not be captured by conventional headcount metrics.

Employee measures are also important for identifying unintended consequences. AI may reduce some forms of workload while increasing the cognitive burden associated with monitoring, checking, or correcting AI-generated outputs. Employee satisfaction and workload should therefore be measured before and after implementation, with particular attention to whether the redesigned process actually improves the employee experience.

7.5 Financial Outcomes

Financial metrics translate operational and customer improvements into an economic assessment of the AI investment. Relevant measures include cost per contact, labour-hour savings, technology costs, implementation costs, avoided costs, incremental revenue, and return on investment.

Cost per contact can reveal whether AI reduces the economic cost of serving individual customer requests. However, this measure should incorporate the full cost of the redesigned process, including AI infrastructure, licences, integration, data preparation, monitoring, maintenance, human oversight, and exception handling. An AI solution that appears inexpensive at the interaction level may have a substantially different total cost once implementation and ongoing governance requirements are included.

Similarly, labour-hour savings should not automatically be interpreted as headcount savings. As discussed in Chapter 6, AI may instead create additional service capacity, reduce the need for future recruitment, accelerate onboarding, or enable employees to handle more complex cases. These benefits can have substantial economic value even when employment levels remain unchanged (Brynjolfsson et al., 2025).

The business case should therefore distinguish between cost reduction, cost avoidance, capacity creation, and revenue effects. This distinction prevents organisations from understating the value of augmentation-oriented use cases simply because they do not produce immediate reductions in staffing.

A credible financial evaluation should also incorporate implementation and transition costs. These may include system integration, data preparation, process redesign, employee training, vendor costs, security controls, compliance assessments, monitoring infrastructure, and ongoing model maintenance. The relevant measure is consequently the net economic contribution of the complete AI-enabled service process rather than the nominal cost of the AI technology.

7.6 Risk and Governance Outcomes

AI performance should finally be evaluated against explicit risk and governance measures. Relevant indicators include privacy incidents, compliance deviations, security incidents, model failures, undocumented or insufficiently traceable decisions, and inappropriate automation.

These measures are particularly important in regulated customer-service environments, where a technically successful system can nevertheless be unacceptable if it processes personal data inappropriately, produces decisions that cannot be adequately explained, or operates outside its authorised scope. Risk metrics should therefore be incorporated into the initial success criteria rather than treated solely as post-implementation compliance checks (Deepsearch, 2026).

The organisation should also monitor the frequency and nature of cases in which AI should have escalated but did not, as well as cases in which AI escalated unnecessarily. These two forms of error have different consequences. Under-escalation can expose customers and the organisation to inappropriate automated decisions, while over-escalation can reduce the efficiency benefits of the system. The objective is therefore not simply to minimise escalation but to achieve appropriate escalation based on the characteristics and risk of the interaction.

Governance metrics should also capture whether AI behaviour remains within its validated scope. Changes in products, tariffs, policies, customer behaviour, or underlying data can cause model performance to deteriorate over time. Continuous monitoring is therefore necessary rather than treating evaluation as a one-time activity conducted before deployment.

7.7 Measurement as a Multi-Dimensional Evaluation System

The six measurement categories should not be treated as independent scorecards. They are interconnected and should be evaluated as a system of outcomes. An improvement in one dimension may generate deterioration in another. For example, increasing automation may reduce labour costs while increasing customer effort. Reducing AHT may decrease employee workload while increasing repeat contacts. Increasing autonomous decision-making may improve efficiency while increasing regulatory risk.

This creates an important methodological principle:

AI performance must be evaluated jointly with business, customer, employee, and risk performance.

A 95% intent-recognition rate is therefore not inherently evidence of success. If the remaining 5% consists primarily of low-risk routine cases, the error rate may be manageable. If those errors disproportionately affect vulnerable customers, financially significant transactions, or regulated decisions, the same aggregate performance level may be unacceptable.

Similarly, an automation rate of 60% should not be interpreted as successful merely because it exceeds a predefined target. The relevant question is whether the 60% of automated interactions are resolved correctly, whether customers are satisfied, whether employees can effectively handle the remaining cases, and whether the resulting economic benefits exceed the costs and risks of implementation.

A rigorous evaluation should therefore establish baseline values before implementation, define target outcomes for each relevant dimension, and measure performance after deployment using comparable data. Where feasible, organisations should use controlled pilots, phased rollouts, or other forms of comparative evaluation to distinguish the effects of AI from unrelated changes in demand, staffing, processes, or customer behaviour. This is particularly important because observed improvements may otherwise be incorrectly attributed to AI.

The resulting measurement architecture can be conceptualised as a hierarchy. At the lowest level, organisations assess technical AI quality, such as accuracy and retrieval performance. These measures determine whether the system performs its immediate computational task. At the next level, organisations assess process outcomes, such as AHT, FCR, response time, and containment. These determine whether AI improves the service workflow. At the highest level, organisations assess customer, employee, financial, and risk outcomes, which determine whether the redesigned service system creates sustainable organisational value.

This hierarchy prevents a common evaluation error: assuming that a technically capable AI system necessarily produces business value. The evidence reviewed throughout this paper suggests that value emerges only when AI capabilities are effectively integrated into tasks, workflows, employee roles, and customer journeys (Huang & Rust, 2018; Xiao & Kumar, 2021; Brynjolfsson et al., 2025).

Accordingly, the Deepsearch (2026) targets should be used as starting hypotheses for management rather than fixed success criteria. Organisations should translate them into use-case-specific targets based on their own baseline, risk tolerance, customer expectations, and strategic objectives. The ultimate criterion is not whether an AI system reaches a particular generic accuracy or automation threshold, but whether it produces a demonstrably better overall outcome than the alternative—while remaining reliable, economically justified, and appropriately governed.

8. Business Case and Economic Value

The economic evaluation of customer-service AI should extend beyond the question of whether the technology can reduce labour costs. As the preceding chapters demonstrate, AI can create value through several mechanisms, including direct automation, employee productivity, capacity creation, faster service, improved quality, reduced errors, and the ability to scale service without proportional increases in staffing. At the same time, AI introduces new costs and risks that must be incorporated into the assessment. A credible business case should therefore evaluate the net economic value of the redesigned service process, rather than the apparent savings generated by the AI component alone.

A useful conceptualisation is that the economic value of an AI use case depends on the relationship between its benefits and its total costs. Benefits may arise from reduced handling time, lower cost per contact, increased service capacity, reduced training requirements, avoided recruitment, lower error rates, improved customer retention, or incremental revenue. Costs include not only AI technology but also implementation, integration, data preparation, employee training, process redesign, monitoring, maintenance, and governance. In addition, organisations should consider the expected cost of AI-related failures, including inappropriate responses, compliance violations, security incidents, and customer remediation.

The Deepsearch (2026) guide illustrates the potential scale of the opportunity through an example in which 40% of 3,000 monthly routine requests are automated, potentially freeing approximately 1,920 hours per year. The calculation is useful for demonstrating the logic of capacity creation, but it should be understood as an illustrative scenario rather than a general economic benchmark. The actual result depends on factors such as the number of requests that are genuinely suitable for automation, the average handling time, the proportion of the process that is eliminated rather than merely transferred, the rate of exceptions and escalations, and whether employees can actually redeploy the resulting capacity.

This distinction is essential because an automated interaction does not necessarily represent an equivalent reduction in organisational cost. If an employee previously spent ten minutes handling a request and AI reduces that requirement to two minutes of human oversight, the organisation has reduced labour input but has not necessarily eliminated the remaining eight minutes of work. Similarly, if the organisation retains the same number of employees after implementation, the resulting productivity improvement may initially appear as additional capacity rather than a direct cash saving.

8.1 Labour Savings Versus Capacity Release

The distinction between labour savings and capacity release is therefore central to the economic evaluation of customer-service AI. Labour savings occur when AI reduces the actual resources required to deliver a given level of service and the organisation can convert that reduction into lower expenditure. Capacity release occurs when AI enables the existing workforce to handle more work or devote more time to higher-value activities without a proportional increase in staffing.

Capacity release can nevertheless have substantial economic value. It may allow an organisation to accommodate higher customer volumes without recruiting additional employees, reduce waiting times during peak periods, improve service levels, support business growth, or redeploy employees toward complex and relationship-oriented interactions. In environments where demand is increasing or skilled labour is difficult to recruit, avoiding future staffing requirements may be economically more significant than reducing current headcount.

The empirical evidence from Brynjolfsson, Li, and Raymond (2025) illustrates this distinction particularly well. Their study of 5,172 customer-support agents found that generative-AI assistance increased issues resolved per hour by approximately 15%, with particularly strong effects among less experienced and lower-performing workers. The study demonstrates substantial productivity improvement but does not imply that organisations should automatically respond by reducing employment. Rather, higher productivity can produce different economic outcomes depending on organisational strategy and customer demand. If demand remains constant and productivity increases, organisations may potentially reduce labour requirements; if demand is growing, the same productivity improvement may enable the organisation to serve more customers without proportionally increasing employment (Brynjolfsson et al., 2025).

This distinction changes how the business case should be constructed. A narrow automation business case might calculate the number of interactions removed from human handling and multiply this figure by an estimated labour cost per interaction. Such an approach risks overstating the financial benefit because it implicitly assumes that every minute of avoided handling time becomes an immediate cash saving. A more realistic assessment should determine what happens to the released capacity.

For each use case, management should therefore identify whether the resulting capacity will be converted into direct cost reduction, additional service volume, shorter waiting times, improved quality, reduced overtime, avoidance of future recruitment, or higher-value employee activities. The economic value of AI depends substantially on this redeployment decision.

8.2 Components of the AI Business Case

A comprehensive business case should incorporate several categories of benefit and cost.

The first category is direct productivity and operating-cost effects. These include reductions in average handling time, manual data entry, information retrieval, documentation, routine correspondence, and other activities that can be automated or accelerated. The value should be calculated using actual organisational process data wherever possible rather than generic assumptions.

The second category is capacity effects. AI may enable the organisation to process a larger number of interactions with existing resources, particularly during demand peaks. Capacity should be valued according to the organisation's alternative use of that resource. If additional capacity prevents the need to hire additional employees, the relevant economic benefit may be the avoided recruitment and employment cost. If it enables higher sales or customer retention, the benefit may instead appear as incremental revenue.

The third category is quality-related value. More consistent information retrieval, improved knowledge access, and decision support may reduce errors, repeat contacts, escalations, and rework. These improvements can reduce operational costs while simultaneously improving customer experience. However, such benefits should be estimated carefully because they may be difficult to isolate from other process changes.

The fourth category is employee-related value. As discussed in Chapter 6, AI can accelerate onboarding, reduce training requirements, support less experienced employees, and allow employees to devote more time to higher-value activities. The findings of Brynjolfsson et al. (2025) are particularly relevant here because they indicate that generative AI can reduce performance differences between employees by making organisational knowledge and effective practices more accessible.

The fifth category is customer-related value. Faster response times, 24/7 availability, improved first-contact resolution, and more consistent information can increase customer satisfaction and potentially improve retention. Such benefits can be economically significant, although they should only be monetised where there is a defensible relationship between the service improvement and customer behaviour.

Against these benefits, the organisation must account for the full cost of ownership. This includes AI licences or infrastructure, implementation, integration with CRM and other enterprise systems, data preparation and migration, security controls, testing, employee training, process redesign, monitoring, model evaluation, maintenance, and ongoing governance. In regulated industries, compliance and audit requirements may represent a significant additional cost and should be included from the beginning rather than treated as an unexpected implementation expense.

The business case should also incorporate expected risk costs. AI systems can generate incorrect information, fail to escalate appropriate cases, expose sensitive data, or produce decisions that require remediation. Not every risk can be assigned a precise monetary value, but material risks should at least be identified, assessed qualitatively, and incorporated into the investment decision. For high-impact use cases, organisations may also estimate expected loss by combining the probability of an adverse event with its potential financial consequence.

8.3 Scenario-Based Business Cases

Because the future effects of AI are uncertain, a single-point estimate of return on investment can create a false sense of precision. A more robust approach is to construct several scenarios reflecting different assumptions about automation, employee redeployment, adoption, customer demand, and implementation costs.

A conservative scenario should assume that most of the productivity improvement is converted into capacity rather than immediate cost reduction. Employees remain in place, and the primary benefits arise from increased service capacity, reduced waiting times, improved employee productivity, and limited avoidance of future staffing requirements. This scenario is particularly appropriate where demand is growing or where employment cannot easily be reduced.

A base-case scenario should assume that AI generates measurable operational and financial improvements under realistic adoption and performance assumptions. Some capacity may be redeployed, some costs may be avoided, and service quality may improve. The base case should use empirical evidence from pilots or comparable processes wherever available rather than relying exclusively on vendor projections.

A transformative scenario should consider the possibility that AI enables a fundamental redesign of the service operating model. This might involve extensive automation of routine interactions, redesigned employee roles, continuous service availability, substantial changes in channel mix, or significant reductions in the marginal cost of serving customers. Such a scenario may create considerable value but also carries greater organisational, technological, and governance uncertainty. It should therefore not be used as the primary justification for investment without supporting evidence.

Scenario analysis can be extended by varying individual assumptions such as automation rate, average handling time, adoption rate, customer demand, AI operating cost, escalation rate, and employee redeployment. This allows management to identify which assumptions have the greatest influence on the business case. For example, if the economic case depends almost entirely on achieving a 70% automation rate, the organisation should recognise that assumption as a major source of uncertainty and test it explicitly through a pilot.

8.4 From Automation ROI to Service-System Value

The broader implication is that the economic evaluation of AI should move from a narrow "automation ROI" perspective toward a service-system value perspective. Automation is one mechanism through which value may be generated, but it is not the ultimate objective.

An AI assistant that increases employee productivity without reducing headcount can still generate significant value if it enables the organisation to absorb demand growth, improve response times, reduce employee turnover, or improve service quality. Similarly, a customer-facing chatbot that resolves only a subset of interactions may be economically attractive if those interactions are high-volume, low-risk, and expensive to handle manually. Conversely, an AI application with a high theoretical automation rate may produce little net value if it requires extensive human checking, generates frequent errors, or increases customer complaints and repeat contacts.

This perspective is consistent with the broader theoretical argument of Huang and Rust (2018), who conceptualise AI adoption as a transformation in the allocation of tasks between humans and machines, and with the empirical findings of Brynjolfsson et al. (2025), which demonstrate that AI can generate value through employee augmentation and knowledge transfer. It is also consistent with the Deepsearch (2026) framework, which places the business case alongside use-case definition, data readiness, stakeholder alignment, and measurable outcomes rather than treating financial value as an independent calculation.

Consequently, the final investment question should not be:

"How much labour can AI eliminate?"

but rather:

"What additional economic and service value can the AI-enabled operating model create relative to the current alternative, after accounting for technology, implementation, organisational, and risk costs?"

This formulation captures both substitution and augmentation effects and allows organisations to recognise forms of value that would otherwise be overlooked. It also creates a stronger basis for comparing alternative AI use cases: a use case with a lower automation rate may be more attractive than one with a higher rate if it produces greater customer value, stronger employee augmentation, lower risk, or a more favourable total cost of ownership.

Ultimately, the business case should connect the quantitative assumptions developed in the preceding sections into a coherent economic model. Volume, task suitability, automation or augmentation potential, handling-time reduction, employee capacity, customer outcomes, implementation costs, ongoing operating costs, and risk must be evaluated together. Only then can an organisation determine whether a customer-service AI use case represents a genuine improvement to the operating model rather than simply an impressive demonstration of technological capability.

9. Regulation, Privacy and Explainability

Regulation, privacy, transparency, and explainability are not peripheral considerations in customer-service AI; they are constraints on the design, deployment, and economic viability of the use case itself. Customer-service systems routinely process information about identifiable individuals and may, depending on the industry and process, handle financial information, health information, authentication data, contractual information, or other sensitive personal data. AI can also influence decisions concerning customers, determine how cases are routed or prioritised, generate communications that customers may rely upon, and increasingly execute actions through connected enterprise systems. Consequently, the regulatory assessment of an AI use case should begin during use-case identification and continue throughout its lifecycle rather than being performed only immediately before deployment.

The Deepsearch (2026) framework appropriately treats compliance and explainability as core elements of AI use-case evaluation. This principle can be strengthened by distinguishing three related but conceptually different questions. First, is the organisation legally permitted to process the relevant data and operate the proposed AI system? Second, are the AI system's outputs sufficiently transparent, controllable, and reliable for the intended application? Third, can the organisation demonstrate appropriate governance and accountability throughout the system's lifecycle? These questions are particularly important in regulated industries, where the consequences of an inappropriate automated response may extend beyond customer dissatisfaction to legal, financial, or reputational harm.

9.1 GDPR and Data Governance

The General Data Protection Regulation (GDPR) provides a central legal framework for customer-service AI where personal data are processed. The GDPR establishes requirements concerning, among other matters, lawfulness, fairness and transparency, purpose limitation, data minimisation, accuracy, storage limitation, integrity and confidentiality, and accountability.

For AI use cases, these principles have practical implications throughout the data and system architecture. The organisation must establish why particular data are being processed, whether the proposed processing is necessary and proportionate for the stated purpose, who is authorised to access the data, how long the data should be retained, and how security and accountability will be maintained. The fact that historical customer interactions are available does not, by itself, establish that they can automatically be used for training, fine-tuning, evaluation, or retrieval purposes. Data availability and data-use permission are fundamentally different questions.

Consequently, the AI Use Case Canvas should record, at a minimum, the categories of personal data involved, the intended processing purposes, the applicable legal basis, retention requirements, data locations, third-party processors, access controls, intended use of data for model training or evaluation, logging requirements, data-subject rights, and whether a Data Protection Impact Assessment (DPIA) or other formal assessment is required.

This should be established before technical architecture is finalised. For example, an organisation considering an AI assistant that retrieves information from historical customer emails should determine not only whether those emails can technically be indexed but also whether the proposed use is compatible with the original purpose of processing, whether access restrictions must be preserved, and whether personal data should be minimised, pseudonymised, or otherwise protected. Similarly, where an external AI provider processes customer information, the organisation must understand the provider's role, processing arrangements, data locations, retention practices, and permitted secondary uses.

Privacy therefore becomes a design variable in use-case selection. Two technically similar AI applications may have very different feasibility profiles because one requires extensive processing of personal information while the other can operate on non-personal or appropriately minimised organisational knowledge.

9.2 Automated Decision-Making and Human Involvement

The GDPR is particularly relevant where AI moves beyond providing information and begins to influence decisions about individual customers. Article 22 addresses automated individual decision-making, including profiling, and establishes protections where a person is subject to a decision based solely on automated processing that produces legal effects or similarly significant effects. Where relevant exceptions apply, the GDPR also requires safeguards including the possibility of human intervention, the ability to express one's point of view, and the ability to contest the decision.

This distinction reinforces an important principle developed throughout this paper: the regulatory implications of an AI use case depend on what the system actually does, not simply on whether it is labelled a chatbot, copilot, or AI assistant. An AI system that retrieves a frequently asked question is materially different from one that recommends whether a customer's claim should be accepted, determines eligibility for a benefit, or autonomously changes contractual conditions.

Use-case descriptions should therefore specify whether AI is merely generating information, recommending an action, supporting a human decision, or executing a decision autonomously. The greater the consequence of the AI's output, the stronger the requirements for validation, human oversight, documentation, traceability, and escalation are likely to become.

This also provides a regulatory rationale for the distinction between automation, augmentation, and decision support developed in Chapter 2. Human involvement is not simply a technological fallback; in some use cases it is an important governance mechanism.

9.3 The EU AI Act

The EU Artificial Intelligence Act adds a further layer of governance through a risk-based regulatory framework. Its relevance to customer service depends on the characteristics and intended use of the AI system rather than on the generic label "customer-service AI." Organisations should therefore avoid treating every chatbot or generative-AI application as subject to the same regulatory requirements.

The Deepsearch guide identifies Article 13 of the AI Act as an important transparency and explainability consideration. This reference requires qualification. Article 13 specifically concerns transparency and provision of information for high-risk AI systems. Its purpose is to ensure that high-risk systems are sufficiently transparent to enable deployers to interpret outputs and use the systems appropriately, including information concerning capabilities, limitations, performance, and relevant human-oversight requirements.

For customer-service applications, however, another provision has become particularly important: Article 50, which establishes specific transparency obligations for certain AI systems. The European Commission's current guidance confirms that Article 50 applies from 2 August 2026.

Article 50 is directly relevant to customer-service chatbots and other AI systems that communicate directly with natural persons. The European Commission explains that providers of AI systems designed to interact directly with people must design them so that individuals are informed that they are interacting with an AI system, unless this is obvious to a reasonably well-informed, observant, and circumspect person. The notification should occur from the beginning of the first interaction and be presented clearly and distinguishably.

This is especially relevant as conversational AI becomes increasingly capable of producing interactions that resemble human communication. The regulatory objective is not to prohibit natural interaction but to prevent customers from being misled about the nature of the system with which they are interacting. Transparency therefore contributes to appropriate calibration of customer trust: customers should be able to understand when they are dealing with an AI system and make informed decisions about whether to continue, provide information, or request human assistance.

The European Commission's current guidance confirms that the Article 50 transparency requirements apply from 2 August 2026. As the present paper is concerned with customer-service AI deployment in the European context, this date should therefore be treated as a current compliance milestone rather than a future consideration.

Importantly, Article 50 should not be conflated with the broader requirements applicable to high-risk AI systems. A customer-service chatbot that directly interacts with customers may be subject to Article 50 transparency requirements without thereby becoming a high-risk AI system under the AI Act. Conversely, where a customer-service AI application performs a function that falls within a high-risk category, additional requirements may apply. The regulatory assessment must therefore be performed at the level of the specific AI system and intended use case.

9.4 Transparency as a Design Principle

The regulatory requirements support a broader managerial principle: transparency should be designed into the customer journey rather than added as a compliance notice at the end of implementation.

For a customer-facing AI system, transparency should encompass at least three questions. First, does the customer know that they are interacting with AI where disclosure is required? Second, does the customer understand what the system can and cannot reliably do? Third, does the customer have an appropriate route to human assistance when the AI cannot resolve the matter?

The third question is particularly important because transparency without meaningful escalation may provide formal disclosure without providing meaningful customer protection. A customer who is told that they are interacting with AI but is then prevented from reaching an appropriate employee may still experience a significant loss of control.

This connects regulation directly with the customer-experience findings discussed in Chapter 5. Research indicates that customers value usefulness, information quality, and successful problem resolution, while poor problem-solving capability and insufficient human interaction can generate negative reactions to AI service (Ashfaq et al., 2020; Zhao et al., 2022). Regulatory transparency should therefore reinforce, rather than substitute for, good service design.

9.5 Explainability, Traceability and Auditability

Explainability is also important beyond the narrow question of whether a customer can receive an explanation of an individual AI output. For organisational governance, the more fundamental requirement is often traceability: the ability to establish what information the system used, what it produced, under which configuration or model version it operated, what human intervention occurred, and why a particular action was taken.

This becomes particularly important when AI is integrated with CRM, ticketing, workflow, telephony, or RPA systems. An AI system that merely drafts an email presents a different governance challenge from one that autonomously changes customer records or initiates a transaction. As the system's degree of autonomy increases, the organisation's need for logging, monitoring, access controls, version management, and auditable decision pathways also increases.

Explainability should consequently be evaluated relative to the purpose and risk of the AI application. Not every customer-service recommendation requires an elaborate technical explanation of the underlying model. However, customers, employees, auditors, regulators, or internal governance functions may need to understand the relevant factors behind a decision or recommendation. In higher-impact applications, the organisation should be able to demonstrate that the system operates within a defined scope and that appropriate human oversight is available.

This perspective is consistent with the task-level approach developed earlier in the paper. An AI system performing low-risk information retrieval may require relatively limited decision transparency, whereas a system recommending or executing consequential customer decisions requires substantially stronger controls.

9.6 Privacy, Regulation and Use-Case Selection

Regulation should therefore not be treated as a final "go/no-go" check after an AI use case has already been designed. It should influence the selection and prioritisation of use cases from the outset.

Consider two otherwise similar customer-service opportunities. The first automatically answers routine questions using publicly available product information. The second uses detailed customer records to make personalised recommendations that could influence a financially significant decision. Both may offer substantial automation potential, but their privacy, regulatory, explainability, human-oversight, and risk profiles are fundamentally different.

The second use case may still be valuable, but it requires a correspondingly stronger governance architecture. The first may be a better initial candidate for deployment because it combines high volume and structural regularity with relatively low consequence of error. This illustrates why regulatory feasibility should be incorporated into the AI Use Case Canvas alongside business value, technical feasibility, and customer impact.

The resulting assessment should establish, for each candidate use case, what data are processed, for what purpose, under which legal basis, where the data and AI services are hosted, who can access them, whether personal data are used for model training, whether automated decision-making is involved, what level of human oversight is required, what transparency obligations apply, and how system behaviour will be monitored and audited.

9.7 Regulation as a Source of Trust and Competitive Advantage

Although regulation is often presented primarily as a constraint, effective governance can also become a source of organisational value. Clear data practices, transparent AI interactions, appropriate human escalation, and auditable decision processes can strengthen customer and employee trust. This is particularly important in customer service because customers may be reluctant to provide personal information or rely on AI-generated advice when they do not understand how the system operates.

The European Commission explicitly links the Article 50 transparency requirements to enabling people to make informed choices and appropriately calibrate their trust in AI-generated content and interactions. In this sense, transparency is not merely a legal disclosure obligation; it is part of the broader architecture of trustworthy AI.

For organisations operating in regulated sectors, this can become a differentiating capability. A customer-service system that is demonstrably accurate, transparent about its limitations, appropriately supervised, and capable of providing human escalation may generate greater long-term trust than a system optimised primarily for maximum automation.

The central implication is therefore that compliance, privacy, explainability, and customer experience should be designed as mutually reinforcing elements of the AI service architecture. GDPR establishes important constraints on the processing and use of personal data; the AI Act introduces additional risk- and transparency-based obligations; and organisational governance determines how these requirements are translated into operational controls.

Accordingly, the regulatory question in AI use-case evaluation should not simply be whether an application is "compliant." It should be whether the proposed AI-enabled process can operate lawfully, transparently, securely, proportionately, and with an appropriate level of human accountability throughout its lifecycle. This reframes regulation from a barrier to innovation into a design criterion for selecting AI applications that are both economically valuable and sustainable in practice.

10. Explainability and Auditability

In regulated customer-service environments, explainability should not be understood exclusively as the ability to describe the mathematical or computational internals of an AI model. While technical interpretability can be important, organisational decision-making requires a broader form of traceability: the ability to reconstruct how an AI-supported process operated, what information was available, what the system produced, how humans interacted with the output, and why the resulting action was taken. Explainability should therefore be considered not only a property of the model but also a governance capability of the overall AI-enabled process.

This distinction is particularly important in customer service because AI systems increasingly operate as components within larger workflows. A generative-AI model may retrieve information from a knowledge base, generate a recommendation, provide it to an employee, and subsequently trigger an action through a CRM, ticketing, workflow, or RPA system. Assessing the explainability of the language model alone would provide an incomplete picture of the organisation's ability to understand and govern the resulting decision.

A more useful concept is therefore decision traceability. For material AI-supported actions, the organisation should, where appropriate and proportionate to the risk of the use case, be able to reconstruct the relevant stages of the process. This includes identifying what information was available to the system, which AI system generated the recommendation, which model or version was used, what output was produced, and what degree of confidence or uncertainty was available. It should also be possible to establish whether a human reviewed the output, what final action was taken, and whether that action was consistent with the applicable policy or process rules.

This traceability requirement transforms explainability from an abstract technical concept into a practical organisational control mechanism. The relevant question is not necessarily "Can we explain every internal computation performed by the model?" but rather "Can we demonstrate, after the fact, how this AI-supported decision was produced and whether it was appropriate under the rules governing the process?"

10.1 From Model Explainability to Process Traceability

The distinction between model explainability and process traceability is particularly important for contemporary generative-AI systems. Large language models can produce complex outputs based on learned statistical relationships, and a complete causal explanation of an individual output may not always be technically available or meaningful to an operational user. Nevertheless, organisations can establish controls around the inputs, retrieval sources, prompts or system instructions where relevant, outputs, model versions, human interventions, and subsequent actions.

For example, consider an AI system supporting an employee handling an insurance-related customer request. The system may retrieve relevant policy information, summarise the customer's history, recommend an appropriate response, and present the recommendation to the employee. The organisation should be able to determine which policy information was retrieved, which AI system and version generated the recommendation, what recommendation was presented, whether the employee accepted or modified it, and what communication was ultimately sent to the customer.

This level of traceability can be more operationally valuable than an abstract explanation of the model's internal parameters. It enables the organisation to investigate errors, respond to customer complaints, identify systematic problems, demonstrate adherence to policy, and determine whether an AI system continues to operate within its validated scope.

The appropriate degree of traceability should nevertheless be proportionate to the consequences of the use case. A system that drafts a low-risk response to a frequently asked question does not necessarily require the same level of audit infrastructure as a system that influences a consequential decision about an individual customer. As the potential impact of an AI-supported action increases, so too should the requirements for logging, review, oversight, and reconstruction of the decision process.

10.2 The Role of Human Oversight

Human oversight is a further component of explainability and auditability. Recording that an AI system generated an output is insufficient if the organisation cannot determine whether a human reviewed it, whether the human had sufficient information and authority to challenge it, and whether the final decision was genuinely under human control.

This is particularly important where employees are expected to validate AI recommendations. A nominal "human-in-the-loop" arrangement can provide limited protection if employees are required to approve large volumes of AI-generated decisions without sufficient time, information, or authority to question the system. Effective human oversight requires employees to understand the system's capabilities and limitations and to have a meaningful opportunity to intervene where appropriate.

The EU AI Act reflects this broader relationship between transparency and human oversight for high-risk AI systems. Article 13 requires high-risk systems to be sufficiently transparent to enable deployers to interpret outputs and use the system appropriately, while the Act also establishes requirements concerning human oversight and record-keeping for high-risk systems. These requirements reinforce the principle that transparency should support effective human control, rather than simply provide technical documentation. The precise obligations depend on whether the particular AI system and intended use fall within the Act's high-risk categories.

This distinction is important for the AI Use Case Canvas developed in this paper. A use case should not merely identify that "human oversight" exists. It should specify who performs the oversight, at what stage, based on what information, with what authority, and under which conditions intervention is mandatory.

10.3 Logging and Audit Trails

Auditability depends fundamentally on appropriate logging. Depending on the risk and purpose of the application, relevant records may include the source information provided to the system, the retrieved knowledge or documents, model and system versions, timestamps, relevant system instructions or configuration, generated outputs, confidence or validation signals where available, human interventions, final decisions, and subsequent process outcomes.

Such records serve several purposes. First, they support incident investigation when an AI system produces an inappropriate output. Second, they support quality assurance by allowing organisations to identify recurring patterns of error. Third, they support regulatory and internal audits by providing evidence of how the system operated. Fourth, they support continuous improvement by identifying weaknesses in knowledge sources, prompts, workflows, or escalation rules.

Logging must, however, be designed consistently with privacy and data-governance requirements. Recording every element of an interaction without considering data minimisation, access controls, retention, and security can itself create regulatory and operational risks. Auditability should therefore not mean unlimited data retention. Rather, organisations should determine what information is necessary to demonstrate accountability and manage risk and should apply appropriate controls to that information.

This creates a direct connection between the present chapter and the GDPR considerations discussed in Chapter 9. Traceability and privacy must be designed together. An audit trail that is technically comprehensive but unnecessarily exposes personal data may create a new governance problem. Conversely, insufficient logging can make it impossible to investigate incidents or demonstrate appropriate control. The appropriate design is therefore a risk-based balance between accountability, data minimisation, security, and operational necessity.

10.4 Auditability as a Vendor-Selection Criterion

These considerations have direct implications for procurement. In regulated customer service, organisations should treat auditability as a core vendor-selection criterion rather than an optional technical feature. Vendor evaluations should examine not only model accuracy and automation performance but also the organisation's ability to monitor, reconstruct, and govern system behaviour.

Relevant questions include whether the vendor provides access to appropriate audit logs, whether model and system versions can be identified retrospectively, whether relevant retrieval sources can be traced, how changes to models or configurations are documented, whether human interventions can be recorded, how long logs can be retained, and whether the organisation can export the information required for its own governance and regulatory obligations.

The organisation should also establish whether the vendor can support meaningful performance monitoring after deployment. AI systems are not static. Models may be updated, knowledge bases may change, prompts or system configurations may be modified, and customer behaviour may evolve. A system that was appropriately validated at deployment may therefore behave differently later. Effective auditability requires the organisation to determine which version of the system produced which outcome and under what operating conditions.

This is especially important when AI services are provided through external vendors or cloud platforms. The organisation may remain accountable for the customer-service process even when parts of the AI infrastructure are operated by a third party. Vendor contracts, therefore, should address appropriate access to logs, incident notification, change management, data handling, security, retention, and cooperation with audits where required.

10.5 Auditability and the AI Use Case Canvas

The principles developed in this chapter suggest that the governance dimension added to the Deepsearch (2026) AI Use Case Canvas should explicitly include explainability, traceability, logging, human oversight, and auditability.

For each candidate use case, organisations should establish what level of traceability is required given the potential consequences of an error. Low-risk informational applications may require relatively lightweight logging, while higher-impact applications may require comprehensive records of inputs, outputs, model versions, human interventions, and final decisions.

This approach also helps organisations compare otherwise similar use cases. Two applications may have similar expected automation rates and financial benefits but very different governance profiles. A use case requiring limited personal data, producing reversible outputs, and operating under straightforward human review may be considerably easier to govern than an application that processes sensitive information and autonomously initiates consequential customer actions.

Auditability therefore becomes part of use-case prioritisation, not merely implementation. A technically feasible and economically attractive use case may nevertheless be a poor initial candidate if the organisation cannot establish adequate controls around its operation.

10.6 From Explainability to Accountable AI

The broader conclusion is that explainability in customer-service AI should be understood as part of a wider concept of accountable AI. Accountability requires more than understanding how a model works. It requires the organisation to be able to demonstrate that the AI system was appropriately selected, trained or configured, integrated, monitored, used within its intended scope, and subject to suitable human oversight.

For this reason, the relevant governance question is not simply:

"Can we explain the model's output?"

It is:

"Can we reconstruct, evaluate, and justify the AI-supported process that produced the outcome?"

This shift is particularly important in regulated industries. An organisation may not be able to provide a complete mathematical explanation of every generative-AI output, but it can and should establish appropriate controls around the system's data sources, operating scope, model version, outputs, human review, decisions, and consequences.

Accordingly, auditability should be treated as a fundamental selection criterion when evaluating customer-service AI vendors. A system that delivers impressive benchmark performance but cannot provide sufficient transparency, logging, version control, or evidence of human intervention may create unacceptable governance risk. Conversely, an AI solution that combines adequate technical performance with strong traceability and controllability can provide a substantially more defensible foundation for deployment.

This reinforces the central proposition of the paper: AI use-case selection should evaluate technological feasibility, customer value, employee impact, economic value, and regulatory requirements as an integrated decision problem. Explainability and auditability form an essential part of that integration because they enable the organisation to demonstrate, rather than merely assume, that an AI-enabled customer-service process remains within its intended operational and regulatory boundaries.

11. Extended AI Customer Service Use Case Canvas

a The evidence reviewed in this paper supports extending the six-dimensional Deepsearch (2026) model into a seven-dimensional framework for identifying, evaluating, and prioritising AI use cases in customer service. The original Deepsearch framework provides a strong managerial foundation by connecting the strategic decision basis, use-case definition, data and system integration, stakeholders, success metrics, and business case. The academic and regulatory literature reviewed in the preceding chapters suggests, however, that governance and risk should be established as a distinct dimension rather than being treated implicitly within the other categories.

This extension is important because AI use-case selection is not a one-dimensional optimisation problem. A technically feasible application may produce significant economic benefits but fail because the required data cannot lawfully be processed. A highly accurate system may nevertheless be inappropriate if customers cannot obtain human assistance when needed. An application with a compelling automation rate may create little net value if employees must extensively review its outputs. Conversely, a use case with relatively modest automation potential may be highly attractive because it substantially improves employee capability, customer experience, or service capacity.

The resulting framework should therefore be understood as an integrated decision architecture. Each dimension addresses a different question, but the dimensions are interdependent. A use case should advance toward implementation only when sufficient evidence exists across all seven dimensions.

11.1 Strategic Basis

The first dimension concerns the fundamental business problem: What problem is the organisation attempting to solve, and why is AI an appropriate means of addressing it?

The starting point should be an evidence-based description of the existing service process rather than a predetermined technology. Organisations should establish baseline information concerning customer-service volumes, request categories, average handling time, response and resolution times, first-contact resolution, escalation rates, customer satisfaction, staffing requirements, peak-demand patterns, error rates, and operating costs. These data provide the reference point against which the effects of an AI intervention can subsequently be evaluated (Deepsearch, 2026).

This dimension prevents technology-led problem definition. "Implement a chatbot" is not, in itself, a meaningful business problem. A stronger formulation might be that the organisation receives a high volume of routine email requests that require substantial employee time despite following standardised procedures, resulting in long response times during peak periods. AI can then be evaluated as one potential intervention for that problem.

The strategic basis should also establish why the problem matters to the organisation. A use case may be strategically attractive because it reduces cost, improves customer experience, increases scalability, addresses capacity constraints, supports employee capability, or strengthens service resilience. The expected contribution should be explicit before the technical solution is selected.

11.2 Use Case, Channel, and Task

The second dimension asks: What precisely will the AI system do, through which channel, for which customer or employee, and under what conditions?

This requires moving from broad descriptions such as "automate customer service" toward clearly bounded tasks and workflows. A candidate use case should identify the customer segment, channel, request type, process trigger, AI action, required data, expected output, human role, and conditions under which escalation occurs.

The task-level perspective developed by Huang and Rust (2018) provides an important theoretical foundation for this dimension. Customer service is not a homogeneous activity; individual interactions consist of multiple tasks with different degrees of suitability for automation or augmentation. Intent classification, information retrieval, summarisation, and response drafting may be relatively suitable for AI, while negotiation, complex complaints, exceptional cases, and emotionally sensitive interactions may require greater human involvement.

The channel should also be considered explicitly. Email, chat, telephone, and employee-facing applications create different technical, operational, and customer-experience requirements. The Deepsearch (2026) guide highlights email and telephone as important opportunities because of their substantial service volumes, but channel volume alone should not determine prioritisation. The relevant combination is volume, structural regularity, task complexity, risk, and expected value.

A well-defined use case should therefore be sufficiently specific to support testing. For example, "AI for customer emails" remains too broad. "AI classifies standard address-change requests, retrieves the applicable procedure, verifies predefined requirements, and either initiates an approved workflow or drafts a response for employee approval" defines a process that can be evaluated in terms of accuracy, automation potential, exception rates, customer outcomes, and risk.

11.3 Data and System Integration

The third dimension asks: Does the AI system have reliable, lawful, and technically appropriate access to the information and systems required to perform the task?

AI performance depends not only on model capability but also on the quality and accessibility of the information surrounding the model. Relevant sources may include historical emails, chat and call transcripts, knowledge bases, FAQs, product and tariff information, policy documents, CRM records, ticketing systems, workflow engines, telephony platforms, and RPA interfaces (Deepsearch, 2026).

Three forms of feasibility should be distinguished.

First, there is data feasibility: whether the system can access sufficiently accurate, current, and relevant information to perform the task. Second, there is technical feasibility: whether the AI can connect reliably to the enterprise systems required to support or execute the process. Third, there is operational feasibility: whether the resulting integrated workflow can function reliably and at the required scale.

Data availability must also be distinguished from permission to use data. Historical customer interactions may exist in large quantities but may not automatically be appropriate for model training, retrieval, or other secondary purposes. GDPR requirements concerning lawfulness, purpose limitation, data minimisation, security, and accountability therefore need to be considered alongside technical data readiness.

The result is that data and integration should be evaluated as an end-to-end capability, rather than simply asking whether the organisation has "enough data."

11.4 Human–AI Roles and Stakeholder Alignment

The fourth dimension asks: What should AI do, what should humans do, and how should responsibility be allocated between them?

This dimension combines two issues that are closely connected in practice: human–AI task allocation and stakeholder alignment. Huang and Rust (2018) provide a theoretical basis for analysing which forms of intelligence and which tasks are more amenable to AI, while Xiao and Kumar (2021) emphasise that service technologies operate within a broader human–technology system involving employee and customer acceptance.

The appropriate role of AI may range from autonomous execution to employee augmentation or decision support. The choice should depend on task complexity, risk, customer expectations, reversibility of errors, and the value of human judgement.

Stakeholder alignment is equally important. Customer-service leadership, frontline employees, IT, data and analytics teams, information security, legal and compliance, data-protection officers, process owners, telecommunications teams, senior management, and employee representatives where applicable may all have legitimate interests in the deployment. Customers are also central stakeholders because the resulting service design directly affects their experience.

A robust use-case assessment should therefore establish clear ownership and accountability. A RACI-type structure can clarify who is responsible for the process, who owns the AI system, who approves changes, who monitors performance, who manages exceptions, and who has authority to suspend the system. The escalation model should also specify when human intervention is mandatory.

The importance of this dimension is reinforced by Brynjolfsson et al. (2025), whose study of 5,172 customer-support agents found substantial productivity gains from generative-AI assistance, particularly among less experienced employees. Their findings demonstrate that AI can create value through augmentation, knowledge transfer, and capability development, not simply through employee replacement. Consequently, the human–AI role allocation should be explicitly evaluated as part of the business case.

11.5 Customer and Operational Metrics

The fifth dimension asks: How will the organisation determine whether the AI-enabled process is actually better?

Success measurement should combine customer outcomes, operational outcomes, and AI-quality indicators. Relevant customer measures include CSAT, Customer Effort Score, NPS, complaint rates, abandonment, and escalation. Operational measures include AHT, FCR, response and resolution time, containment or automation rate, and transfer rates. AI-quality measures may include intent-recognition accuracy, answer accuracy, retrieval accuracy, hallucination rates, inappropriate-response rates, and escalation precision.

The academic evidence indicates why these dimensions must be evaluated together. Ashfaq et al. (2020) show that information and service quality influence customer satisfaction with chatbots, while Xie et al. (2024) identify utilitarian gratification as a particularly important determinant of chatbot satisfaction. Zhao et al. (2022), in contrast, demonstrate that poor problem-solving capability and insufficient human interaction can generate substantial negative customer reactions.

The implication is that automation rate cannot serve as a sufficient success criterion. A system that contains a high proportion of interactions but generates repeat contacts, customer frustration, or inappropriate responses may perform well according to a narrow operational metric while reducing overall service value.

Metrics should therefore be defined before implementation and linked to a baseline. The Deepsearch (2026) targets, such as 50–60% automation and greater than 84% intent recognition, can provide initial management hypotheses but should be adapted to the specific use case rather than treated as universal benchmarks.

11.6 Business Case

The sixth dimension asks: What financial and strategic value does the AI-enabled process generate after accounting for its complete cost and implementation requirements?

The economic evaluation should incorporate direct productivity effects, capacity release, avoided recruitment, improved service levels, potential revenue effects, implementation costs, technology costs, integration, training, maintenance, monitoring, and risk-related costs.

The distinction between labour savings and capacity release is particularly important. AI may reduce the time required to handle customer requests without reducing the number of employees. In such circumstances, the immediate financial effect may be limited, but the organisation may gain additional capacity that can be used to accommodate demand growth, reduce waiting times, avoid future recruitment, or enable employees to focus on higher-value activities.

This is consistent with Brynjolfsson et al. (2025), whose evidence shows substantial productivity improvement without establishing that AI necessarily leads to employment reductions. The economic value of increased productivity depends on how organisations redeploy the resulting capacity and how customer demand evolves.

Business cases should therefore include conservative, base-case, and transformative scenarios. This approach avoids treating ambitious automation assumptions as certain outcomes and makes explicit which assumptions drive the expected return.

11.7 Governance and Risk

The seventh dimension asks: Is the application legally, ethically, securely, transparently, and operationally acceptable?

This dimension represents the principal extension of the Deepsearch framework. The evidence reviewed in Chapters 9 and 10 demonstrates why governance should be made explicit. A use case may be technically feasible, strategically attractive, and economically compelling while nevertheless being unsuitable because the required data cannot lawfully be processed, the system's decisions cannot be adequately controlled, the risks of error are disproportionate, or the organisation cannot provide sufficient auditability.

The governance assessment should therefore consider GDPR requirements, the applicable provisions of the EU AI Act, data protection, security, human oversight, transparency, explainability, auditability, and operational resilience. It should establish what personal data are processed, for what purpose, under which legal basis, where data and AI services are hosted, whether data are used for model training, what transparency obligations apply, whether automated decision-making is involved, and how customers and employees can escalate problematic cases.

Auditability should also be treated as a core selection criterion. For material AI-supported actions, organisations should be able to determine, where proportionate to the use case, what information was available, which AI system and model version generated the output, what was produced, whether human intervention occurred, and what final action was taken. This converts explainability from an abstract technical objective into a practical accountability mechanism.

The governance dimension should therefore operate as both a constraint and a prioritisation mechanism. High-risk use cases may still be strategically valuable, but they require stronger controls and may be less suitable as initial deployment candidates than lower-risk applications with comparable economic potential.

11.8 The Seven Dimensions as an Integrated Decision Model

The seven dimensions should not be interpreted as seven independent checklists. Their principal value lies in their interaction. Strategic relevance without data feasibility produces an attractive but impractical idea. Technical feasibility without customer value produces an efficient solution to the wrong problem. Economic value without governance creates an unsustainable business case. Strong AI performance without appropriate human roles may generate employee resistance or customer harm. Conversely, a use case that performs well across all seven dimensions is more likely to produce sustainable value.

The framework can therefore be understood as a sequence of increasingly demanding questions:

First, is there a sufficiently important business problem?

Second, can the problem be decomposed into a clearly defined AI-suitable task?

Third, are the required data and system integrations available and permissible?

Fourth, is there an appropriate allocation of responsibilities between AI and humans?

Fifth, can improvement be demonstrated through customer, operational, and AI-quality measures?

Sixth, does the expected value exceed the full economic cost and risk-adjusted investment?

Seventh, can the resulting system operate within acceptable legal, ethical, security, transparency, and governance boundaries?

Only when these questions can be answered with sufficient evidence should an organisation proceed from use-case identification to experimentation, procurement, and implementation.

The resulting seven-dimensional model therefore extends the practical value of the Deepsearch (2026) AI Use Case Canvas while incorporating the principal insights from the academic literature reviewed in this paper. Huang and Rust (2018) provide the task-level foundation for determining where AI can substitute for or augment human capabilities. Xiao and Kumar (2021) emphasise the importance of the wider human–technology service system. Ashfaq et al. (2020), Xie et al. (2024), Zhao et al. (2022), and Zhang et al. (2026) demonstrate why customer experience and interaction characteristics must be incorporated into use-case evaluation. Brynjolfsson et al. (2025) demonstrate the potential for AI to augment employee productivity and distribute organisational knowledge. The GDPR and EU AI Act provide the regulatory context within which these applications must operate.

The central proposition of the extended framework is consequently that AI use-case selection should be treated as a multi-dimensional organisational decision rather than a technology-selection exercise. The objective is not to identify the use cases with the highest theoretical automation potential. It is to identify those applications in which the combination of strategic relevance, task suitability, data and integration readiness, human–AI complementarity, measurable service improvement, economic value, and governance feasibility creates a sufficiently strong case for action.

In this sense, the seventh dimension is not an additional compliance layer placed on top of the original framework. Governance and risk are what determine whether value identified in the other six dimensions can be realised sustainably. A use case that is economically attractive and technically feasible but legally or operationally unacceptable is not a viable AI use case. Conversely, a use case that satisfies all seven dimensions provides a substantially stronger foundation for responsible scaling and long-term organisational value.

12. Prioritisation Model

Once candidate AI use cases have been identified and evaluated across the seven dimensions, organisations require a method for determining which use cases should be pursued first. Identification alone does not establish priority: organisations typically face constraints in budget, data availability, technical capacity, change-management resources, and governance capability. A prioritisation model should therefore integrate expected value with feasibility and risk rather than ranking use cases according to technological sophistication or automation potential alone.

The logic developed in the preceding chapters suggests that prioritisation should be based on five aggregated dimensions: value, feasibility, data readiness, compatibility, and risk. A general prioritisation score for use case can be expressed as:

where represents expected business and customer value, represents technical and operational feasibility, represents data readiness, represents organisational and customer compatibility, and represents regulatory and operational risk. The terms represent organisation-specific weights reflecting strategic priorities.

The model should be understood as a decision-support mechanism rather than a mathematically objective measure of AI value. The weights necessarily reflect managerial judgement and organisational priorities. A regulated financial-services organisation, for example, may assign substantially greater weight to governance and operational risk than an organisation deploying low-risk informational chatbots. Similarly, an organisation experiencing severe service-capacity constraints may place greater weight on operational value and scalability.

12.1 Business and Customer Value

The first component, V, captures the expected value generated by the use case. This should extend beyond direct cost reduction. As established in the preceding chapters, AI can generate value through reduced handling time, faster responses, increased service capacity, improved first-contact resolution, greater availability, improved customer experience, employee augmentation, and potentially additional revenue.

The customer dimension is particularly important. Research on AI-mediated service demonstrates that customers evaluate AI according to factors such as information quality, service quality, usefulness, problem-solving capability, and the nature of the interaction (Ashfaq et al., 2020; Xie et al., 2024; Zhao et al., 2022). Consequently, a use case should receive a high value score only where the expected operational improvement translates into meaningful customer or organisational outcomes.

The value assessment should ideally be based on the baseline established under Dimension 1 of the extended Use Case Canvas. For example, a use case processing 10,000 requests per month with long response times may have substantially greater potential value than a technically sophisticated application affecting only a few hundred low-impact interactions.

12.2 Technical and Operational Feasibility

The second component, F, captures whether the organisation can realistically implement and operate the proposed AI application. This includes model capability, system integration, API availability, workflow compatibility, infrastructure requirements, scalability, reliability, and the organisation's ability to operate and monitor the resulting solution.

Feasibility should not be confused with the existence of an AI model capable of performing the task in a demonstration environment. A customer-service AI application may perform well in isolation while failing to create operational value because it cannot access the required CRM information, execute transactions, handle exceptions, or integrate reliably with existing workflows.

The distinction between technical feasibility and operational feasibility is therefore important. A use case should score highly only when the complete process—not merely the AI component—can operate reliably at the required scale.

12.3 Data Readiness

The third component, D , assesses whether the organisation possesses the information required to develop, evaluate, and operate the AI application. Relevant considerations include data availability, completeness, accuracy, consistency, timeliness, representativeness, accessibility, and governance.

Data readiness should also incorporate the legal and organisational conditions governing data use. As discussed in Chapters 9 and 10, the existence of historical customer data does not automatically imply that the data can be freely reused for AI training or retrieval. Data minimisation, purpose limitation, access rights, retention, security, and other governance requirements may materially affect the feasibility of a proposed use case.

This makes data readiness a particularly important differentiator between superficially attractive use cases. Two customer-service processes may have similar automation potential, but the one supported by high-quality, well-governed, readily accessible data may be substantially more suitable for initial implementation.

12.4 Organisational and Customer Compatibility

The fourth component, C, captures the degree to which the proposed AI-enabled process fits the organisation's operating model and the expectations of its customers and employees.

Organisational compatibility includes employee acceptance, management support, process ownership, required skills, change-management capacity, governance maturity, and the extent to which existing responsibilities and workflows must be redesigned. Customer compatibility concerns whether customers are likely to accept the proposed form of AI interaction and whether the technology is appropriate for the nature of the service interaction.

This dimension follows directly from the human–AI complementarity perspective developed by Huang and Rust (2018) and Xiao and Kumar (2021). AI adoption changes the distribution of tasks between people and technology rather than simply adding another software component. The findings of Brynjolfsson et al. (2025) further suggest that AI can generate substantial benefits by augmenting employees and transferring knowledge, particularly among less experienced workers. A use case that strengthens employee capability may therefore be more valuable than one that maximises autonomous automation but creates significant resistance or weakens service quality.

Customer compatibility should similarly distinguish between transactional and relational interactions. Zhang et al. (2026) find that AI can perform particularly strongly in competency-oriented interactions, whereas humans retain advantages in relational contexts requiring warmth and trust. Consequently, customer compatibility should reflect not only whether customers technically can use an AI system but whether the proposed interaction is appropriate for the customer's underlying need.

12.5 Regulatory and Operational Risk

The fifth component, R, represents the downside risk associated with the use case and is subtracted from the overall score. Relevant risks include privacy violations, security incidents, inappropriate outputs, hallucinations, discriminatory outcomes, regulatory non-compliance, inadequate human oversight, poor auditability, reputational damage, and operational failure.

Risk should be assessed in relation to both the probability and consequence of failure. A low-probability error may still make a use case inappropriate if the consequences are severe. Similarly, a use case with frequent but easily reversible errors may remain attractive if strong monitoring and human escalation mechanisms are available.

This is consistent with the governance dimension introduced in Chapter 11. Regulatory risk should not simply be represented by a generic compliance score. The assessment should consider the specific data processed, the nature of the AI-supported decision, the potential consequences for customers, the applicable GDPR and EU AI Act requirements, the degree of autonomy, the availability of human intervention, and the organisation's ability to audit system behaviour.

Importantly, risk should not necessarily be interpreted as a reason to reject every high-risk use case. Rather, it should influence sequencing and control requirements. A strategically important high-risk application may be pursued after appropriate governance capabilities have been established, while lower-risk applications may provide better candidates for an initial deployment.

12.6 Scoring and Weighting

For practical implementation, each component can be scored on a five-point scale. The organisation might, for example, assess expected value, feasibility, data readiness, and compatibility from 1 (very low) to 5 (very high), while scoring risk from 1 (low risk) to 5 (very high risk). The weights should then be established before individual use cases are scored wherever possible to reduce the risk that managers unconsciously manipulate the criteria to favour a preferred project.

The numerical result should nevertheless be treated as a structured conversation rather than an automatic decision rule. A weighted score can create transparency concerning why one use case ranks above another, but it cannot eliminate managerial judgement. Some criteria may also represent hard constraints rather than trade-offs. For example, a use case that cannot lawfully process the required personal data should not necessarily be rescued by a sufficiently high business-value score.

A useful approach is therefore to combine the weighted prioritisation score with minimum eligibility criteria. Legal permissibility, minimum data quality, security requirements, and essential human-oversight capabilities can function as gates. Only use cases that satisfy these conditions enter the final ranking.

12.7 From Ranking to Portfolio Management

The model also supports portfolio-level decision-making. Organisations should avoid assuming that the single highest-scoring use case is necessarily the only or best investment. A balanced AI portfolio may contain several types of initiatives.

Low-risk, high-feasibility applications can provide early operational wins and organisational learning. Employee-augmentation applications can build internal AI capability and demonstrate productivity benefits. More complex autonomous applications can subsequently be pursued once data, governance, and change-management capabilities have matured.

This portfolio perspective is consistent with the evidence from Brynjolfsson et al. (2025), which demonstrates that substantial value can arise from relatively straightforward augmentation rather than full automation. It also reflects the task-level logic of Huang and Rust (2018): AI adoption can progress incrementally as organisations learn which forms of intelligence can reliably be supported or substituted by technology.

Consequently, prioritisation should consider not only the stand-alone attractiveness of a use case but also its contribution to organisational learning. A relatively modest AI assistant may create valuable knowledge about data quality, employee adoption, model monitoring, governance, and customer acceptance that enables more ambitious applications later.

12.8 The Principle of Risk-Adjusted Value

The central principle emerging from the model is therefore:

Prioritise use cases that combine high value, high feasibility, strong data readiness, appropriate human and customer fit, and manageable risk—not those that simply demonstrate the most advanced AI technology.

This principle addresses a common strategic failure mode in AI programmes: selecting applications because they are technologically impressive rather than because they solve important organisational problems. Generative AI, conversational agents, and autonomous workflows can appear compelling in demonstrations, but their strategic value depends on the context in which they are deployed.

A use case involving a large volume of structured requests, reliable data, established workflows, measurable customer pain, strong employee acceptance, and relatively contained risk may therefore deserve priority over a more sophisticated application involving fewer interactions, uncertain data, significant regulatory exposure, and unclear customer value.

The prioritisation model consequently provides the bridge between the seven-dimensional evaluation framework and practical investment decisions. The seven dimensions establish what must be understood about a use case; the prioritisation model provides a structured mechanism for deciding which opportunities warrant experimentation, pilot deployment, scaling, or rejection.

Ultimately, the objective is not to maximise the number of customer-service interactions handled by AI. It is to maximise the risk-adjusted organisational and customer value generated by AI-enabled service, while preserving appropriate human judgement, customer trust, and regulatory accountability.

13. A Four-Quadrant Portfolio

The prioritisation model developed in Chapter 12 can be complemented by a simpler portfolio perspective that maps AI use cases according to two fundamental dimensions: expected business and customer value and implementation feasibility. While the weighted scoring model provides a more comprehensive assessment by incorporating data readiness, organisational compatibility, and risk, a two-dimensional portfolio offers a useful managerial visualisation of where different initiatives should sit within an organisation's broader AI programme.

The four-quadrant model should therefore not replace the seven-dimensional assessment. Rather, it provides a portfolio-level representation of its results, helping decision-makers distinguish between use cases that should be implemented immediately, developed strategically, pursued opportunistically, or rejected.

13.1 High Value and High Feasibility: Immediate Priorities

The first quadrant contains use cases that combine high expected value with relatively high implementation feasibility. These should generally constitute the first wave of an organisation's AI customer-service programme because they offer the strongest combination of business relevance and execution probability.

Typical examples include automated intent classification, email triage, knowledge retrieval, response drafting, routine status requests, and agent-assist summarisation. These applications often operate on comparatively structured tasks and can be implemented as either automation or employee augmentation. They also provide relatively clear opportunities for measurement through indicators such as handling time, response time, accuracy, containment, first-contact resolution, and employee productivity.

The academic literature provides support for prioritising augmentation-oriented applications within this quadrant. Brynjolfsson et al. (2025), for example, found that generative-AI assistance increased customer-support productivity by approximately 15% on average, with particularly strong benefits for less experienced employees. This suggests that organisations do not necessarily need to begin with fully autonomous customer-service processes to realise significant value. Agent-assist applications can create measurable productivity and capability benefits while retaining human responsibility for the final customer interaction.

Similarly, the task-level perspective of Huang and Rust (2018) suggests that activities such as information retrieval, classification, and other relatively mechanical or analytical tasks are natural candidates for early AI adoption. These characteristics make such applications particularly suitable for controlled pilots and subsequent scaling.

The strategic objective in this quadrant should therefore be rapid but controlled value realisation. Organisations should use these initiatives to establish operational experience, develop AI governance capabilities, generate evidence about customer and employee acceptance, and create reusable technical infrastructure.

13.2 High Value and Low Feasibility: Strategic Development

The second quadrant contains use cases with high potential value but substantial implementation barriers. These initiatives should not necessarily be rejected. Instead, they should be treated as strategic development opportunities requiring investment in data, integration, organisational capabilities, or governance before they can be deployed responsibly.

Examples include complex voicebots capable of handling multi-step interactions, autonomous cross-system resolution, automated claims-related decisions, and highly personalised next-best-action systems. Such applications may affect a larger portion of the customer journey and therefore offer substantial potential value, but they also require significantly more sophisticated data, system integration, exception handling, monitoring, and governance.

The distinction between value and feasibility is particularly important in regulated industries. A use case may promise substantial savings or service improvements but remain unsuitable for immediate deployment because the organisation lacks the necessary APIs, sufficiently reliable data, monitoring capabilities, or governance processes. In such circumstances, the appropriate response is not necessarily to abandon the opportunity but to identify the capability gaps that prevent implementation.

For example, an organisation may identify autonomous resolution across CRM, billing, and workflow systems as a high-value opportunity but discover that its systems cannot currently exchange the information required to complete the process reliably. The AI initiative can then become part of a broader digital-transformation programme focused on API enablement, data quality, workflow standardisation, and governance.

This quadrant therefore represents a strategic development pipeline. Use cases should progress toward the high-value/high-feasibility quadrant as the organisation removes the constraints identified during assessment.

13.3 Low Value and High Feasibility: Opportunistic Automation

The third quadrant consists of applications that are relatively easy to implement but offer limited strategic value. These initiatives can be attractive because they can often be deployed quickly and with comparatively little technical effort. However, their ease of implementation should not be confused with strategic importance.

Examples might include low-volume informational automations, minor administrative tasks, or narrowly defined service functions that require little integration but have limited effect on customer outcomes, operating costs, or employee capacity.

Such applications can nevertheless be worthwhile when implementation costs are negligible, when they provide useful learning, or when they contribute to a broader platform capability. They may also serve as low-risk experiments through which employees become familiar with AI-enabled workflows.

The appropriate management principle is therefore opportunistic rather than strategic investment. These initiatives should not consume disproportionate management attention or scarce technical resources simply because they are easy to implement. Where a higher-value application is blocked by limited resources, an organisation should generally prefer developing the capability required for the more valuable use case.

This reflects the broader argument of the paper that technological feasibility alone is insufficient as a basis for prioritisation. The purpose of AI adoption is not to maximise the number of tasks automated but to create sustainable customer and organisational value.

13.4 Low Value and Low Feasibility: Reject or Defer

The fourth quadrant contains use cases characterised by low expected value and low implementation feasibility. These applications should normally be rejected or deferred because neither the expected benefits nor the probability of successful implementation provides a compelling justification for investment.

Such initiatives may require substantial data preparation, complex integration, extensive organisational change, or significant governance controls while affecting relatively few customers or generating limited operational benefit. Even if technically interesting, they are unlikely to justify the associated opportunity costs and risk.

Deferral should nevertheless be distinguished from permanent rejection. Changes in technology, customer behaviour, data availability, regulation, or organisational infrastructure may alter the feasibility or value of a use case over time. A use case placed in Quadrant IV today could move into another quadrant if the underlying constraints change. Periodic portfolio reassessment is therefore appropriate, particularly as AI capabilities and enterprise technology architectures evolve.

13.5 The Portfolio as a Dynamic Management Tool

The four-quadrant model is most useful when treated as a dynamic portfolio rather than a one-time classification exercise. Use cases can move between quadrants as organisations improve data quality, establish integrations, develop employee capabilities, address regulatory requirements, or learn more about customer acceptance.

For example, a high-value voicebot may initially be classified as high value/low feasibility because speech recognition, system integration, escalation, and compliance requirements are insufficiently mature. Following investment in telephony integration, knowledge management, monitoring, and human escalation, the same use case may become high value/high feasibility.

Similarly, an application initially classified as high feasibility may be downgraded if a pilot demonstrates that customers strongly prefer human interaction for the relevant type of request. This illustrates why feasibility should encompass customer and organisational feasibility, not merely technical implementation.

The portfolio should therefore be reviewed using evidence from pilots and operational deployments. Metrics such as customer satisfaction, resolution rates, AI accuracy, employee productivity, exception rates, and risk incidents can provide new evidence that changes a use case's position.

13.6 Relationship to the Seven-Dimensional Framework

The four-quadrant model provides a useful simplification, but it should not obscure the broader seven-dimensional assessment. In particular, risk should not be treated as merely another component of feasibility.

A high-value/high-feasibility use case may still be unsuitable if it creates unacceptable legal, ethical, security, or operational risks. For this reason, the portfolio should be subject to governance gates before a use case is classified as an immediate priority. Similarly, a use case should not be considered genuinely feasible if its data cannot lawfully be processed or if the organisation lacks the ability to monitor and control its operation.

The relationship between the models can therefore be expressed conceptually as follows. The seven-dimensional framework determines whether a use case is viable and why. The prioritisation model determines its relative attractiveness. The four-quadrant portfolio determines how the organisation should manage it over time.

This creates a coherent progression from analysis to action. An organisation first defines the business problem and use case, evaluates data and integration requirements, determines appropriate human–AI roles, establishes customer and operational metrics, develops the business case, and assesses governance and risk. It can then score and rank candidate applications and finally position them within a portfolio according to value and feasibility.

13.7 Strategic Implications

The central strategic implication of the four-quadrant model is that AI investment should be sequenced rather than pursued indiscriminately. Organisations should generally begin with high-value/high-feasibility applications because these offer the greatest opportunity to generate evidence and benefits with comparatively manageable implementation risk. High-value/low-feasibility initiatives should form the strategic development pipeline, with targeted investment aimed at removing the barriers that currently prevent deployment. Low-value/high-feasibility opportunities can be pursued selectively when costs are low or learning benefits are significant, while low-value/low-feasibility applications should normally be deferred or rejected.

This sequencing also supports a more realistic approach to AI transformation. Organisations do not need to choose between maintaining the existing human service model and pursuing complete automation. Instead, they can develop a portfolio in which relatively low-risk augmentation applications create immediate value while more ambitious forms of automation are developed progressively as technological, organisational, and governance capabilities mature.

The resulting portfolio is therefore not simply a method for categorising AI ideas. It is a mechanism for allocating scarce organisational resources to the customer-service applications most likely to generate sustainable, risk-adjusted value. This reinforces the central argument of the paper: the most important question is not where AI can technically be deployed, but where AI can be deployed in a way that simultaneously improves customer outcomes, strengthens employee capability, creates economic value, and remains operationally and regulatorily defensible.

14. Implementation Roadmap

The preceding framework provides a basis for identifying and prioritising AI use cases, but successful adoption also requires a structured approach to implementation. Customer-service AI should not be treated as a conventional software deployment in which a technology is selected, configured, and subsequently rolled out across the organisation. The evidence reviewed in this paper indicates that AI outcomes depend on the interaction between model capability, data quality, workflow design, employee behaviour, customer acceptance, system integration, and governance (Huang & Rust, 2018; Xiao & Kumar, 2021; Brynjolfsson et al., 2025). Implementation should therefore be conceived as an iterative organisational learning process rather than a one-time technology project.

A suitable implementation programme can be organised into six interconnected phases: baseline, use-case discovery, evaluation, pilot, scale, and governance-driven continuous improvement. The phases are sequential in their logic but should not be understood as strictly linear. Evidence generated during piloting may change the assessment of a use case, reveal previously unidentified data requirements, or demonstrate that the intended human–AI division of labour needs to be redesigned. The organisation should consequently be prepared to move backwards and forwards between phases as new evidence emerges.

14.1 Phase 1: Establish the Baseline

Implementation should begin with a detailed understanding of the existing customer-service operation. The purpose is not simply to document current processes but to establish the quantitative and qualitative baseline against which AI-related improvements can be assessed.

The organisation should map relevant customer journeys and service processes, including channels, request categories, volumes, handling times, response and resolution times, first-contact resolution, escalation rates, abandonment, customer satisfaction, staffing requirements, peak-demand patterns, and error rates. Where possible, these measures should be disaggregated by request type and channel because aggregate averages can conceal substantial differences between individual processes.

Process mapping should also identify the tasks performed within each interaction. This is particularly important given the task-level perspective of Huang and Rust (2018). A process such as handling a customer email may involve classification, authentication, information retrieval, policy interpretation, response generation, transaction execution, documentation, and escalation. These tasks may have very different levels of suitability for automation or augmentation.

The baseline should therefore provide two forms of evidence: what the process currently costs and how well it performs, and what employees actually do to deliver the service. Without this information, organisations risk defining AI success in terms of technical metrics rather than meaningful business outcomes.

14.2 Phase 2: Use-Case Discovery

Once the baseline has been established, organisations should identify a broad portfolio of potential AI applications. A useful initial target is approximately 10–20 candidate use cases, although the appropriate number will depend on the size and complexity of the service organisation.

At this stage, the organisation should deliberately avoid constraining the list by a predetermined technology. The objective is to identify business problems and potentially valuable tasks first, and only subsequently determine whether conversational AI, generative AI, machine learning, automation, retrieval systems, voice technology, or another approach is appropriate.

Candidate use cases should be generated from multiple sources, including customer pain points, employee pain points, high-volume processes, long response times, repetitive workflows, error-prone activities, knowledge-intensive tasks, peak-demand constraints, and opportunities for employee augmentation. Frontline employees should play an important role in this discovery process because they possess detailed knowledge of exceptions, informal workarounds, and sources of customer frustration that may not be visible in management-level process data.

The resulting list should include both automation and augmentation opportunities. This is important because the evidence from Brynjolfsson et al. (2025) indicates that AI assistance can generate substantial productivity improvements without requiring the complete automation of customer-service work. Similarly, Huang and Rust (2018) suggest that AI adoption should be considered at the level of individual tasks rather than entire occupations.

14.3 Phase 3: Structured Use-Case Evaluation

The candidate portfolio should then be evaluated using the seven-dimensional framework developed in this paper. Each use case should be assessed in terms of its strategic relevance, task and channel characteristics, data and system readiness, human–AI role allocation, customer and operational outcomes, economic value, and governance and risk.

The prioritisation model introduced in Chapter 12 can provide a structured scoring mechanism, while the four-quadrant portfolio in Chapter 13 can provide a visual representation of relative value and feasibility. Importantly, these tools should support rather than replace managerial judgement.

Evaluation should include explicit assumptions and evidence. For example, a projected 50% automation rate should be based on the proportion of requests that genuinely meet the defined automation criteria, rather than on a generic vendor benchmark. Similarly, expected labour savings should distinguish between actual headcount reduction and capacity released through productivity improvements.

Governance should be assessed at this stage rather than after the preferred use case has already been selected. GDPR requirements, AI Act applicability, data location, security, human oversight, transparency, auditability, and escalation mechanisms should be incorporated into the evaluation. This prevents organisations from investing heavily in technically promising applications that subsequently prove difficult or impossible to deploy within the required regulatory and operational boundaries.

14.4 Phase 4: Controlled Pilot

The next phase should involve selecting one or two use cases that combine high expected value with manageable implementation complexity and risk. The objective of the pilot is not merely to demonstrate that the AI system works technically. It is to determine whether the complete AI-enabled service process produces measurable improvement under realistic operating conditions.

Where feasible, organisations should establish a control group, comparison group, or robust pre/post baseline. The appropriate experimental design will depend on the process and operational constraints. Relevant outcomes should include customer metrics such as CSAT, CES, complaint rates, escalation, and resolution; operational metrics such as AHT, FCR, response time, and containment; AI-quality measures such as accuracy, hallucination, retrieval quality, and inappropriate-response rates; and employee outcomes such as productivity, workload, training requirements, and satisfaction.

The importance of measuring employee outcomes is supported by Brynjolfsson et al. (2025), whose findings demonstrate heterogeneous effects of AI assistance across employee experience levels. An average productivity improvement can therefore conceal important differences between employee groups. Pilot evaluation should examine not only whether AI improves aggregate performance but who benefits, under what conditions, and whether quality changes alongside productivity.

Customer acceptance should also be assessed. Research on chatbot service demonstrates that customers respond to information and service quality, usefulness, problem-solving capability, and the degree of human interaction (Ashfaq et al., 2020; Zhao et al., 2022; Xie et al., 2024). A technically successful pilot that produces poor customer outcomes should therefore not proceed automatically to scale.

The pilot should also test exception and escalation behaviour. The critical question is not simply whether AI can handle typical requests, but whether it reliably recognises when it should not act autonomously.

14.5 Phase 5: Scale and Integrate

A successful pilot should not be interpreted as proof that the application can immediately be deployed across the entire organisation. Scaling introduces new requirements concerning system integration, process standardisation, capacity, monitoring, security, and organisational change.

At this stage, the AI application should be integrated with the systems required to support the complete customer-service workflow. Depending on the use case, these may include CRM platforms, ticketing systems, knowledge-management systems, telephony infrastructure, identity and authentication services, workflow engines, and RPA interfaces.

The integration objective should be to eliminate unnecessary manual handoffs while preserving appropriate control points. An AI system that generates an accurate response but requires employees to manually copy information between several systems may produce less value than expected. Conversely, excessive automation can increase operational risk if the system is given authority to execute actions without adequate validation.

Scaling should therefore proceed through controlled increments. Organisations can initially expand the volume of interactions, customer segments, channels, or process variants while monitoring whether performance remains within validated boundaries. New use cases should not automatically inherit the risk assumptions of the original deployment because different tasks may involve different data, customer expectations, and regulatory consequences.

14.6 Phase 6: Governance and Continuous Improvement

AI deployment should be treated as the beginning of an ongoing management process rather than the end of the implementation programme. Customer behaviour, product information, policies, knowledge bases, model capabilities, system configurations, and regulatory requirements can all change over time.

Continuous monitoring should therefore cover at least five areas: AI performance, customer outcomes, employee outcomes, operational reliability, and regulatory compliance.

AI monitoring should assess whether accuracy, retrieval quality, hallucination rates, escalation performance, and other relevant quality indicators remain within acceptable limits. Customer monitoring should examine whether service satisfaction, effort, complaints, and resolution outcomes remain at or above the intended baseline. Employee monitoring should consider productivity, workload, adoption, training requirements, and the distribution of benefits across employee groups.

Governance monitoring should establish whether the system continues to operate within its approved scope and whether material changes require reassessment. Model updates, changes in underlying knowledge sources, new integrations, or changes in the business process can alter system behaviour and should therefore be subject to appropriate change-management controls.

Auditability is particularly important at this stage. As established in Chapter 10, organisations should be able to reconstruct material AI-supported actions sufficiently to investigate failures, demonstrate accountability, and support continuous improvement. Logging should therefore be integrated into the operating model rather than added retrospectively.

14.7 The Importance of Iteration

The six-phase approach should ultimately be understood as an iterative learning cycle. The initial baseline provides assumptions about the process; use-case discovery converts these observations into candidate interventions; structured evaluation identifies the most promising opportunities; pilots generate empirical evidence; scaling tests whether that evidence generalises; and continuous governance determines whether the resulting system remains effective and acceptable over time.

This iterative approach is preferable to a "big bang" deployment for several reasons. First, AI performance is dependent on the quality and structure of the underlying data. Second, integration with enterprise systems can reveal constraints that are not visible in isolated prototypes. Third, employee adoption and customer acceptance cannot always be predicted reliably before real-world use. Fourth, pilot results may reveal that an apparently suitable automation task is better implemented as employee augmentation. Finally, regulatory and governance requirements may change as the use case becomes more autonomous or consequential.

The evidence from Brynjolfsson et al. (2025) illustrates why empirical learning is particularly valuable. Their study found substantial average productivity gains but also significant heterogeneity across workers. Such findings suggest that organisations should avoid assuming that a single AI intervention will affect all employees or customer interactions uniformly.

14.8 Implementation as Organisational Transformation

The implementation roadmap ultimately reinforces a central conclusion of this paper: customer-service AI is an organisational transformation problem, not merely a technology implementation problem.

Successful deployment requires the alignment of data, technology, workflows, employees, customers, management processes, and governance. The role of AI may also evolve over time. An organisation may begin with employee-facing knowledge retrieval and response drafting, progress to semi-automated customer interactions, and eventually introduce autonomous execution for narrowly defined, well-controlled processes. Each stage generates experience and infrastructure that can support the next.

The objective should therefore not be to automate as much customer service as technically possible as quickly as possible. Rather, organisations should develop an evidence-based pathway in which each subsequent level of AI autonomy is justified by demonstrated value, validated performance, appropriate human oversight, and acceptable risk.

The resulting implementation logic can be summarised as:

Baseline the current service → discover candidate tasks → evaluate value and risk → pilot under controlled conditions → scale proven applications → continuously monitor, govern, and improve.

This approach translates the Deepsearch (2026) use-case methodology into an implementation process while incorporating the academic evidence on task-level AI adoption (Huang & Rust, 2018), human–AI complementarity (Xiao & Kumar, 2021), employee productivity and heterogeneous AI effects (Brynjolfsson et al., 2025), and customer acceptance and service quality (Ashfaq et al., 2020; Zhao et al., 2022; Xie et al., 2024). It thereby provides a practical pathway from AI opportunity identification to responsible, measurable, and scalable customer-service transformation.

15. Vendor Evaluation

The use-case framework developed in this paper should extend directly into the vendor-selection and procurement process. A recurring risk in customer-service AI programmes is that organisations begin procurement by comparing vendors according to product features, benchmark claims, or demonstrations rather than first establishing whether a particular solution can reliably address the organisation's specific use case. This reverses the logic developed throughout the paper. Vendor selection should follow use-case definition and evaluation, not substitute for them.

The appropriate procurement question is therefore not:

"Which AI vendor has the best chatbot?"

but:

"Which solution can demonstrably deliver the required customer-service outcome, within our data, technology, organisational, economic, and regulatory constraints?"

This distinction is particularly important because AI performance is highly context-dependent. A vendor may report impressive aggregate accuracy while achieving substantially different results on an organisation's specific intents, languages, products, policies, customer segments, or exception patterns. The findings of Brynjolfsson et al. (2025) further demonstrate that AI effects can vary across types of customer-support work and employee experience. Vendor evaluation should therefore focus on evidence generated under conditions that approximate the intended deployment environment.

15.1 From Vendor Features to Use-Case Evidence

The Deepsearch (2026) approach provides a useful foundation for procurement because the AI Use Case Canvas can be translated into a structured set of vendor requirements and evaluation criteria. The vendor should be assessed against the actual process, data, integrations, performance targets, human roles, and governance requirements established during use-case analysis.

For example, rather than asking whether a vendor offers "intent recognition," the organisation should test how accurately the system recognises the organisation's own customer intents, including ambiguous, multilingual, incomplete, and historically difficult cases. Similarly, rather than accepting a generic automation-rate claim, the organisation should determine what proportion of its own interactions can be safely completed without human intervention and under what conditions escalation is required.

This approach turns procurement into an extension of the experimental logic developed in Chapter 14. Vendors should be required, wherever feasible, to demonstrate performance using representative organisational data or carefully controlled evaluation datasets, subject to applicable privacy and security requirements.

15.2 Accuracy and Quality

Technical accuracy should be evaluated at the level of the intended use case. Relevant measures may include intent-recognition accuracy, retrieval accuracy, answer correctness, inappropriate-response rates, hallucination rates, escalation precision, and successful task completion.

A vendor's headline accuracy figure is meaningful only when its measurement methodology is transparent. Organisations should establish what dataset was used, how the dataset was constructed, which intents were represented, how ambiguous cases were treated, whether the data were representative of the target deployment, and whether the evaluation was performed by the vendor or independently validated.

The organisation should also distinguish between syntactic correctness and operational correctness. A response can be linguistically fluent while providing an incorrect policy interpretation or failing to resolve the customer's underlying problem. Research on chatbot service reinforces this distinction: customers respond strongly to information quality, service quality, usefulness, and problem-solving capability (Ashfaq et al., 2020; Zhao et al., 2022).

For this reason, vendor testing should prioritise end-to-end task success rather than language quality alone.

15.3 Uncertainty, Failure and Escalation

A particularly important procurement criterion is how the AI system behaves when it is uncertain. High-quality customer-service AI should not be evaluated only according to its ability to produce correct answers. It should also be evaluated according to its ability to recognise when it should not answer or act autonomously.

Vendor evaluation should therefore investigate how uncertainty is represented, how confidence thresholds are established, and what triggers escalation. Organisations should determine whether the system can distinguish between routine cases and exceptional situations and whether escalation can be configured according to business rules, customer characteristics, risk levels, or process states.

This is closely related to the human–AI complementarity perspective developed earlier. The objective is not necessarily to maximise autonomous completion but to establish an appropriate division of labour. A system that safely resolves 60% of interactions and correctly escalates the remainder may be substantially more valuable than a system that claims 80% automation but produces inappropriate responses in the residual cases.

Vendor demonstrations should therefore include failure testing, including ambiguous requests, incomplete information, contradictory information, unusual cases, adversarial inputs, and situations where the correct response is to transfer the interaction to a human employee.

15.4 Data Architecture and Integration

The vendor must also demonstrate compatibility with the organisation's existing data and technology environment. The assessment should cover APIs, authentication, CRM integration, ticketing, knowledge-management systems, telephony, workflow engines, RPA, identity management, monitoring infrastructure, and relevant enterprise architecture requirements.

This criterion follows directly from the data and integration dimension of the Use Case Canvas. An AI model can perform exceptionally well in a standalone environment while creating limited business value if it cannot access the information or systems necessary to complete the customer's request.

The organisation should therefore ask not simply whether an integration "exists" but what level of functionality the integration provides. Read-only access to a CRM system may support knowledge retrieval but may not enable the AI to execute a customer transaction. Similarly, an API may technically exist while imposing latency, throughput, authentication, or availability constraints that make the intended service workflow impractical.

Integration requirements should consequently be tested against the complete process architecture established during use-case evaluation.

15.5 Data Location, Privacy and Security

Where customer-service systems process personal information, vendors should provide sufficiently detailed information about where data are processed, stored, transferred, and retained. Organisations should establish which subprocessors are involved, whether data leave the European Economic Area where this matters to the organisation's legal and risk assessment, what security measures apply, and whether customer data can be used for model training or other secondary purposes.

These questions should be assessed against the organisation's GDPR obligations and internal data-governance requirements rather than treated as generic vendor-compliance claims. Relevant considerations include data minimisation, purpose limitation, access controls, retention, deletion, encryption, authentication, logging, and support for data-subject rights.

Where an organisation requires data to remain within the EU for legal, contractual, or risk-management reasons, the vendor should be required to provide sufficiently specific evidence concerning the relevant processing locations and data flows. Statements such as "EU compliant" should not substitute for an examination of the actual architecture and contractual arrangements.

15.6 Model Updates and Change Management

AI systems are dynamic technologies. Models, retrieval mechanisms, prompts, safety controls, knowledge sources, and other system components may change over time. Consequently, vendor evaluation should establish how changes are governed and how organisations are informed about material changes.

Important questions include how model versions are documented, whether customers can control the timing of significant updates, how changes are tested, whether regression testing is available, how performance changes are communicated, and whether the organisation can determine retrospectively which model or configuration generated a particular output.

This requirement follows directly from the auditability principles developed in Chapter 10. If the organisation cannot identify which version of an AI system produced a material customer-facing decision or response, its ability to investigate incidents and demonstrate accountability may be significantly constrained.

Vendor contracts should therefore address change notification, version management, incident response, service-level requirements, and appropriate access to operational records.

15.7 Auditability and Human Override

Auditability should be treated as a fundamental procurement criterion, particularly in regulated environments. Organisations should establish whether the vendor can provide appropriate records of AI inputs, outputs, relevant retrieval sources, system and model versions, human interventions, escalations, and final actions.

The organisation should also examine the human-override mechanisms available. Human intervention should not be limited to a generic "contact an agent" function. For higher-impact processes, employees may need the ability to stop an automated action, modify a recommendation, reject an AI-generated response, or take over the entire interaction.

The design of these mechanisms should correspond to the role assigned to AI in the use case. An employee-facing drafting assistant may require relatively simple acceptance and editing controls, whereas an autonomous transaction system may require explicit approval gates, role-based permissions, and the ability to reverse actions.

The procurement process should therefore evaluate whether human oversight is technically meaningful and operationally usable, rather than merely present in product documentation.

15.8 GDPR and AI Act Support

Vendor evaluation should also establish how the supplier supports the organisation in meeting applicable regulatory requirements. For GDPR, this may include documentation concerning data processing, security, subprocessors, data-subject rights, retention, and relevant contractual arrangements.

For the EU AI Act, the organisation should determine which regulatory obligations apply to the specific AI system and use case and what responsibilities are allocated between provider and deployer. This distinction is important because compliance responsibilities cannot necessarily be transferred to the vendor simply through procurement. The organisation deploying an AI system remains responsible for managing its own obligations within the applicable regulatory framework.

For customer-facing systems, the vendor should be able to support applicable transparency requirements, including appropriate disclosure where customers are directly interacting with AI. For systems that fall within higher-risk categories, additional requirements concerning transparency, human oversight, record-keeping, risk management, and other controls may apply.

The relevant procurement question is therefore not simply whether a vendor claims to be "AI Act compliant," but which obligations apply to the proposed system, which party is responsible for satisfying each obligation, and what technical and contractual mechanisms support compliance.

15.9 Data and Knowledge Portability

Another important procurement consideration is portability. Organisations should avoid unnecessary dependence on a vendor if the AI solution becomes deeply embedded in customer-service operations.

The evaluation should establish whether the organisation can export its customer-service knowledge, configurations, evaluation datasets where legally permissible, interaction records, prompts or system instructions where applicable, analytics, and other relevant assets. It should also establish what happens to organisational data and knowledge when the contract terminates.

Portability is strategically important because AI architectures and vendors are evolving rapidly. A use case that is appropriate for one technology today may be better served by another technology in the future. Maintaining control over organisational knowledge and operational data therefore reduces switching costs and strengthens the organisation's negotiating position.

15.10 Total Cost of Ownership

Vendor comparison should also move beyond headline licence or subscription prices. The relevant economic measure is the total cost of ownership over the expected life of the use case.

This should include software or usage fees, implementation, integration, data preparation, customisation, model evaluation, security reviews, training, employee change management, monitoring, support, infrastructure, ongoing maintenance, compliance activities, and eventual migration or termination costs.

The business-case logic developed in Chapter 8 should be applied consistently to vendor selection. A solution with a higher nominal licence cost may produce greater overall value if it achieves substantially better task performance and requires less manual review. Conversely, a low-cost model may prove expensive if it generates high exception rates, requires extensive employee intervention, or creates substantial integration and governance costs.

The relevant question is therefore cost per successfully resolved customer need, not simply cost per AI interaction.

15.11 A Structured Procurement Process

The vendor evaluation process can consequently be structured around the same evidence hierarchy used throughout this paper. First, the organisation defines the use case and its required outcomes. Second, it establishes acceptance criteria for accuracy, customer experience, employee impact, integration, governance, and economics. Third, vendors are evaluated against these criteria using comparable evidence. Fourth, shortlisted solutions are tested using representative scenarios and, where possible, organisational data. Finally, the preferred solution is validated through a controlled pilot before large-scale contractual and operational commitment.

This process reduces the risk of vendor-led problem definition, in which a compelling product demonstration creates demand for a use case that was never independently established as strategically valuable.

It also allows organisations to distinguish between three different types of vendor claim:

Capability claims concern what the technology can theoretically do.

Performance claims concern how well it performs on specified evaluation tasks.

Business-value claims concern the operational and financial outcomes that customers can realistically achieve.

Only the latter two should materially influence investment decisions, and both should be validated in the organisation's own context.

15.12 From Feature Comparison to Evidence-Based Procurement

The central implication is that vendor evaluation should become an extension of the AI Use Case Canvas and prioritisation framework, rather than a separate procurement exercise.

A vendor should be able to provide credible evidence concerning performance on the organisation's own intents, hallucination and incorrect-response rates, benchmark methodology, data and system architecture, processing locations, data-use practices, model-update procedures, auditability, human override, uncertainty handling, APIs, regulatory support, failure management, data portability, and complete implementation and operating costs.

These criteria also provide a basis for competitive comparison. Rather than comparing vendors primarily on the number of features they advertise, organisations can compare them on their ability to satisfy predefined use-case requirements under realistic operating conditions.

The resulting procurement principle can be stated simply:

Select the vendor that provides the strongest evidence of achieving the required customer-service outcome within the organisation's technical, economic, organisational, privacy, regulatory, and governance constraints—not the vendor with the most impressive product demonstration.

This approach aligns procurement with the broader argument of the paper. AI adoption should begin with the business problem, proceed through task-level use-case evaluation, incorporate customer and employee outcomes, establish the economic case, and integrate governance from the outset. Vendor selection then becomes the mechanism for finding the technology partner capable of delivering that validated use case—not the starting point for deciding what the organisation should automate.

16. Discussion

The central finding of this paper is that the adoption of AI in customer service should not be conceptualised primarily as an automation programme. It is better understood as a service-system redesign problem in which organisations reconsider how customer needs are identified, resolved, escalated, and supported through a combination of AI systems and human capabilities. The relevant strategic question is therefore not how many customer-service tasks can be automated, but how AI can be allocated across tasks and interactions in ways that improve customer outcomes, employee performance, organisational economics, and regulatory robustness simultaneously.

The Deepsearch (2026) framework provides a valuable practical starting point because it places the use case rather than the technology at the centre of decision-making. Its emphasis on the business problem, customer-service channel, data and integration requirements, stakeholders, performance measures, and business case provides a pragmatic structure for moving from broad AI interest to concrete applications. The academic literature reviewed in this paper broadly supports this orientation, but also introduces several qualifications that are important when the framework is applied in practice. In particular, the literature suggests that automation potential does not automatically translate into customer value, that human–AI complementarity is often preferable to substitution, that productivity effects are heterogeneous, that data readiness is a strategic capability, and that governance must be incorporated into use-case selection from the beginning.

16.1 Automation Potential Is Not Equivalent to Customer Value

A fundamental implication of the literature is that technological feasibility and customer value are distinct constructs. A process may be highly suitable for automation from an operational perspective while producing limited or even negative customer value.

This issue is particularly important because customers may interpret the introduction of automation as evidence that an organisation is prioritising its own cost reduction over customer welfare. Research published in the Journal of Consumer Research demonstrates that consumers can evaluate bot-provided service more negatively than equivalent human-provided service when they infer that the automation primarily benefits the firm through cost reduction. Importantly, this negative response can be attenuated or eliminated when consumers perceive that automation also creates benefits for them, such as faster or otherwise superior service.

This finding adds an important dimension to the Deepsearch framework. The business case for a use case cannot be limited to the question of how much labour cost can potentially be removed. It should also establish who receives the benefits of automation and how those benefits are experienced by the customer. Faster responses, 24/7 availability, reduced waiting, improved consistency, easier access to information, and more effective problem resolution can provide genuine customer value. By contrast, automation that primarily transfers effort from employees to customers—for example, forcing customers through an inflexible chatbot before they can reach a human—may generate an attractive internal cost calculation while simultaneously damaging customer experience.

Customer value should therefore be treated as an explicit selection criterion rather than an assumed consequence of operational efficiency. This reinforces the importance of the customer and operational metrics proposed in the extended Use Case Canvas, including CSAT, customer effort, resolution rates, escalation, abandonment, and complaint rates.

16.2 Human–AI Complementarity Rather Than a Binary Choice

The evidence also challenges a simplistic interpretation of AI adoption as a choice between human service and automated service. Huang and Rust (2018) conceptualise AI adoption at the task level, distinguishing different forms of intelligence and suggesting that the potential for substitution varies across tasks. This provides a theoretical basis for decomposing customer-service processes into individual activities rather than treating the employee's entire role as the unit of automation.

The customer-service literature similarly indicates that AI is particularly suitable for structured, information-intensive, and transactional activities, whereas human capabilities remain important where interactions require empathy, discretion, negotiation, contextual interpretation, or relationship management (Xiao & Kumar, 2021; Zhang et al., 2026). Chatbot research reinforces this distinction: customers value usefulness, information quality, service quality, and effective problem resolution, but negative reactions can emerge when AI is unable to solve problems or when customers perceive insufficient access to human assistance (Ashfaq et al., 2020; Zhao et al., 2022).

The strategic implication is that organisations should design deliberate human–AI role allocations. AI can classify interactions, retrieve information, summarise cases, draft responses, recommend actions, or autonomously resolve narrowly defined routine requests. Employees can remain responsible for exceptions, emotionally sensitive interactions, complex decisions, and situations where customer trust or discretion is particularly important.

This perspective also changes how AI success should be evaluated. The objective should not necessarily be maximum automation or maximum containment. Instead, organisations should seek an optimal allocation of tasks between AI and employees, taking into account the complexity and risk of the task, customer expectations, the cost of errors, and the comparative strengths of humans and AI.

16.3 Productivity Effects Are Heterogeneous

The findings of Brynjolfsson, Li and Raymond (2025) provide particularly important empirical support for this complementary perspective. Their study of 5,172 customer-support agents found that access to a generative-AI conversational assistant increased issues resolved per hour by approximately 15% on average. However, the average effect conceals substantial heterogeneity: less experienced and lower-skilled workers benefited considerably more, while highly experienced workers experienced smaller gains and, in some cases, small reductions in quality.

The implication is that the widely cited 15% figure should not be treated as a universal productivity benchmark. It represents an observed effect within a particular organisation, occupation, technology, employee population, and operating environment. Differences in task structure, AI capability, data quality, employee skills, process design, and implementation quality can produce substantially different outcomes elsewhere.

At the same time, the study provides a more important strategic insight than the headline percentage. AI can function as a mechanism for knowledge diffusion and capability enhancement, reducing performance gaps between employees rather than simply eliminating work. Less experienced employees may benefit because the AI system provides access to information and response patterns that previously depended on experience.

This finding strengthens the case for including employee outcomes within the AI business case. Organisations should consider not only labour-hour reduction but also faster onboarding, reduced training requirements, increased capacity, improved consistency, and the ability to redeploy employees towards activities requiring greater judgement and interpersonal capability.

It also suggests that implementation should examine distributional effects. An AI system that increases average productivity but disproportionately increases workload for particular employees, reduces quality among experienced staff, or creates new forms of monitoring pressure may not produce the organisational benefits implied by aggregate productivity measures.

16.4 Data Readiness as a Strategic Capability

Another important finding is that data availability should be treated as a strategic capability rather than merely a technical prerequisite. The existence of large quantities of historical customer-service data does not guarantee that an AI application will perform effectively. Data must be sufficiently relevant, representative, accurate, accessible, current, and legally usable for the intended task.

The findings of Brynjolfsson et al. (2025) are particularly informative in this respect. The strongest productivity improvements occurred for moderately uncommon customer problems: cases for which employees had relatively limited prior experience, but for which the AI system nevertheless had sufficient training information. Highly routine cases generated smaller incremental benefits because employees already possessed substantial experience, while extremely rare problems could lack sufficient data for the AI system to provide reliable assistance.

This suggests that the relationship between task frequency and AI value is non-linear. A use case should not be prioritised simply because it represents the largest volume of customer contacts. The more relevant question is whether AI can provide capabilities that meaningfully improve the current human process.

Consequently, data readiness assessment should examine not only volume but also coverage of intents, quality of labels, representativeness of customer language, frequency of exceptions, availability of authoritative knowledge sources, and the organisation's ability to maintain the underlying information over time. In customer service, knowledge governance may therefore become as important as model selection.

This also has implications for the portfolio model developed in this paper. Some high-value applications should initially be classified as strategically important but infeasible because the organisation lacks the data or knowledge infrastructure required for reliable deployment. Investment in data quality, knowledge management, APIs, and governance can subsequently move those use cases towards implementation readiness.

16.5 Regulation as a Design Constraint and Selection Criterion

The final major implication concerns regulation and governance. The analysis in this paper suggests that compliance should not be treated as a post-selection validation step in which legal teams review an already selected technology. Regulatory requirements can fundamentally influence whether a use case is appropriate, how it should be designed, and how much value it can realistically generate.

This is particularly relevant under the EU regulatory environment. The EU AI Act establishes transparency, human oversight, documentation, and other obligations depending on the nature and risk classification of the AI system. For customer-facing applications, transparency concerning interaction with AI is particularly relevant, while higher-risk applications may involve substantially more extensive governance requirements. GDPR requirements simultaneously influence what customer data can be processed, for what purpose, under which legal basis, and with what safeguards.

The consequence is that risk and governance can change the economic attractiveness of a use case. A theoretically high-value application may require extensive monitoring, human oversight, documentation, security controls, or data-governance infrastructure. These requirements should be incorporated into the business case from the outset rather than treated as external constraints after the financial analysis has been completed.

This supports the seventh dimension of the extended Use Case Canvas: governance and risk should stand alongside strategy, use case, data, stakeholders, metrics, and economics. A use case should not be considered viable merely because it is technically feasible and financially attractive. It must also be legally permissible, operationally controllable, auditable, and compatible with the organisation's risk appetite.

16.6 Implications for the Extended Use Case Framework

Taken together, these findings support the extension of the Deepsearch framework proposed in this paper. The original six dimensions provide a strong managerial structure for moving from customer-service problems to AI use cases. The academic literature indicates, however, that three additional principles should guide interpretation of those dimensions.

First, value must be multidimensional. Economic value should be evaluated alongside customer value and employee value. A use case that reduces costs while increasing customer effort or employee workload may not represent genuine service-system improvement.

Second, AI capability must be evaluated relative to human capability. The appropriate question is not whether AI can perform a task, but whether AI can perform it sufficiently well and at sufficient scale to create more value than the existing human process, either independently or in combination with employees.

Third, risk-adjusted value should determine prioritisation. Technical feasibility, data readiness, and financial potential are insufficient if the organisation cannot operate the system safely and accountably.

These principles lead to a more nuanced understanding of AI customer-service transformation. The organisation is not choosing between a human operating model and an AI operating model. It is designing a hybrid service system in which different activities are allocated to humans, AI, or human–AI collaboration according to their respective capabilities and constraints.

16.7 Theoretical and Managerial Implications

Theoretically, the framework developed in this paper contributes by connecting three strands of research that are often considered separately: AI task substitution and augmentation, customer-service quality and customer acceptance, and organisational governance and economic evaluation. Combining these perspectives suggests that the appropriate unit of AI strategy is neither the technology nor the occupation, but the customer-service task embedded within a broader service process.

Managerially, this implies that organisations should resist technology-first strategies. The sequence should instead begin with the customer and business problem, identify the underlying tasks, assess the suitability of different forms of AI involvement, establish the required data and system architecture, define the human–AI division of labour, quantify customer and business outcomes, and incorporate governance before committing to deployment.

The resulting decision logic can be expressed as a series of questions:

Does the use case create genuine customer or organisational value?

Can AI perform the relevant task reliably enough to improve the existing process?

Is the necessary data and system infrastructure available and appropriately governed?

What should AI do, and where should humans remain responsible?

Can the resulting service be measured and audited?

Is the risk-adjusted economic value sufficient to justify implementation?

Only when these questions can be answered satisfactorily should an organisation proceed to vendor selection and deployment.

16.8 Limitations and Directions for Further Research

Several limitations should nevertheless be acknowledged. First, the empirical evidence underlying the framework is heterogeneous. The Brynjolfsson et al. (2025) study provides unusually strong evidence concerning generative AI in customer support, but its findings should not be assumed to generalise automatically across organisations, technologies, countries, or service contexts. Similarly, chatbot studies examine particular technologies, customer populations, and interaction settings.

Second, much of the customer-service AI literature remains focused on individual interactions rather than the long-term organisational consequences of AI deployment. Future research should examine how AI changes customer relationships, employee skill development, organisational learning, service design, and cost structures over longer periods.

Third, further empirical work is needed to identify when augmentation consistently outperforms automation and vice versa. The distinction is likely to depend on task complexity, error costs, employee expertise, customer expectations, and the degree of process standardisation.

Finally, future research should examine how AI governance requirements affect the economic value of different use cases. The costs of monitoring, auditability, human oversight, data governance, and regulatory compliance are likely to vary substantially between applications and should become more systematically incorporated into AI investment models.

16.9 Overall Assessment

Overall, the evidence supports a shift from automation-oriented AI adoption toward evidence-based service-system design. The Deepsearch (2026) framework offers a practical mechanism for structuring this process, while the academic literature demonstrates why the framework must incorporate customer value, human–AI complementarity, heterogeneous productivity effects, data readiness, and governance.

The strongest AI opportunities are therefore unlikely to be those that simply maximise automation. They are more likely to be applications in which AI addresses a clearly defined customer or operational problem, possesses access to the information required to perform the task, complements rather than unnecessarily displaces valuable human capabilities, produces measurable improvements for customers and employees, and can be governed within the organisation's legal and operational constraints.

The central conclusion can consequently be stated as follows:

The strategic value of AI in customer service depends less on how much work can be automated than on how effectively organisations redesign the service system around the complementary capabilities of AI and humans.

This perspective provides the foundation for the paper's extended seven-dimensional Use Case Canvas, prioritisation model, four-quadrant portfolio, implementation roadmap, and evidence-based vendor evaluation approach. Together, these elements transform AI use-case identification from a technology-selection exercise into a structured process of strategic, operational, economic, customer-oriented, and responsible service innovation.

17. Conclusions

This paper has argued that the central challenge of artificial intelligence in customer service is not determining whether AI can automate customer-service activities, but determining where AI should be introduced, in what role, and under what conditions it creates sustainable value. The distinction is consequential. Customer service is characterised by substantial volumes of repetitive and information-intensive work, making it an attractive domain for AI. Yet the technical potential for automation does not by itself establish customer value, employee value, economic value, or regulatory acceptability.

The analysis therefore reframes customer-service AI as a service-system redesign problem. Rather than treating the technology or the occupation as the primary unit of analysis, organisations should examine the individual tasks embedded within customer-service processes. Huang and Rust (2018) provide an important theoretical foundation for this approach by distinguishing different forms of intelligence and showing why the potential for AI substitution varies across tasks. In customer service, activities such as classification, information retrieval, summarisation, and response drafting may be highly suitable for AI, while negotiation, exceptional cases, emotional support, and relationship-sensitive interactions may benefit more from continued human involvement.

The empirical evidence further supports a shift away from a simple substitution logic. Brynjolfsson et al. (2025) demonstrate that generative-AI assistance can increase customer-support productivity, with particularly strong effects among less experienced employees. The finding is important not simply because of the reported average productivity improvement, but because it illustrates the potential for AI to augment human capability and distribute organisational knowledge. AI can reduce knowledge gaps, accelerate learning, improve consistency, and increase employee capacity without necessarily requiring a corresponding reduction in employment.

The customer perspective is equally important. Research on chatbots and AI-mediated service demonstrates that customers value usefulness, information quality, service quality, and effective problem resolution (Ashfaq et al., 2020; Xie et al., 2024). At the same time, customers can respond negatively when AI is perceived as primarily serving the organisation's cost interests or when it fails to provide effective problem resolution (Zhao et al., 2022). This implies that customer-service AI should not be designed around anthropomorphism or automation rates alone. The fundamental objective should be better service: accurate information, successful resolution, appropriate speed, transparency, consistency, and easy access to human assistance when required.

This leads to the first major conclusion of the paper: automation potential is not equivalent to customer value. A use case should be prioritised when AI improves the overall service system, not simply when AI can technically perform the underlying task. Customer benefit must therefore be incorporated explicitly into use-case evaluation alongside operational and financial considerations.

The second conclusion is that human–AI complementarity should be treated as a strategic design choice. The relevant alternatives are not simply "human" and "AI." Organisations can choose among autonomous automation, employee augmentation, and AI-based decision support, with the appropriate model depending on task complexity, customer expectations, error consequences, and regulatory requirements. This provides a more flexible foundation for AI adoption than strategies based on maximising automation or minimising headcount.

Third, data readiness is a strategic capability. Successful AI deployment depends not merely on having large quantities of historical customer data but on having relevant, reliable, representative, current, and appropriately governed information. The evidence that AI benefits can vary substantially across different types of customer problems reinforces the importance of evaluating data readiness at the level of the individual use case. Organisations should consequently invest not only in models but also in knowledge management, data quality, system integration, APIs, and information governance.

Fourth, economic value should be understood as more than direct labour-cost reduction. AI can generate value through reduced handling time, increased capacity, faster response, improved employee productivity, reduced training requirements, higher service volumes, and avoidance of future recruitment. The distinction between labour savings and capacity release is therefore essential. An organisation that maintains its workforce after introducing AI may still realise substantial economic value if additional capacity enables higher service quality, growth, or redeployment towards higher-value activities.

Fifth, governance and risk must be integrated into use-case selection from the beginning. The extension of the Deepsearch framework from six to seven dimensions reflects this conclusion. A technically feasible and financially attractive use case may nevertheless be inappropriate because of privacy, security, regulatory, ethical, or operational risks. GDPR requirements and applicable EU AI Act obligations should therefore be considered during use-case definition and evaluation rather than after vendor selection. Transparency, human oversight, logging, auditability, data governance, and appropriate escalation mechanisms should form part of the initial system design.

On this basis, the paper extends the Deepsearch (2026) AI Use Case Canvas into a seven-dimensional framework comprising: strategic decision basis; use case and channel; data and system integration; human–AI roles and stakeholders; customer and operational metrics; business case; and governance and risk. The framework provides a structured mechanism for moving from an abstract AI opportunity to an operationally defined and measurable use case.

The accompanying prioritisation model and four-quadrant portfolio provide a means of translating this analysis into investment decisions. High-value, high-feasibility applications should generally form the first implementation wave, while high-value but currently infeasible applications should become part of a strategic development pipeline. Low-value applications should not consume disproportionate resources merely because they are technically easy to implement, and low-value, low-feasibility initiatives should generally be rejected or deferred. Importantly, governance should operate as a decision gate rather than merely another scoring variable: unacceptable risk cannot necessarily be compensated for by high economic value.

The implementation roadmap further suggests that organisations should proceed iteratively. Establishing a baseline, identifying candidate use cases, evaluating them systematically, conducting controlled pilots, scaling validated applications, and continuously monitoring performance provides a more robust approach than large-scale "big bang" deployment. AI performance depends on the interaction of technology with data, workflows, employees, customers, and organisational processes. Consequently, evidence from pilots should be used to revise assumptions and refine the human–AI division of labour before scaling.

The same logic should govern vendor selection. Organisations should move from feature-based procurement to evidence-based use-case validation. Vendor claims regarding accuracy, automation, security, compliance, or return on investment should be evaluated against the organisation's own intents, data, workflows, customer expectations, and regulatory requirements. The preferred vendor should therefore not necessarily be the one with the most sophisticated product or the highest headline benchmark. It should be the provider that can demonstrate the strongest evidence of delivering the required outcome within the organisation's technical, economic, customer, employee, privacy, and governance constraints.

Overall, the paper supports a fundamental shift in the way organisations should approach customer-service AI. The objective should not be to automate the maximum possible proportion of interactions. Nor should AI adoption be justified solely through projected labour savings. Instead, organisations should seek to design service systems in which AI and humans perform the tasks for which they are comparatively best suited.

The central proposition of this paper can therefore be summarised as follows:

The strategic value of AI in customer service depends less on how much work can be automated than on how effectively organisations redesign the service system around the complementary capabilities of AI and humans.

This perspective reconciles the practical value of the Deepsearch use-case methodology with the broader academic evidence on AI adoption, service quality, human–AI complementarity, customer acceptance, productivity, and responsible AI. It also provides a basis for moving beyond technology-driven experimentation towards disciplined AI investment.

For practitioners, the implication is straightforward: start with the customer and business problem, decompose the process into tasks, determine the appropriate human–AI role, validate the data and integration requirements, measure customer and employee outcomes, quantify risk-adjusted economic value, and only then select the technology and vendor. For researchers, the framework highlights the need for further empirical work on long-term human–AI complementarity, heterogeneous productivity effects, customer responses to different levels of automation, and the economic consequences of AI governance.

Ultimately, responsible customer-service AI is not defined by the absence of humans. It is defined by the quality of the allocation between human and artificial intelligence. The organisations most likely to realise sustainable value will therefore be those that treat AI not as an isolated automation technology, but as a capability for systematically redesigning how customer service is delivered.

References

Ashfaq, M., Yun, J., Yu, S., & Loureiro, S. M. C. (2020). I, Chatbot: Modeling the determinants of users' satisfaction and continuance intention of AI-powered service agents. Telematics and Informatics, 54, 101473.

Brynjolfsson, E., Li, D., & Raymond, L. R. (2025). Generative AI at Work. The Quarterly Journal of Economics, 140(2), 889–942.

Deepsearch. (2026). KI-Anwendungsfälle im Kundenservice identifizieren & bewerten: Leitfaden. Deepsearch.

European Union. (2024). Regulation (EU) 2024/1689 of the European Parliament and of the Council laying down harmonised rules on artificial intelligence (Artificial Intelligence Act). Official Journal of the European Union.

European Commission. (2026). Guidelines on transparency obligations for providers and deployers of AI systems.

Huang, M.-H., & Rust, R. T. (2018). Artificial intelligence in service. Journal of Service Research, 21(2), 155–172.

Park, Y., Kim, J., Jiang, Q., & Kim, K. H. (2024). Impact of artificial intelligence (AI) chatbot characteristics on customer experience and customer satisfaction. Journal of Global Scholars of Marketing Science, 34(3), 439–457

Shin, H. (2023). The influence of chatbot humour on consumer evaluations of services. International Journal of Consumer Studies, 47(2), 545–562

Xiao, L., & Kumar, V. (2021). Robotics for customer service: A useful complement or an ultimate substitute? Journal of Service Research, 24(1), 9–29.

Xie, C., Wang, Y., & Cheng, Y. (2024). Does artificial intelligence satisfy you? A meta-analysis of user gratification and user satisfaction with AI-powered chatbots. International Journal of Human–Computer Interaction, 40(3), 613–623.

Zhang, Y., Özsomer, A., Gürhan-Canlı, Z., & Baghirov, F. (2026). Artificial intelligence chatbots versus human agents in customer satisfaction: The role of warmth and competence. Journal of Advertising Research, 61(2)

Zhao, T., Cui, J., Hu, J., Dai, Y., & Zhou, Y. (2022). Is artificial intelligence customer service satisfactory? Insights based on microblog data and user interviews. Cyberpsychology, Behavior, and Social Networking, 25(2)

Contact

Reach out via email for inquiries.

Email

Subscribe to newsletter

info@grcadvisory.ch

© 2025. All rights reserved.