From AI Output to Human Judgement: Reframing Artificial Intelligence Governance in Banking
AI can make banking decisions faster, but only judgement-centred governance can ensure that faster decisions remain accountable, contestable and worth trusting.
Sanchez P.
9/11/2026134 min read


Abstract
Abstract
Artificial intelligence (AI) is increasingly embedded in banking, supporting credit assessment, fraud detection, financial crime monitoring, risk management and customer interaction. Existing AI governance frameworks appropriately emphasise accuracy, robustness, fairness, transparency and explainability. However, these controls do not, by themselves, ensure that AI-assisted decisions are appropriate, defensible or accountable. The central governance challenge is therefore not only whether an AI system produces reliable outputs, but whether the institution remains capable of exercising sound judgement over those outputs. This is particularly important where a formal “human in the loop” provides little substantive protection because reviewers lack the information, competence, time, independence or authority required to meaningfully assess and challenge AI recommendations.
Drawing on literature on explainable AI, algorithmic governance, accountability, automation bias and human–AI collaboration, this paper develops a Judgement-Centred AI Governance (JCAIG) framework for banking. The framework shifts the primary unit of governance from the AI model to the AI-assisted decision process and comprises six interconnected pillars: contextual risk classification; traceability and reconstructability; contestability; meaningful human override authority; independent challenge; and organisational learning. These pillars are operationalised through risk-proportionate human–AI workflows, role-specific training, decision-useful explanations, controlled experimentation, monitoring of human–AI interaction and lifecycle governance.
The paper argues that meaningful human oversight should be evaluated not by whether a human is formally present, but by whether the organisation preserves the conditions necessary for humans to understand, evaluate, question and, where appropriate, change AI-supported decisions. This reframes AI governance from a primarily model- and output-centred activity towards decision governance, in which technical performance is considered alongside human judgement, organisational capability and institutional accountability. The paper concludes that sustainable AI adoption in banking depends not on maximising either automation or human intervention, but on establishing an appropriate institutional relationship between the two: AI can inform institutional decisions, but should not silently acquire institutional authority.
Keywords: Artificial intelligence; AI governance; banking; human oversight; algorithmic accountability; explainable AI; judgement; financial regulation
1. Introduction
Artificial intelligence (AI) is increasingly embedded in the operational and decision-making infrastructure of the banking sector. Banks now employ AI and machine-learning systems across a wide range of functions, including credit assessment, fraud detection, financial crime monitoring, risk management and customer interaction. These technologies promise substantial improvements in efficiency, analytical capacity and the ability to process complex and high-volume information. However, the increasing use of AI in banking also creates a fundamental governance challenge: the technical quality or apparent plausibility of an AI-generated output does not, by itself, guarantee the quality, legitimacy or appropriateness of the decision that follows.
Much of the existing debate surrounding AI governance has focused on the characteristics of the technology itself. Accuracy, robustness, fairness, transparency and explainability are commonly treated as essential conditions for responsible AI. These concerns are particularly significant in banking, where AI-supported decisions may affect access to credit, the detection of potentially fraudulent or suspicious activity, customer treatment and institutional risk exposure. Nevertheless, a focus on technical performance alone provides an incomplete account of effective governance. Even a highly accurate or technically explainable system may be poorly governed if its outputs are accepted uncritically, interpreted incorrectly or applied without adequate consideration of context.
This limitation is particularly evident in the literature on explainable artificial intelligence (XAI). XAI is frequently presented as a means of improving transparency, trust and accountability by making the outcomes of complex AI systems more intelligible to human stakeholders. However, recent research suggests that explainability does not automatically translate into meaningful human understanding or effective organisational oversight. Tuge and Msweli (2026), in their systematic literature review of XAI in the banking sector, identify a significant gap between the technical promise of explainability and its operational implementation. Their review demonstrates that the adoption of XAI in banking is constrained not only by technical challenges, such as the trade-off between accuracy and interpretability, but also by organisational, cultural and governance barriers. These include resistance to change, misalignment between technical specialists and business units, stakeholder diversity and the absence of standardised frameworks for evaluating explanations. The authors therefore argue that XAI should not be understood as a purely technical solution but as an organisational capability requiring interdisciplinary collaboration, context-sensitive explanation design and human-centred governance (Tuge and Msweli, 2026).
This insight has important implications for AI governance more broadly. If an explanation does not ensure understanding, then the existence of a human reviewer does not necessarily ensure meaningful oversight. A decision-maker may formally review an AI-generated recommendation while lacking the knowledge, information, authority or institutional incentives required to challenge it. Human involvement may therefore become procedural rather than substantive: the human is present within the process but does not exercise genuine independent judgement. In such circumstances, the apparent presence of a “human in the loop” may provide a false sense of accountability while leaving the practical authority of the AI system largely uncontested.
The distinction between AI output and human judgement is therefore central to this paper. An AI system may generate a prediction, recommendation or explanation that appears credible and technically sophisticated, yet the institutional decision based upon that output may still be flawed. Decision quality depends not only on the properties of the AI system but also on how its outputs are interpreted, contextualised, challenged and ultimately incorporated into organisational action. This is particularly important in banking, where decisions are rarely reducible to a single prediction or score. Financial decision-making frequently involves uncertainty, competing objectives, legal obligations, institutional risk appetite and contextual factors that may not be fully captured by an algorithmic model.
The importance of this broader socio-technical perspective is supported by the growing literature on explainability in finance. Research has increasingly recognised that the value of XAI lies not simply in exposing the internal logic of complex models but in supporting appropriate human understanding and decision-making. Systematic reviews of XAI in finance have emphasised the importance of transparency and interpretability for strengthening trust, risk assessment and accountability, particularly where complex and opaque models are deployed in consequential financial decisions. However, the literature also demonstrates that the effectiveness of explainability depends on context, the needs of different stakeholders and the ability of institutions to translate technical information into meaningful organisational practice.
This paper builds on these insights while advancing a broader argument. It proposes that effective AI governance in banking should be evaluated not primarily by the quality of AI outputs, nor solely by the presence of technical safeguards, but by the quality of the human and institutional judgement surrounding AI-assisted decisions. The central concern is therefore not simply whether an AI system can produce an accurate or explainable answer. Rather, it is whether the organisation using that system remains capable of recognising when the answer should be questioned, contextualised or rejected.
This perspective develops the central argument advanced in AI Governance in Banking: Why Judgement Beats Output (Kamm, 2026), which challenges the assumption that a persuasive or apparently credible AI output constitutes evidence of good decision-making. The article highlights a growing governance risk: as AI systems become increasingly capable of producing fluent, coherent and professionally convincing outputs, users may mistake plausibility for reliability. In high-stakes environments such as banking, this risk is particularly significant because an apparently authoritative output may discourage critical scrutiny rather than stimulate it.
The present paper therefore shifts the analytical focus from output governance to decision governance. Output governance is primarily concerned with the characteristics of an AI system: whether it is accurate, robust, fair or explainable. These characteristics remain essential and should not be treated as secondary. However, they are insufficient to explain whether an institution is capable of making responsible decisions with AI. Decision governance considers the wider socio-technical process through which an AI-generated output is transformed into organisational action. This includes model development and validation, the presentation and interpretation of outputs, human review, escalation mechanisms, override authority, documentation and the allocation of accountability.
The paper addresses the following research question:
How can banks design AI governance frameworks that preserve meaningful human judgement and accountability rather than reducing human oversight to the formal approval of AI-generated outputs?
To address this question, the paper develops the concept of judgement-centred AI governance. Under this approach, human oversight is not assessed according to whether a person formally participates in an AI-assisted process. Instead, it is assessed according to whether the individual or group responsible for the decision has the practical capacity to exercise independent judgement. Meaningful oversight requires, at a minimum, sufficient competence to understand the limitations of the AI system, access to relevant information and evidence, genuine authority to challenge or override automated recommendations, and an organisational environment in which critical scrutiny is encouraged rather than treated as an obstacle to efficiency.
The paper makes three contributions. First, it extends the discussion of AI governance beyond technical performance and explainability by distinguishing between the quality of an AI output and the quality of the decision made on the basis of that output. Second, it conceptualises meaningful human oversight as an organisational capability rather than a simple procedural requirement. This directly responds to the implementation gap identified by Tuge and Msweli (2026), whose findings demonstrate that technically available solutions may fail when the organisational conditions necessary for their effective use are absent. Third, the paper proposes a judgement-centred framework for AI governance in banking, structured around contextual risk assessment, traceability, contestability, meaningful override authority, independent challenge and organisational learning.
The argument developed in this paper is not that human judgement is inherently superior to AI. Human decision-making is itself subject to bias, inconsistency, limited attention and error. Rather, the argument is that responsible banking governance requires an institutional capacity to combine the analytical capabilities of AI with meaningful human responsibility. AI systems should therefore be treated neither as autonomous decision-makers nor as neutral technical instruments whose outputs can simply be accepted at face value. Their use must be embedded within governance structures capable of examining assumptions, identifying uncertainty, challenging recommendations and assigning responsibility for consequential decisions.
The remainder of the paper is organised as follows. Section 2 examines the distinction between AI output and decision quality and considers why technically credible outputs may nevertheless produce poor institutional decisions. Section 3 analyses the limitations of nominal human oversight and develops the conditions necessary for meaningful human judgement. Section 4 considers the particular significance of these issues in the banking sector, where AI is increasingly used in consequential and regulated decision-making processes. Section 5 presents the proposed framework of judgement-centred AI governance. Section 6 examines the role of controlled experimentation and organisational learning in strengthening governance capabilities. Finally, Section 7 discusses the broader implications of the framework, and Section 8 concludes by arguing that the effectiveness of AI governance should ultimately be judged not only by what an AI system produces, but by an institution's capacity to determine when its output should not be accepted.
2. AI Output and the Problem of Decision Quality
The increasing sophistication of artificial intelligence has made it necessary to distinguish more carefully between the quality of an AI output and the quality of the decision made on the basis of that output. These concepts are related but should not be treated as equivalent. An AI system may produce a prediction, recommendation, classification or explanation that appears accurate, coherent and professionally credible, while the institutional decision based on that output may nevertheless be inappropriate, unjustified or harmful. Effective AI governance must therefore extend beyond evaluating whether an AI system produces technically satisfactory outputs and examine the wider process through which those outputs are interpreted, contextualised and transformed into organisational action.
This distinction is particularly important in banking. Financial institutions increasingly rely on AI systems to process large volumes of information and generate insights that would be difficult or impossible for human decision-makers to produce at comparable speed. AI is now applied across functions such as credit risk assessment, fraud detection, anti-money-laundering monitoring, customer analytics and risk management (Tuge and Msweli, 2026). In each of these areas, however, the output of an AI system does not automatically constitute the final decision. A credit-risk model may estimate the probability of default, but a lending decision involves broader legal, institutional and contextual considerations. A transaction-monitoring system may identify suspicious activity, but the identification of a pattern is not equivalent to establishing wrongdoing. Similarly, a generative AI system may produce a convincing explanation or recommendation, but fluency and plausibility do not necessarily demonstrate factual reliability.
The governance challenge therefore arises at the point where algorithmic output becomes institutional action. A technically sophisticated system can support better decisions, but it can also create a false sense of certainty if its outputs are treated as authoritative rather than provisional. This risk becomes increasingly significant as AI systems become more capable of producing results that are persuasive to human users. The apparent quality of an output may influence how critically it is examined. In such circumstances, the governance problem is not necessarily that the system has failed technically. Rather, the problem may be that human decision-makers have failed to recognise the limits of the system's contribution.
2.1 Output Quality Does Not Equal Decision Quality
AI performance is commonly evaluated using technical measures such as predictive accuracy, precision, recall, robustness and consistency. These measures are essential for determining whether a system performs its intended function. However, they do not provide a complete assessment of decision quality.
A technically accurate model may still contribute to a poor decision if it is used outside its intended context, interpreted incorrectly or combined with inappropriate organisational processes. Similarly, a model that performs well under testing conditions may behave differently when deployed in a changing operational environment. Decision quality therefore depends not only on whether a model performs accurately but also on whether decision-makers understand the circumstances under which its outputs can be relied upon.
This distinction is especially relevant to complex AI systems. As Tuge and Msweli (2026) demonstrate in their systematic literature review, banking institutions face a persistent implementation gap between the technical availability of explainability methods and their effective integration into organisational practice. The existence of tools capable of explaining an AI output does not guarantee that users will understand those explanations or use them appropriately. Their review identifies the accuracy–interpretability trade-off and the problem of “explaining the explanation” as important technical challenges. More fundamentally, however, these difficulties are compounded by organisational and cultural barriers, including misalignment between technical teams and business units, skills shortages, resistance to change and the absence of standardised evaluation frameworks.
These findings are significant because they demonstrate that a technically generated explanation is not equivalent to effective human understanding. A system may provide a feature-importance score, visualisation or other form of explanation while the individual responsible for the decision remains unable to assess its practical significance. The explanation may therefore create the appearance of transparency without necessarily producing meaningful comprehension.
The same principle applies more broadly to AI outputs. A recommendation is not self-validating merely because it has been produced by an advanced system. Its value depends on the context in which it is generated, the quality and relevance of the underlying data, the assumptions embedded within the model and the way in which the recommendation is subsequently interpreted.
Consequently, the question of AI governance cannot be limited to: “Is the output correct?” It must also include:
What is the output intended to represent?
What assumptions and limitations underpin it?
In what context can it be relied upon?
What information may not have been captured by the system?
Who is responsible for evaluating its relevance?
Can the output be challenged or rejected?
How is the final institutional decision justified?
These questions move the analysis beyond the technical properties of AI and towards the broader governance of decision-making.
2.2 The Problem of Plausibility and Automated Authority
A central risk in AI-assisted decision-making is the tendency to associate technological sophistication with epistemic authority. When a system produces an output that appears precise, coherent or highly sophisticated, users may assume that the underlying result is reliable. Yet plausibility is not evidence of correctness.
This issue is particularly important in the context of contemporary AI systems, including advanced machine-learning and generative models. Such systems can produce outputs that appear authoritative even when the underlying information is incomplete, uncertain or incorrect. In banking, where decisions are often made under conditions of uncertainty, the persuasive quality of an AI output may therefore create a governance problem in its own right.
The problem is not limited to demonstrably incorrect outputs. An output may be broadly accurate while still being inappropriate for the specific decision being considered. For example, a model may correctly identify statistical relationships in historical data while failing to account for a significant change in economic conditions. A credit model may perform strongly across a large population while producing problematic outcomes for particular cases. An explanation may accurately describe the behaviour of a model but remain unintelligible to the stakeholder expected to rely upon it.
The increasing use of AI therefore introduces a distinction between technical authority and decision authority. AI systems may possess considerable technical authority in the sense that they can analyse information and identify patterns beyond normal human capacity. However, technical authority should not automatically become decision authority. The capacity to generate a prediction does not necessarily include the authority to determine how that prediction should influence an institutional decision.
This distinction is essential for banking governance. Banks operate within legal, regulatory and ethical environments in which responsibility for decisions cannot simply be transferred to a technological system. The use of AI may change the process through which decisions are made, but it does not eliminate the need for institutional accountability.
2.3 Explainability as a Necessary but Insufficient Condition
Explainable AI has emerged as one of the principal responses to the problem of opaque algorithmic systems. XAI seeks to make AI decisions more understandable to relevant stakeholders and has been associated with objectives such as transparency, accountability, trust and regulatory compliance (Tuge and Msweli, 2026).
However, explainability should not be treated as a complete solution to the problem of decision quality. An explanation may improve transparency while still failing to produce meaningful understanding. Tuge and Msweli's (2026) review is particularly important in this regard because it demonstrates that the effectiveness of XAI depends on the interaction between technical methods, stakeholder requirements and organisational conditions.
Different stakeholders require different forms of explanation. Technical specialists may require detailed information about model behaviour and performance. Auditors and regulators may require standardised and verifiable evidence. Business managers may require explanations that support risk assessment and strategic decision-making, while customers may require accessible explanations of how decisions affecting them were reached. A single explanation is therefore unlikely to satisfy all relevant audiences.
This stakeholder diversity creates an important governance challenge. Explanations must be not only technically valid but also contextually appropriate. An explanation that is mathematically precise but incomprehensible to the intended decision-maker may contribute little to meaningful oversight. Conversely, a simplified explanation may be easier to understand but fail to capture important limitations or uncertainty.
The governance objective should therefore not be the production of explanations for their own sake. Rather, it should be to enable stakeholders to exercise appropriate judgement. Explainability should be understood as one component of a broader decision-making environment rather than as a self-contained governance mechanism.
This point is consistent with Tuge and Msweli's (2026) conclusion that XAI must evolve from a technical add-on into an embedded organisational capability. Such a transformation requires structural changes in governance, communication and ethical oversight. The implications extend directly to the present argument: if explainability is to support meaningful decision-making, banks must ensure that explanations are integrated into processes that permit questioning, challenge and accountability.
2.4 From Model Governance to Decision Governance
The limitations of a purely technical approach suggest the need to distinguish between model governance and decision governance.
Model governance concerns the AI system itself. It includes issues such as:
data quality;
model development;
validation;
accuracy and performance;
bias and fairness;
robustness;
explainability; and
ongoing monitoring.
These elements are essential. However, model governance alone does not determine whether an institution makes good decisions with AI.
Decision governance, by contrast, concerns the wider socio-technical process through which an AI output becomes an organisational action. It includes:
how AI outputs are presented to users;
what contextual information is available;
who reviews the output;
what expertise those reviewers possess;
whether they can challenge the system;
whether they have authority to override it;
how disagreements are escalated;
how decisions are documented; and
who remains accountable for the final outcome.
The distinction is significant because an institution may possess strong model governance while having weak decision governance. A bank may rigorously validate its models, document their performance and implement sophisticated explainability techniques, yet still create poor outcomes if employees treat AI recommendations as default decisions.
Conversely, effective decision governance can provide an additional layer of protection against unavoidable model limitations. No AI system is entirely free from error, uncertainty or changing environmental conditions. The capacity of human decision-makers to recognise anomalies, introduce contextual information and challenge automated outputs can therefore strengthen institutional resilience.
This does not imply that human judgement is inherently superior to algorithmic analysis. Human decision-making is also vulnerable to bias, inconsistency, limited information and cognitive error. The objective should not be to replace AI with human judgement or to assume that humans automatically provide a corrective to technological limitations. Instead, governance must be designed to combine the complementary strengths of both.
AI may provide speed, consistency and the ability to identify complex patterns. Human decision-makers may contribute contextual understanding, normative judgement, institutional accountability and the ability to question whether a technically valid output is appropriate for a particular situation. The central governance challenge is therefore not to determine whether AI or humans should make decisions. It is to determine how responsibilities should be structured when both contribute to the decision-making process.
2.5 Decision Quality as a Socio-Technical Outcome
The findings of Tuge and Msweli (2026) provide an important theoretical basis for understanding decision quality as a socio-technical outcome. Their review demonstrates that the successful implementation of XAI depends on the interaction of technical, organisational and governance factors. AI systems do not operate independently of institutions; they are embedded within processes, professional roles, incentives and regulatory environments.
This insight should also apply to the governance of AI-assisted decisions more broadly. A decision cannot be evaluated solely by examining the algorithm that contributed to it. The wider decision environment must also be considered.
A judgement-centred approach therefore treats decision quality as the product of interaction between:
the technical system, including the model's performance and limitations;
the information environment, including the data and contextual evidence available to decision-makers;
the human actors, including their competence, experience and capacity for critical evaluation;
the organisational structure, including roles, incentives, escalation procedures and authority; and
the governance framework, including accountability, documentation, monitoring and independent challenge.
Weakness in any of these components can undermine the quality of the final decision. A technically reliable model may be misused by an inadequately trained employee. A knowledgeable reviewer may lack sufficient authority to challenge an automated recommendation. A well-designed governance policy may fail if commercial incentives reward rapid approval over critical scrutiny.
The quality of AI-assisted decision-making must therefore be assessed at the level of the whole decision system, rather than at the level of the algorithm alone.
2.6 Implications for Banking AI Governance
The distinction between output quality and decision quality has significant implications for banking. Banks should not evaluate AI governance exclusively by asking whether a system performs as intended. They should also examine whether the institution has created the conditions necessary for AI outputs to be used responsibly.
This requires attention to the points at which human and machine judgement interact. Institutions should consider whether decision-makers:
understand the purpose and limitations of the AI system;
receive information that allows meaningful interpretation;
can identify circumstances in which the model may be unreliable;
have access to evidence beyond the AI output;
possess genuine authority to challenge recommendations;
can escalate concerns without undue organisational pressure; and
remain clearly accountable for consequential decisions.
These questions shift the focus of governance from controlling the output to governing the decision process.
The distinction also reveals a limitation of governance frameworks that rely heavily on formal compliance mechanisms. An organisation may satisfy a requirement for human oversight by ensuring that an employee approves an AI recommendation. However, if the employee lacks the competence, information or authority to disagree with the system, the approval may have little substantive value.
The next chapter develops this argument further by examining the limitations of formal or nominal human oversight. It argues that the presence of a human within an AI-assisted process should not be treated as sufficient evidence of accountability. Instead, meaningful oversight requires the practical capacity to understand, challenge and, where necessary, reject an AI-generated recommendation. The central question is therefore not simply whether a human is present in the loop, but whether that human retains genuine judgement and influence over the final decision.
3. The Limitations of “Human in the Loop”: From Formal Oversight to Meaningful Judgement
The concept of human oversight occupies a central position in contemporary approaches to responsible artificial intelligence (AI). In high-impact domains such as banking, the continued involvement of human decision-makers is commonly regarded as an important safeguard against algorithmic error, bias and inappropriate automation. The underlying rationale is straightforward: where an AI system can influence consequential decisions, a human should retain the ability to assess its output and intervene where necessary.
However, the existence of a human within an AI-assisted process does not, in itself, establish meaningful oversight. A human may formally review, approve or sign off an AI-generated recommendation while exercising little independent judgement. The result is a potentially important distinction between human presence and human control. The former is a structural feature of a decision process; the latter is a substantive governance capability.
This distinction is particularly important in banking because AI systems are increasingly deployed in decision-support roles rather than operating as completely autonomous systems. The resulting decision is therefore produced through an interaction between machine-generated analysis and human interpretation. As Yeung (2018) argues in the broader context of algorithmic regulation, computational systems can increasingly participate in processes traditionally associated with human judgement and institutional decision-making. The governance challenge is consequently not simply to regulate algorithms as technical objects, but to understand how algorithmic systems alter the distribution of knowledge, discretion and authority within organisations.
The central argument of this chapter is that “human in the loop” should not be treated as a sufficient governance condition. Meaningful human oversight requires the practical capacity to understand, interrogate, challenge and, where appropriate, reject an AI-generated recommendation. This requires more than procedural involvement. It requires competence, information, time, authority, independence and organisational incentives that support critical judgement.
3.1 From Human Presence to Human Agency
The phrase “human in the loop” can conceal substantial variation in the actual role performed by the human decision-maker. At one extreme, the human may independently assess the available evidence, use the AI output as one input among several, challenge the system when appropriate and ultimately determine the outcome. At the other extreme, the human may simply confirm an automated recommendation because the system is assumed to be more accurate, objective or efficient.
These two arrangements may look similar from the perspective of a governance checklist: in both cases, a human participates in the process. Yet they are fundamentally different from an accountability perspective.
The distinction can be expressed as follows:
Human-in-the-loop ≠ meaningful human oversight.
Meaningful oversight exists only where human participation creates a realistic possibility that the AI recommendation will be questioned or changed.
This distinction is consistent with research demonstrating that passive approval of automated recommendations is insufficient to constitute effective human review. Experimental research on AI-assisted decision-making has shown that informing decision-makers about the possibility of system errors and explicitly reminding them of their responsibility for the final decision can influence how actively they process AI recommendations (Malgieri and Comandé, 2017; Bader et al., 2023). The implication is important: human oversight is partly a function of how responsibility is structured and communicated, not merely whether a person is technically involved.
Tuge and Msweli (2026) provide complementary evidence from the banking-specific XAI literature. Their systematic review identifies an implementation gap between the technical development of explainability methods and their effective integration into banking workflows. The authors emphasise that organisational and cultural barriers—including resistance to change, misalignment between technical and business functions, skills limitations and the absence of standardised evaluation frameworks—can prevent technically available XAI capabilities from becoming effective governance mechanisms.
This suggests that the effectiveness of human oversight depends on the surrounding organisational environment. A technically sophisticated AI system cannot create meaningful human judgement if the institution has not created the conditions under which judgement can be exercised.
3.2 Automation Bias and the Risk of Passive Acceptance
One of the principal concerns associated with human oversight is automation bias: the tendency for people to place excessive reliance on automated recommendations, potentially overlooking contradictory information or failing to conduct an independent assessment.
The concept is important because it challenges a simplistic assumption that humans naturally provide an effective corrective to AI. If a human reviewer systematically defers to an AI system, then the addition of a human may do little to mitigate algorithmic error. In extreme cases, human involvement can even legitimise an automated decision without materially increasing its quality.
Recent research reinforces the need to take this risk seriously while also cautioning against treating automation bias as universal. Romeo and Conti (2026), reviewing the literature on automation bias in human–AI collaboration, identify over-reliance on automated recommendations as a recurring challenge across high-stakes applications. Their review also highlights the interaction between trust calibration, cognitive load, timing of intervention and organisational constraints. These factors are particularly relevant to banking, where employees may be expected to review large volumes of AI-generated recommendations under considerable time pressure.
However, empirical research also indicates that human responses to algorithmic advice are more complex than a simple tendency towards blind deference. Alon-Barkat and Busuioc (2023), for example, found no general evidence across their experimental studies that participants were automatically more likely to follow algorithmic advice than equivalent human advice. They did, however, identify selective adherence to algorithmic recommendations when those recommendations aligned with pre-existing stereotypes. Their findings therefore suggest that the governance problem should not be conceptualised solely as universal automation bias. Human–AI interaction can produce more complex patterns of reliance, resistance and selective acceptance.
This qualification is important for a banking governance framework. The objective should not be to assume that humans will inevitably over-trust AI, nor to assume that humans will naturally correct it. Both assumptions are inadequate. Instead, institutions should design governance mechanisms that support calibrated reliance: humans should rely on AI when the system provides useful evidence, but retain the capacity to identify circumstances in which its output may be unreliable or inappropriate.
The relevant governance question is therefore not:
“Will humans trust the AI?”
but:
“Under what conditions should humans trust, question or reject the AI?”
This reframing places judgement at the centre of the governance problem.
3.3 The Competence Problem
Meaningful human oversight requires appropriate competence. A reviewer cannot effectively challenge an AI recommendation if they do not understand the system sufficiently to identify its limitations or interpret its output.
This does not imply that every banking employee involved in an AI-assisted decision must become a machine-learning expert. Rather, competence should be proportionate to the decision's consequences and the complexity of the AI system. Employees responsible for high-impact decisions should understand, at minimum:
what the AI system is designed to do;
what data and assumptions influence its output;
the difference between prediction and certainty;
the system's known limitations and error patterns;
when the model may be operating outside its intended conditions;
what additional evidence should be considered; and
how and when the recommendation should be challenged or escalated.
The competence requirement is particularly important in relation to explainability. Tuge and Msweli (2026) demonstrate that XAI implementation involves a significant communication and interpretation challenge. Different stakeholders require different forms of explanation, and the technical production of an explanation does not guarantee that the recipient can use it effectively.
This creates what might be termed an interpretation gap: the system may be capable of explaining its output, while the human decision-maker remains unable to determine what that explanation means for the decision at hand.
Consequently, training should not focus exclusively on how to operate an AI system. It should also develop AI-critical judgement. Employees need to understand when AI should be trusted, when additional evidence is necessary and when a recommendation should be rejected.
This is fundamentally different from conventional technology training. The objective is not merely competent system use, but competent system interrogation.
3.4 The Information Problem
Competence alone is insufficient if human reviewers lack access to the information necessary to make an independent assessment. Human involvement in an AI-supported decision does not, by itself, guarantee meaningful human judgement or responsibility. For a human decision-maker to be genuinely responsible for an AI-supported decision, they must have sufficient epistemic access to the relevant information and to the basis of the system's recommendation (Baum et al., 2022). This is particularly important where AI systems are used as decision-support tools rather than as fully autonomous decision-makers.
Consider a hypothetical credit decision in which an AI system produces a high-risk classification. If the employee receives only the classification and a brief automated explanation, their capacity to exercise meaningful judgement is severely constrained. They may technically be authorised to override the recommendation, but they lack the evidential basis for doing so. This reflects a broader concern about automation bias, whereby decision-makers may place undue reliance on algorithmic recommendations rather than independently assessing the underlying case (Alon-Barkat and Busuioc, 2023; Kupfer et al., 2023). Consequently, the formal presence of a human decision-maker should not be confused with substantive human oversight.
Meaningful oversight therefore requires access to information beyond the final AI output. Baum et al. (2022) argue that a human decision-maker needs access to an explanation of an AI system's recommendation in order to assess disagreement between their own judgement and that of the system and to make a responsible decision. More generally, explainability is intended to make otherwise opaque model outputs sufficiently understandable to support human assessment and scrutiny (Adadi and Berrada, 2018; Barredo Arrieta et al., 2020).
Depending on the application, this may include:
relevant input variables;
model confidence or uncertainty indicators;
information about missing or anomalous data;
relevant contextual information not captured by the model;
explanations appropriate to the user's role;
known limitations and historical performance;
information about model version and deployment conditions; and
relevant human or external evidence.
The precise information required will depend on the decision context and the type of model being used. This is important because there is no universally appropriate form of explanation: explanations need to be fit for purpose and for the intended recipient, rather than simply providing technical information about the underlying model (Adadi and Berrada, 2018; Barredo Arrieta et al., 2020). In high-stakes settings, this also raises the question of whether explaining a complex black-box model is sufficient at all, since interpretable models may provide a more appropriate basis for consequential decisions (Rudin, 2019).
This principle is particularly important because AI systems necessarily operate through a process of selection. A model does not simply “see” reality. It processes particular data according to particular assumptions and objectives. Information excluded from that process may nevertheless be relevant to the final decision. The limitations of algorithmic decision-making partly arise from the fact that models operate on selected representations of a phenomenon rather than on the phenomenon in its entirety. Consequently, the information considered by a model may not exhaust the information relevant to a human decision-maker. This is one reason why human judgement remains important in AI-supported decision-making, particularly where contextual information or competing considerations are not adequately represented in the model (Barredo Arrieta et al., 2020; Rudin, 2019).
Human oversight should therefore function partly as a mechanism for reintroducing contextual information that may not be represented in the model. This conception of oversight goes beyond simply checking whether the AI's output has been followed. Rather, the human should be capable of assessing the recommendation against other relevant evidence and identifying situations in which the model's assessment may be inappropriate. This distinction is important given evidence that human decision-makers can exhibit automation bias and may therefore accept algorithmic recommendations without sufficiently independent scrutiny (Alon-Barkat and Busuioc, 2023; Kupfer et al., 2023).
This is one reason why explainability should be treated as a decision-support mechanism rather than an end in itself. An explanation becomes valuable when it helps a decision-maker determine whether an AI output is sufficiently supported for the decision being considered. Baum et al. (2022) make a particularly strong version of this argument: explanations can provide the human decision-maker with the epistemic access necessary to evaluate the AI's recommendation, particularly where the human and AI disagree. From this perspective, the purpose of explainability is not merely to make an algorithm technically transparent, but to enable meaningful human judgement, contestability and responsible decision-making (Baum et al., 2022; Barredo Arrieta et al., 2020).
3.5 The Authority Problem
Even a competent and well-informed human reviewer cannot exercise meaningful oversight without decision authority. Human involvement is not meaningful merely because a person is formally placed within the decision-making process; meaningful human oversight requires the capacity to assess, question and, where appropriate, depart from an AI recommendation (Baum et al., 2022; Kupfer et al., 2023). This distinction is particularly important where AI is used as decision support, since a nominal human-in-the-loop can otherwise become little more than a formal checkpoint.
This appears obvious but creates an important practical governance problem. Organisations may formally allocate responsibility to employees while simultaneously establishing incentives that make disagreement with AI difficult. Research on human–AI decision-making demonstrates that decision-makers can exhibit automation bias, whereby they accept algorithmic recommendations without sufficiently independent scrutiny (Alon-Barkat and Busuioc, 2023; Kupfer et al., 2023). If AI-generated recommendations are treated as the institutional default, an employee who overrides them may be required to provide substantially more justification than an employee who accepts them. Over time, this asymmetry can encourage passive compliance, particularly where organisational or technological conditions encourage heuristic rather than systematic information processing (Kupfer et al., 2023).
Meaningful oversight therefore requires an explicit and credible override authority. Human oversight should involve active assessment and verification rather than passive approval of an automated recommendation (Kupfer et al., 2023). The ability to override an AI recommendation is particularly important because a human reviewer can only exercise meaningful judgement if they retain genuine control over the final decision (Baum et al., 2022).
3.6 The Independence Problem
Meaningful judgement also requires a degree of independence from the AI system and from organisational pressures surrounding its use. Human involvement alone does not guarantee independent judgement: decision-makers may be influenced by algorithmic recommendations and may selectively adhere to them depending on factors such as the nature of the recommendation and the decision context (Alon-Barkat and Busuioc, 2023). Similarly, research on automation bias indicates that the presence of AI decision-support can affect whether human reviewers critically evaluate recommendations or simply confirm them (Kupfer et al., 2023).
Independence does not mean that the human reviewer must operate without institutional constraints. Rather, it means that the reviewer must be capable of reaching a conclusion that differs from the AI recommendation without the review process being structurally biased towards acceptance. This distinction is important because meaningful human responsibility requires more than formal participation in a decision: the human decision-maker must have sufficient epistemic access and the practical capacity to exercise independent judgement (Baum et al., 2022). Where the organisational or technological environment systematically encourages acceptance of algorithmic recommendations, the formal presence of a human reviewer may therefore provide only limited substantive oversight (Alon-Barkat and Busuioc, 2023; Kupfer et al., 2023).
This is particularly important where AI systems are introduced to achieve efficiency gains. If the principal performance measure for employees is the number of cases processed, intensive scrutiny of AI outputs may become economically or organisationally undesirable. The use of AI can consequently alter the conditions under which human judgement is exercised: while automation may increase the volume and speed of decisions, it can also create incentives for humans to rely more heavily on algorithmic recommendations (Alon-Barkat and Busuioc, 2023; Kupfer et al., 2023). The technology may therefore create a paradox: the more efficient the AI system becomes, the less organisational capacity there may be to scrutinise its recommendations.
The result is a form of oversight compression, in which increasing volumes of AI-assisted decisions are combined with decreasing amounts of human attention per decision. This term can be understood as a governance risk arising from the interaction between automation, increased decision volume and constrained human attention. It is particularly relevant to the concern that human oversight can become nominal rather than substantive when humans are required to review large numbers of algorithmic outputs without sufficient time, information or institutional support to assess them critically (Baum et al., 2022; Kupfer et al., 2023).
This problem reinforces the need for risk-based governance. Not every AI output requires the same degree of human scrutiny. The appropriate level of human involvement depends partly on the potential consequences of the decision, the characteristics of the application and the possibility of error or uncertainty (Barredo Arrieta et al., 2020). A proportionate approach is therefore preferable to treating all AI-assisted decisions as requiring an identical level of human intervention. In regulatory contexts, risk-sensitive approaches similarly recognise that the intensity of oversight should correspond to the potential risks and consequences associated with the regulated activity (Yeung, 2018).
Low-risk, routine applications may appropriately involve limited intervention, while high-impact or uncertain decisions require greater human attention. This is consistent with the broader principle that explainability and human oversight should be designed in relation to the purpose, context and potential consequences of an AI system rather than treated as uniform requirements applicable in the same way to every application (Barredo Arrieta et al., 2020).
Meaningful oversight should therefore be proportionate to consequence and uncertainty, rather than uniformly applied. Such an approach allows organisational resources to be concentrated where human judgement is most consequential, while reducing the risk that nominal human oversight becomes diluted across large volumes of routine AI-assisted decisions.
3.7 Explainability Does Not Automatically Produce Oversight
The relationship between XAI and human oversight deserves particular attention because explainability is often presented as the mechanism through which humans are empowered to challenge AI.
The logic is intuitively attractive: if people can understand why an AI system produced a particular recommendation, they should be better able to evaluate and challenge it. Yet this relationship is not automatic.
Tuge and Msweli (2026) demonstrate that the banking literature contains significant challenges concerning both the production and interpretation of explanations. Technical methods may involve trade-offs between model accuracy and interpretability, while explanations may themselves require interpretation. In addition, the absence of standardised evaluation frameworks makes it difficult to determine whether an explanation is genuinely useful to a particular stakeholder.
The implication is that explainability should be evaluated by its decision value, not simply by its technical availability.
An explanation is governance-relevant if it helps an authorised human determine whether the AI recommendation is appropriate. If it merely describes model behaviour without enabling meaningful evaluation, its contribution to accountability may be limited.
This distinction is particularly important in banking because different stakeholders occupy different positions in the decision process. A model developer, risk manager, compliance officer, front-line employee, internal auditor and customer may all require different forms of information. Effective governance must therefore align explanation with decision responsibility.
The purpose of explanation should ultimately be to facilitate informed challenge, rather than merely to provide transparency.
3.8 Accountability Cannot Be Delegated to the Algorithm
The limitations discussed above lead to a broader question concerning accountability.
If a human employee approves an AI recommendation, who is accountable for the resulting decision? If the employee simply follows the recommendation, responsibility may become blurred between the individual, the institution and the technology provider. This is particularly problematic when the system is complex enough that no single participant fully understands its behaviour.
Yeung (2018) highlights the broader governance challenge created when algorithmic systems become embedded within processes of regulation and decision-making. As computational systems increasingly influence how decisions are made, conventional assumptions about responsibility and discretion can become difficult to apply.
Banking institutions therefore need to avoid what may be described as accountability diffusion. Responsibility should not become fragmented across model developers, vendors, data scientists, business users and final decision-makers to the point where no actor can clearly explain or justify the final outcome.
A judgement-centred governance model addresses this problem by distinguishing between algorithmic contribution and institutional responsibility.
An AI system may contribute evidence, predictions or recommendations. It does not thereby become the accountable institutional decision-maker. Where a human retains formal decision authority, the governance framework should ensure that this authority is substantive rather than symbolic.
This requires clear allocation of responsibilities across the AI lifecycle, including:
model development;
validation;
deployment;
monitoring;
human review;
escalation;
override;
incident management; and
post-decision evaluation.
The objective is not to attribute every error to an individual human reviewer. Rather, accountability should reflect the distributed nature of AI-assisted decision-making while ensuring that responsibility remains identifiable and actionable.
3.9 Designing for Meaningful Human Judgement
The analysis above suggests that meaningful human oversight can be understood through six conditions, synthesised from the literature on explainability, human–AI decision-making, automation bias and responsible AI (Adadi and Berrada, 2018; Barredo Arrieta et al., 2020; Baum et al., 2022; Alon-Barkat and Busuioc, 2023; Kupfer et al., 2023).
1. Competence
The reviewer has sufficient knowledge to understand the AI system's purpose, limitations and relevant sources of uncertainty. Human oversight requires decision-makers to possess an appropriate level of understanding of the AI system and its outputs; otherwise, formal human involvement may not translate into meaningful judgement (Adadi and Berrada, 2018; Barredo Arrieta et al., 2020). In high-stakes contexts, this requirement is particularly important because the limitations and potential errors of complex models may not be apparent from their outputs alone (Rudin, 2019).
2. Information
The reviewer has access to sufficient evidence and contextual information to assess the AI output independently. This condition follows directly from the argument that responsible human decision-making requires sufficient epistemic access to the basis of an AI recommendation (Baum et al., 2022). Explainability can contribute to this access by enabling decision-makers to understand and evaluate algorithmic recommendations, although the appropriate form of explanation depends on the user, context and purpose of the system (Adadi and Berrada, 2018; Barredo Arrieta et al., 2020).
3. Authority
The reviewer has genuine power to challenge, reject or override the AI recommendation. Human oversight is substantively meaningful only where the human retains the ability to exercise independent judgement rather than merely confirming an automated output (Baum et al., 2022). This is particularly important given evidence that humans may otherwise defer to algorithmic recommendations as a result of automation bias (Alon-Barkat and Busuioc, 2023; Kupfer et al., 2023).
4. Independence
The reviewer can exercise judgement without excessive pressure to accept the automated recommendation. The presence of an AI recommendation can influence human judgement, including through automation bias and selective adherence to algorithmic advice (Alon-Barkat and Busuioc, 2023). Organisational and technological conditions can further shape whether decision-makers critically evaluate AI recommendations or rely on them as a default (Kupfer et al., 2023). Meaningful oversight therefore requires not only formal authority to disagree but also conditions in which disagreement is practically possible.
5. Time and attention
The decision process provides sufficient opportunity for meaningful review, particularly in high-impact cases. Human oversight cannot be meaningful if reviewers are expected to process AI-assisted decisions at a volume or speed that prevents substantive consideration of the evidence. This follows from the broader concern that human involvement can become merely procedural where decision-makers lack the opportunity or capacity to critically assess algorithmic recommendations (Baum et al., 2022; Kupfer et al., 2023). The appropriate allocation of human attention should therefore reflect the significance and potential consequences of the decision (Barredo Arrieta et al., 2020; Yeung, 2018).
6. Accountability
The roles and responsibilities of humans and AI systems are clearly defined, documented and auditable. Human oversight should not be understood merely as the physical presence of a person within an automated decision process. Rather, responsibility requires clarity about who makes the decision, what role the AI system plays, and the basis on which the human decision-maker is expected to exercise judgement (Baum et al., 2022). This is particularly important in complex algorithmic systems, where responsibility can otherwise become difficult to locate and where formal human involvement may obscure rather than resolve questions of accountability (Yeung, 2018).
These six conditions transform human oversight from a procedural requirement into an organisational capability. Rather than asking simply whether a human has reviewed an AI recommendation, organisations should consider whether the human possessed the competence, information, authority, independence and time necessary to exercise meaningful judgement, and whether the resulting decision can be attributed and scrutinised through clear accountability arrangements (Baum et al., 2022; Barredo Arrieta et al., 2020).
They also suggest that the effectiveness of human oversight should itself be subject to governance and measurement. Banks should not simply document that a human reviewed an AI recommendation. They should evaluate whether review is actually substantive. This follows from the distinction between formal human involvement and meaningful human judgement, particularly in contexts where automation bias may cause reviewers to accept algorithmic recommendations without sufficient independent scrutiny (Alon-Barkat and Busuioc, 2023; Kupfer et al., 2023).
Potential indicators include:
rates and patterns of human overrides;
frequency of escalation;
recurring disagreement between human and AI decisions;
the quality of documented justifications;
evidence of consideration of contradictory information;
employee understanding of model limitations;
incidents involving inappropriate reliance on AI; and
changes in decision outcomes following human intervention.
These indicators should not, however, be interpreted mechanically. A low override rate, for example, does not necessarily demonstrate effective oversight: it could indicate that the AI system performs well, but it could equally reflect automation bias, insufficient authority or organisational pressure to accept its recommendations (Alon-Barkat and Busuioc, 2023; Kupfer et al., 2023). Conversely, a high override rate does not automatically demonstrate better oversight. The more meaningful question is whether overrides and other interventions are appropriately justified, evidence-based and responsive to the circumstances of individual cases.
Such measures would therefore provide a more meaningful assessment of human oversight than the existence of a formal sign-off procedure. They shift governance attention from whether a human was involved to whether the human actually exercised meaningful judgement.
3.10 From “Human in the Loop” to “Human on the Hook”
The central argument of this chapter can be summarised through a conceptual shift from human in the loop to human on the hook (Nair, 2026).
“Human in the loop” describes the architecture of a system: a human occupies a position somewhere within the decision process.
“Human on the hook” describes the governance reality: a human or institution retains meaningful responsibility for the consequences of the decision and possesses the practical means to exercise judgement.
The latter concept is more demanding. It requires organisations to ensure that responsibility is accompanied by authority, information and capability. It also recognises that accountability without control is problematic. An employee should not be held responsible for a decision they had no realistic ability to influence.
This distinction provides an important bridge between the conceptual argument of this paper and the governance framework developed in later chapters. If meaningful human judgement is the objective, governance mechanisms must be designed around the conditions that make such judgement possible.
3.11 Conclusion
Human oversight is an essential component of responsible AI governance in banking, but its existence cannot be established merely by placing a person within an AI-assisted decision process. The central governance challenge is to ensure that human involvement remains substantive.
The literature demonstrates that human–AI interaction is complex. Automation bias can create risks of excessive reliance on algorithmic recommendations, but empirical evidence does not support the assumption that humans invariably defer to algorithms. Instead, reliance can vary according to context, information, incentives and the relationship between algorithmic advice and human judgement (Alon-Barkat and Busuioc, 2023; Romeo and Conti, 2026). This complexity reinforces rather than weakens the case for judgement-centred governance.
Tuge and Msweli (2026) further demonstrate that even technically sophisticated explainability mechanisms can fail to deliver effective governance when organisational and cultural conditions are inadequate. Explainability is therefore best understood as an enabler of human judgement rather than a substitute for it.
Meaningful oversight requires competence, information, authority, independence, time and accountability. These conditions must be supported by organisational incentives and governance structures that make challenge legitimate and practical.
The key implication is that banks should stop treating “human in the loop” as an end-state. The relevant question is whether the human reviewer can genuinely influence the decision.
A governance framework that merely requires human approval may preserve the appearance of control while leaving substantive authority with the algorithm. A judgement-centred framework seeks something more demanding: a decision process in which AI can inform judgement without silently becoming judgement itself.
The following chapter applies this distinction specifically to banking and examines why the consequences of weak human oversight are particularly significant in areas such as credit assessment, financial crime monitoring and risk management.
4. AI Governance in Banking: The Consequences of Weak Human Judgement
The governance challenges associated with artificial intelligence (AI) become particularly consequential when AI systems are deployed within banking. Unlike many lower-risk applications of AI, banking systems can directly influence access to financial services, the treatment of customers, the detection and reporting of suspected financial crime, and the management of institutional and systemic risk (Das et al., 2023; Černevičienė and Kabašinskas, 2024). AI is increasingly applied across areas including credit assessment, fraud detection, risk management, compliance and other financial decision-making processes, creating governance challenges that extend beyond the technical performance of individual models (Sailer, 2026; Tuge and Msweli, 2026). Errors may therefore have consequences that extend beyond the immediate interaction between a user and an algorithm. They can affect individuals, financial institutions, regulators and, in some circumstances, the stability and integrity of the wider financial system (Das et al., 2023; García-Llorente and Olmeda, 2026).
The preceding chapters established that the quality of an AI output cannot be equated with the quality of the decision made on its basis, and that formal human involvement does not necessarily constitute meaningful oversight. This chapter applies those arguments to the banking sector. It argues that the consequences of weak human judgement are particularly significant in banking because AI is increasingly embedded in high-impact, regulated and interconnected decision processes (Das et al., 2023; García-Llorente and Olmeda, 2026). The governance of AI in financial institutions must therefore extend beyond questions of model accuracy and technical performance to include questions of accountability, explainability, human oversight and organisational control (Barredo Arrieta et al., 2020; Lioliou et al., 2026).
The issue is not that AI is inherently unsuitable for banking. On the contrary, AI can significantly strengthen banking decision-making by processing large datasets, identifying patterns and supporting consistency at a scale that would be difficult to achieve through human analysis alone (Černevičienė and Kabašinskas, 2024; Sailer, 2026). AI-based systems can therefore provide substantial benefits in areas such as financial analysis, risk assessment, fraud detection and decision support (Das et al., 2023; Tuge and Msweli, 2026). The governance challenge arises when the analytical capabilities of AI are combined with institutional processes that do not provide adequate mechanisms for interpretation, challenge and accountability (Barredo Arrieta et al., 2020; Lioliou et al., 2026). Technical capability may improve the production of predictions or recommendations without necessarily ensuring that those outputs are appropriately interpreted or acted upon by human decision-makers.
This creates a central paradox. The more deeply AI becomes embedded in banking, the less adequate purely technical approaches to AI governance become. As AI moves from experimental applications towards consequential operational decisions, governance must increasingly address not only model performance but also how human judgement is structured around the technology (García-Llorente and Olmeda, 2026; Lioliou et al., 2026). In other words, increasing technical sophistication does not eliminate the need for human governance; it can make the institutional conditions surrounding human judgement more important. Effective AI governance therefore requires attention to the relationship between technical systems, organisational processes, human decision-makers and accountability structures, rather than treating the AI model as an isolated technological object (Barredo Arrieta et al., 2020; Yeung, 2018; Lioliou et al., 2026).
4.1 Banking as a High-Impact AI Environment
The banking sector presents a distinctive environment for AI governance because decisions are simultaneously data-intensive, economically consequential and subject to extensive regulatory expectations. Banks process large volumes of structured and unstructured information and operate sophisticated systems for assessing risk, identifying anomalous activity and allocating resources (Černevičienė and Kabašinskas, 2024; Sailer, 2026). AI and machine-learning techniques are consequently being applied across a growing range of financial activities, including credit assessment, fraud detection, risk management and other forms of financial decision support (Das et al., 2023; Tuge and Msweli, 2026).
These characteristics make banking particularly attractive for AI deployment. Machine-learning systems can identify relationships in large datasets that may not be readily apparent to human analysts, while AI can automate repetitive processes, improve monitoring and potentially increase the consistency and efficiency of decision-making (Černevičienė and Kabašinskas, 2024; Sailer, 2026). The ability to process large and complex datasets is therefore one of the principal advantages of AI in financial services.
However, these same characteristics create governance risks. The scale of banking operations means that a relatively small model weakness can affect large numbers of customers or transactions. The complexity and interconnectedness of financial systems may also allow errors or inappropriate decisions to propagate across related processes, increasing the potential consequences of model failure (Das et al., 2023; García-Llorente and Olmeda, 2026). Moreover, historical financial data may contain existing social and economic inequalities, meaning that models trained on such data can reproduce or amplify problematic patterns (Das et al., 2023). The use of historical data therefore creates a governance challenge that cannot be resolved simply by improving predictive accuracy.
Das, Stanton and Wallace (2023) highlight the importance of algorithmic fairness in financial services, particularly in credit markets. Their analysis demonstrates that algorithmic decision-making can create important fairness questions because financial models operate on data that may reflect historical patterns of unequal access to economic opportunities (Das et al., 2023). This means that apparently neutral variables and predictive relationships can have different consequences across groups, particularly where underlying financial data reflect pre-existing structural inequalities. Consequently, improving predictive performance does not necessarily eliminate concerns about fairness or distributive consequences (Das et al., 2023).
This reinforces the distinction established in Chapter 2. A model can be statistically effective while the decision process in which it is embedded remains ethically, legally or institutionally problematic. Explainability and responsible AI literature similarly emphasise that technical properties such as accuracy and interpretability cannot, by themselves, determine whether an AI system is appropriate for a particular decision context (Barredo Arrieta et al., 2020). In high-stakes applications, the choice of model and the way its outputs are incorporated into decision-making may therefore be as important as predictive performance itself (Rudin, 2019).
AI governance in banking must therefore consider the consequences of both model error and model use. Model governance must address whether the system produces reliable and appropriate outputs, but governance must also examine how those outputs are interpreted, acted upon and incorporated into institutional decision processes (Barredo Arrieta et al., 2020; García-Llorente and Olmeda, 2026). This shifts the focus from the AI model as an isolated technical artefact towards the wider socio-technical system in which the model operates, including human decision-makers, organisational processes and accountability structures (Yeung, 2018; Lioliou et al., 2026).
4.2 Credit Assessment and the Distribution of Opportunity
Credit assessment provides perhaps the clearest example of why human judgement remains important in AI-supported banking decisions. Traditional lending decisions have always involved a combination of quantitative analysis and professional judgement. AI can increase the sophistication of this process by identifying patterns across large datasets and estimating the probability of particular outcomes. Machine-learning models may therefore improve predictive performance relative to conventional approaches in some circumstances, particularly where large and complex datasets contain relationships that are difficult to identify through traditional methods (Černevičienė and Kabašinskas, 2024; Sailer, 2026).
However, credit decisions are not merely predictions. They allocate access to financial resources and can therefore have significant consequences for individuals. A decision to approve, restrict or deny credit can affect an individual's ability to purchase a home, establish a business, manage financial difficulties or participate in the wider economy. The use of algorithmic systems in credit assessment consequently raises questions not only about whether a model predicts repayment behaviour effectively, but also about how its predictions are translated into decisions and whether those decisions are considered fair and appropriate (Das, Stanton and Wallace, 2023).
This creates an important distinction between predictive accuracy and decision legitimacy. A model may accurately predict repayment behaviour while still producing outcomes that raise questions about fairness, discrimination or the appropriate use of information. Das, Stanton and Wallace (2023) demonstrate that algorithmic fairness in financial markets, including credit, cannot be reduced to the predictive accuracy of a model. Fairness involves normative judgements concerning which outcomes are acceptable, how competing objectives should be balanced and what forms of differential treatment can be justified. Algorithmic performance and fairness are therefore related but distinct dimensions of decision quality (Das, Stanton and Wallace, 2023).
Human judgement therefore remains important even when an AI system demonstrates strong predictive performance. The role of the human decision-maker is not simply to verify that the model has produced an output, but to assess whether that output is sufficiently supported and appropriate in the circumstances. This requires the decision-maker to be capable of asking whether:
the information used by the model is appropriate for the decision;
the applicant's circumstances have been adequately represented;
unusual or exceptional circumstances warrant additional review;
the model is operating within the conditions under which its performance has been validated;
the outcome provides evidence of a systematic disparity or other potentially problematic pattern; and
the AI recommendation should be accepted, challenged or escalated.
The precise questions required will depend on the particular system and decision context; these questions should therefore be understood as an operationalisation of the broader governance requirements identified in the literature, rather than as a checklist directly prescribed by any single source. Explainability and access to relevant information are important in this respect because a human reviewer cannot meaningfully assess or challenge an AI recommendation without sufficient epistemic access to the basis and limitations of that recommendation (Baum et al., 2022; Adadi and Berrada, 2018). Similarly, research on automation bias indicates that human decision-makers may place excessive weight on algorithmic recommendations, making active assessment and challenge important components of effective oversight (Alon-Barkat and Busuioc, 2023; Kupfer et al., 2023).
This does not mean that human discretion should simply replace algorithmic assessment. Human decision-makers can themselves introduce inconsistency, subjectivity and bias. The governance objective should instead be to combine the consistency and analytical capabilities of algorithmic systems with structured and meaningful human challenge. This reflects a broader view of AI governance in which responsibility cannot be addressed solely through model performance, but must also consider how AI outputs are interpreted, contested and incorporated into organisational decisions (Barredo Arrieta et al., 2020; Yeung, 2018).
The appropriate role of the human is therefore neither to blindly accept the model nor to disregard it. Rather, the human decision-maker should determine how much weight the model's recommendation should receive in the particular decision context, taking account of the quality and relevance of the underlying information, the circumstances of the individual case, the model's limitations and the potential consequences of the decision. In this sense, meaningful human oversight is not the rejection of algorithmic decision-making but the preservation of human judgement over how algorithmic evidence is interpreted and acted upon.
4.3 Financial Crime and the Problem of Context
AI is also increasingly used in financial crime prevention, including transaction monitoring, fraud detection and anti-money-laundering processes. These applications illustrate another important limitation of purely algorithmic decision-making: pattern recognition is not the same as contextual understanding. AI and machine-learning techniques can support financial institutions in identifying patterns, anomalies and potentially suspicious activity across large volumes of financial data, making them particularly attractive for fraud detection and financial crime monitoring (Černevičienė and Kabašinskas, 2024; Tuge and Msweli, 2026).
An AI system may identify transactions or behavioural patterns that differ from historical norms. This capability is valuable because financial crime can involve complex and evolving patterns that may be difficult to identify through static rules alone. AI can therefore support more dynamic monitoring and potentially reduce the volume of routine analytical work undertaken by investigators, allowing human resources to be directed towards cases requiring greater attention (Černevičienė and Kabašinskas, 2024; Sailer, 2026).
However, an anomalous transaction is not necessarily an illicit transaction.
The distinction matters because financial crime systems can produce false positives. A customer may be flagged because their behaviour differs from historical patterns for legitimate reasons. If such alerts are treated as evidence of wrongdoing without adequate investigation, AI-supported monitoring may impose unnecessary costs on customers and employees and may result in inappropriate decisions or interventions. This illustrates the broader governance problem that predictive or classificatory performance does not, by itself, determine whether the resulting decision is appropriate (Das, Stanton and Wallace, 2023).
Human judgement is therefore required to introduce contextual information that may not be adequately represented in the model's output. Investigators may need to consider the customer's circumstances, transaction history, relevant documentation and other available evidence before reaching a conclusion. Meaningful human involvement, however, requires more than simply assigning an employee to review an alert. The investigator must have sufficient information and understanding to assess the recommendation and determine whether the available evidence supports further action (Baum et al., 2022; Adadi and Berrada, 2018).
The governance requirement is consequently not simply that an AI system achieves a high detection rate. Institutions must also establish whether investigators have:
sufficient information to interpret and contextualise alerts;
appropriate training to understand the system's purpose, limitations and potential sources of error;
adequate time and attention to investigate alerts rather than merely process them;
authority to question or challenge model outputs;
mechanisms for escalating uncertain or consequential cases; and
processes for identifying systematic patterns in false positives and false negatives.
These requirements should be understood as an operationalisation of the broader principles of meaningful human oversight rather than as a checklist directly prescribed by a single source. Research on automation bias demonstrates that human decision-makers can place excessive reliance on algorithmic recommendations, while research on explainability emphasises the importance of providing users with information that enables them to understand and assess AI outputs (Alon-Barkat and Busuioc, 2023; Kupfer et al., 2023; Adadi and Berrada, 2018). In financial crime prevention, this means that investigators should be positioned as active assessors of AI-generated alerts rather than merely as procedural approvers.
This is another example of the distinction between AI-assisted detection and AI-determined judgement. The former can strengthen institutional capability by extending the scale and sophistication of monitoring, while retaining human responsibility for interpreting evidence and determining the appropriate response. The latter can create significant governance risks where an anomalous pattern is treated as sufficient evidence of wrongdoing and relevant contextual information is excluded from the decision. AI should therefore be understood as supporting the detection and prioritisation of potentially significant patterns, rather than automatically determining their meaning or the appropriate institutional response (Barredo Arrieta et al., 2020; Yeung, 2018).
4.4 Customer Decisions and the Problem of Explainability
The customer-facing use of AI introduces an additional governance dimension. Customers increasingly encounter AI systems through chatbots, personalised financial recommendations, automated service decisions and processes involving credit or account access. The increasing application of AI across banking therefore means that explainability is not solely an internal governance concern but can also affect how customers understand and respond to decisions made or supported by AI systems (Tuge and Msweli, 2026; Černevičienė and Kabašinskas, 2024).
In these contexts, explainability is particularly important because customers may need to understand why a decision affecting them has occurred. This is especially significant where an AI-assisted decision has material consequences for access to financial services or where the individual may need to question, contest or seek further information about the outcome (Lui, Lamb and Durodola, 2025). Yet, as Tuge and Msweli (2026) demonstrate, explainability is not simply a matter of producing more technical information. Explanations must be understandable and appropriate to the stakeholder receiving them, reflecting the different purposes and informational needs associated with different users of AI systems (Tuge and Msweli, 2026; Barredo Arrieta et al., 2020).
This creates an important distinction between technical transparency and meaningful explanation. Technical transparency may involve exposing information about the model, its variables, its architecture or its statistical behaviour. Such information can be valuable for technical validation and auditing, but it does not necessarily enable a non-technical recipient to understand why a particular decision was reached. Meaningful explanation, by contrast, requires communicating relevant information in a form that enables the recipient to understand the basis of the decision and, where appropriate, question or challenge it (Adadi and Berrada, 2018; Barredo Arrieta et al., 2020).
For banks, this means that explanations should be designed around the needs of the decision recipient rather than treated as a single, standardised disclosure. A customer does not necessarily need the same information as a model validator. A regulator does not necessarily require the same explanation as a front-line employee. The appropriate form and level of explanation will therefore depend on the recipient's role, knowledge, rights and responsibilities (Barredo Arrieta et al., 2020; Tuge and Msweli, 2026).
Tuge and Msweli (2026) emphasise precisely this challenge in their review of XAI in banking. Their analysis identifies stakeholder diversity and the absence of standardised approaches to evaluating explanations as important barriers to effective implementation. This suggests that banks should not assume that the provision of an explanation automatically satisfies the requirements of transparency or accountability. An explanation may exist in a formal sense while remaining insufficiently understandable, relevant or actionable for the person receiving it (Tuge and Msweli, 2026).
The governance objective should instead be to provide decision-relevant explanations: explanations that enable the relevant stakeholder to understand the basis and limitations of an AI-assisted decision sufficiently to exercise their appropriate rights and responsibilities. For customers, this may involve understanding the principal factors contributing to a decision and the avenues available for clarification or challenge; for employees, it may require sufficient information to assess and question an AI recommendation; and for regulators or technical reviewers, it may require more detailed information concerning model behaviour, limitations, validation and governance. Explainability should therefore be understood not as an end in itself, but as a mechanism for enabling appropriate understanding, scrutiny and accountability within the particular decision context (Adadi and Berrada, 2018; Tuge and Msweli, 2026; Lui, Lamb and Durodola, 2025).
4.5 Model Risk and Changing Environments
Another reason human judgement remains important in banking is that AI systems operate within changing environments. Financial markets, customer behaviour, economic conditions and regulatory requirements are not static. A model trained and validated using historical data may therefore perform differently when the underlying environment changes. This problem is commonly associated with model drift, distributional change or changes in the data-generating environment, and represents an important challenge for the continued reliability of AI systems (Černevičienė and Kabašinskas, 2024; Tuge and Msweli, 2026).
The governance implication is significant. A model cannot be considered permanently reliable simply because it performed well during development and initial validation. Model performance is established under particular data, assumptions and operating conditions, and changes in those conditions can affect the appropriateness of an AI system's outputs. Consequently, effective AI governance requires continued assessment rather than treating initial validation as a permanent assurance of reliability (Barredo Arrieta et al., 2020; Tuge and Msweli, 2026).
Ongoing monitoring is therefore essential. However, monitoring should not be limited to technical performance indicators. Banks should also consider whether the meaning and consequences of model outputs are changing. This requires attention not only to whether a model continues to meet aggregate performance thresholds, but also to whether its outputs remain appropriate for the particular contexts and populations in which they are being used (Das, Stanton and Wallace, 2023; Tuge and Msweli, 2026).
For example, a model may continue to achieve acceptable aggregate predictive performance while becoming less appropriate for a particular customer segment or changing economic environment. Aggregate metrics may therefore conceal important localised problems, particularly where performance or outcomes differ across groups or contexts. This reinforces the distinction between overall model performance and the broader fairness and appropriateness of the decisions produced through its use (Das, Stanton and Wallace, 2023).
Human judgement can provide an additional layer of resilience by identifying circumstances that are difficult to capture through standard monitoring metrics. Employees working directly with AI systems may encounter unusual patterns, recurring exceptions or changes in customer behaviour that prompt questions about whether the model remains appropriate in practice. Such human observations can complement formal monitoring by providing contextual information that may not be captured by aggregate technical indicators (Baum et al., 2022; Barredo Arrieta et al., 2020).
This does not mean that employees should replace formal model monitoring or validation. Rather, human observations can function as an additional source of evidence through which institutions identify emerging limitations and determine whether further investigation, recalibration, retraining or changes to the use of the system may be necessary. Meaningful oversight therefore extends beyond the assessment of an individual AI output to include the capacity to recognise when the conditions surrounding the system's use have changed.
This creates a feedback relationship:
AI supports human judgement → human judgement identifies limitations → institutional learning improves AI governance.
The relationship should therefore be iterative rather than linear. AI systems should be subject to continuing observation, challenge and organisational learning, with information generated through human interaction with the system feeding back into monitoring, validation and governance processes. In this sense, human judgement is not simply a control applied after an AI system has produced an output; it can also form part of the institutional learning process through which the continued appropriateness of AI systems is assessed.
4.6 The Risk of Scaling Poor Decisions
The scale of banking operations introduces another distinctive governance issue: AI can scale both good and bad decisions.
If an AI system improves the identification of fraudulent transactions, its use across millions of transactions may generate significant benefits. Conversely, if the system systematically produces inappropriate classifications, the same scale can multiply the consequences of the underlying weakness. The significance of an AI system therefore depends not only on the accuracy of its individual outputs, but also on the scale, frequency and context in which those outputs are deployed (Das, Stanton and Wallace, 2023; García-Llorente and Olmeda, 2026).
This creates a fundamental difference between individual human error and algorithmically mediated institutional error. A human employee may make a poor decision affecting one customer. An embedded AI system, by contrast, may reproduce the same problematic classification across thousands of cases before the underlying problem is identified. The resulting risk is therefore not necessarily that any individual AI-assisted decision is more erroneous than a human decision, but that a systematic weakness can be reproduced consistently and at a much greater scale (García-Llorente and Olmeda, 2026).
This does not mean that AI necessarily creates greater risk than human decision-making. Automation can also reduce inconsistency and certain forms of human error, while enabling institutions to process volumes of information that would be difficult to manage through purely manual processes (Černevičienė and Kabašinskas, 2024; Sailer, 2026). The important point is that the scale of deployment changes the consequences of governance failure. Where an AI system is embedded within a large-scale banking process, even a relatively limited weakness can produce substantial aggregate consequences if it is repeatedly applied across customers or transactions.
As a result, banks should place particular emphasis on early detection, monitoring and escalation. The objective should be to identify systematic problems before they become embedded at scale. This requires governance arrangements that extend beyond initial model validation and include continuing monitoring of performance, outcomes and the conditions in which the system operates (García-Llorente and Olmeda, 2026; Tuge and Msweli, 2026).
This strengthens the case for controlled experimentation and risk-based deployment. High-impact systems should not simply move from development to full-scale operational deployment without appropriate testing and governance controls. Institutions should have mechanisms for testing how systems behave under realistic conditions, identifying potential failure modes and assessing whether human reviewers can effectively recognise and respond to system limitations (Yeung, 2018; Kupfer et al., 2023).
The level of oversight should also be proportionate to the potential consequences of system failure. Systems capable of affecting large numbers of customers or producing significant financial, legal or institutional consequences warrant greater scrutiny, stronger escalation mechanisms and more demanding conditions for deployment than lower-impact applications. This reflects a risk-based approach to AI governance in which the intensity of oversight is connected to the potential consequences of error rather than being determined solely by the technical characteristics of the model (Yeung, 2018; García-Llorente and Olmeda, 2026).
Human oversight is particularly important within this framework because testing the model itself is not sufficient. Institutions must also establish whether the humans responsible for reviewing AI outputs can recognise inappropriate recommendations, access the information necessary to assess them and exercise meaningful authority when concerns arise (Baum et al., 2022; Kupfer et al., 2023). The governance question is therefore not simply whether the AI system performs adequately in testing, but whether the wider human and organisational system is capable of detecting and containing systematic failures before they are reproduced at scale.
4.7 The Organisational Problem: Efficiency Versus Challenge
Efficiency, human judgement and risk-based oversight
AI adoption is often justified in terms of efficiency. Automation can reduce processing time, lower costs and enable employees to manage larger volumes of work. These benefits are legitimate and may be important to the competitiveness and operational capacity of financial institutions (Černevičienė and Kabašinskas, 2024; Sailer, 2026).
However, efficiency can create tension with meaningful oversight. A system designed to process decisions rapidly may create organisational expectations that employees should review AI recommendations quickly. If employees are required to process increasing numbers of cases without sufficient time for investigation, human oversight can become increasingly superficial. Research on human–AI interaction indicates that organisational and technological conditions can influence the extent to which individuals actively verify algorithmic recommendations, while excessive reliance on AI can contribute to automation bias (Alon-Barkat and Busuioc, 2023; Kupfer et al., 2023).
This creates a potential efficiency–judgement trade-off. The more organisations optimise processes around AI-generated recommendations, the greater the possibility that critical human review becomes perceived as an obstacle to operational efficiency. Employees may have incentives to accept an AI recommendation rather than challenge it when acceptance is quicker, requires less justification or is more consistent with established organisational expectations. This represents a potential organisational mechanism through which efficiency pressures can weaken the practical exercise of human judgement, even where formal human oversight remains in place (Alon-Barkat and Busuioc, 2023; Kupfer et al., 2023).
This is precisely why human oversight must be treated as an organisational capability rather than a procedural checkbox. The presence of a human reviewer does not, by itself, establish meaningful oversight. The reviewer must have sufficient information, competence, time and authority to assess the AI output and, where necessary, depart from it (Baum et al., 2022; Kupfer et al., 2023). Banks therefore need to allocate sufficient organisational resources for meaningful review in proportion to the consequences and uncertainty associated with the decision.
Risk-based approaches are therefore essential. A low-impact customer-service recommendation may require limited human intervention, whereas a decision involving credit denial, account restriction or a significant compliance action may warrant substantially greater scrutiny. This reflects a broader risk-based approach to algorithmic governance in which the intensity of oversight is proportionate to the potential consequences of system failure or inappropriate use (Yeung, 2018; García-Llorente and Olmeda, 2026).
The appropriate governance question is consequently not whether humans should review every AI output to the same degree. It is whether the level of human judgement is proportionate to the potential consequences of error. Such an approach avoids two extremes: treating every AI output as requiring the same degree of manual intervention, or allowing efficiency considerations to reduce human oversight to nominal approval. Meaningful governance instead requires institutions to determine where human attention adds the greatest value and to ensure that sufficient capacity exists for that attention to be exercised effectively.
4.8 The Governance Gap Between Technical and Business Functions
The implementation challenges identified by Tuge and Msweli (2026) also highlight a broader organisational issue: AI governance often involves multiple professional communities with different forms of expertise.
Data scientists may understand model architecture and performance. Risk professionals may understand institutional exposure. Compliance specialists may understand regulatory requirements. Business teams may understand operational context. Front-line employees may understand customer circumstances.
No single group necessarily possesses all the knowledge required to govern AI effectively.
This creates a risk of governance fragmentation. Technical teams may focus on model performance while business teams focus on operational utility. Compliance teams may focus on formal requirements while employees focus on practical workflow constraints.
Tuge and Msweli (2026) identify misalignment between technical teams and business units as an important barrier to effective XAI implementation. Their findings reinforce the need for interdisciplinary governance structures in which technical and non-technical perspectives are integrated.
A judgement-centred approach therefore requires governance mechanisms capable of connecting these forms of expertise.
This may involve cross-functional AI governance committees, clearly defined ownership structures, independent model validation and regular communication between technical, risk, compliance and business functions.
The objective is not to eliminate professional specialisation but to ensure that AI governance does not become isolated within one function.
4.9 Accountability in High-Impact Banking Decisions
The issues discussed above ultimately converge on accountability.
When AI contributes to a banking decision, responsibility can become distributed across multiple actors: the model developer, data scientist, vendor, business unit, reviewer, risk function and senior management. Without clear governance, this distribution can make it difficult to determine who is responsible when something goes wrong. Responsibility for AI-enabled decisions may extend across organisational and technical networks rather than resting with a single individual, creating potential gaps between those who design, deploy, oversee and ultimately act on AI-generated outputs (Yeung, 2018; Baum et al., 2022; Lioliou et al., 2026).
The problem is particularly serious where human reviewers are formally responsible for final decisions but have limited ability to influence them. A human-in-the-loop arrangement does not, by itself, establish meaningful responsibility: the individual must have sufficient epistemic access to understand the AI recommendation and a genuine capacity to exercise judgement over it (Baum et al., 2022). Assigning responsibility without corresponding authority therefore risks creating what can be described as accountability without control. This is particularly problematic where organisational processes encourage acceptance of algorithmic recommendations, as automation bias can lead decision-makers to accept AI outputs without sufficient independent verification (Alon-Barkat and Busuioc, 2023; Kupfer et al., 2023).
A judgement-centred approach addresses this problem by requiring alignment between responsibility, authority and capability. This alignment can be understood as a governance principle derived from the literature on meaningful human oversight rather than as a requirement established by any single source. Where a person is accountable for a consequential decision, that person should have:
sufficient information to understand the basis and limitations of the decision;
sufficient competence to evaluate the AI contribution;
sufficient time to conduct an appropriate review;
authority to challenge or override the recommendation; and
access to escalation mechanisms when uncertainty remains.
These conditions reflect the broader requirement that human oversight should involve genuine assessment rather than merely formal approval. In particular, explanations are relevant to accountability because they can provide the epistemic basis required for a human decision-maker to assess whether an AI recommendation should be accepted, challenged or rejected (Baum et al., 2022; Adadi and Berrada, 2018). Similarly, research on automation bias suggests that effective oversight depends not simply on the presence of a reviewer but on the conditions under which that reviewer verifies and responds to AI recommendations (Alon-Barkat and Busuioc, 2023; Kupfer et al., 2023).
At the institutional level, senior management and governance functions must also remain responsible for determining whether AI systems are appropriate for their intended uses. Accountability therefore extends beyond individual decision-makers to the organisational decisions that determine where AI is deployed, what level of oversight is required, and how responsibility is allocated around the system (Yeung, 2018; García-Llorente and Olmeda, 2026; Lioliou et al., 2026).
The use of an external AI vendor cannot transfer this institutional responsibility. Nor can model complexity justify the absence of accountability. Where responsibility is distributed across developers, vendors and organisational users, governance must make those relationships sufficiently clear to identify who has decision authority, who is responsible for monitoring and intervention, and who remains accountable for the consequences of deployment (Baum et al., 2022; Lioliou et al., 2026).
4.10 Towards Risk-Proportionate Human Judgement
The analysis of AI applications across banking suggests that human oversight should be risk-proportionate rather than uniform. Not every use of AI carries the same potential consequences, and therefore not every application requires the same level of human involvement, review or challenge. Governance arrangements should reflect the nature and significance of the decision being supported by AI, consistent with broader approaches to risk-sensitive algorithmic governance and emerging banking-specific approaches that distinguish between different levels of risk and accountability (Yeung, 2018; García-Llorente and Olmeda, 2026).
A useful conceptual way of expressing this relationship is:
Impact × Uncertainty × Scale = Governance Intensity
This formulation is not intended to represent a mathematical risk model or empirically validated measure. Rather, it provides a conceptual principle for determining the appropriate level of governance. Where the potential impact of an AI-assisted decision is significant, the behaviour or limitations of the system are uncertain, and the system operates at considerable scale, the intensity of governance should increase accordingly. This is consistent with risk-based approaches to algorithmic governance, in which regulatory and organisational controls should be sensitive to the potential consequences of algorithmic systems rather than applied identically across all uses (Yeung, 2018; García-Llorente and Olmeda, 2026). Greater governance intensity may involve more extensive human review, stronger documentation and traceability requirements, greater opportunities for challenge and override, independent assessment, and more frequent monitoring. These specific mechanisms represent an operationalisation of the broader risk-proportionate principle rather than a checklist directly established by the cited literature.
Conversely, applications with limited consequences, well-understood behaviour and relatively low scale may require less intensive human intervention. Routine administrative automation, for example, may require only limited human judgement, while customer-service assistance may warrant a moderate level of review depending on the nature of the interaction. By contrast, applications such as fraud-alert prioritisation, credit assessment, account restriction or closure, and significant risk-management decisions can have material consequences for customers and the institution and therefore require substantially stronger human oversight. AI is increasingly used across financial activities including credit assessment, fraud detection, risk management and customer-facing applications, but the governance implications of these uses differ according to the decision being supported and its potential consequences (Černevičienė and Kabašinskas, 2024; Sailer, 2026; Tuge and Msweli, 2026). Prudential and strategic decisions, where the potential institutional and systemic consequences are particularly significant, warrant the highest level of human judgement and governance scrutiny (Das et al., 2023; García-Llorente and Olmeda, 2026).
The purpose of differentiating governance in this way is to avoid two opposing failures. The first is under-governance, in which consequential AI-assisted decisions receive insufficient scrutiny and may therefore allow errors or inappropriate recommendations to pass into institutional decisions. This is particularly important where algorithmic outputs influence consequential financial decisions or where automation bias may reduce the extent to which human reviewers independently evaluate recommendations (Alon-Barkat and Busuioc, 2023; Kupfer et al., 2023). The second is over-governance, in which low-risk applications are subjected to excessive human intervention and procedural controls that increase costs and reduce the efficiency benefits of AI without materially improving decision quality. The latter is primarily a governance trade-off identified by this analysis rather than a proposition that should be presented as an established empirical finding.
The objective should therefore not be to maximise human involvement in every AI-assisted process. Rather, it should be to ensure that the intensity and nature of human judgement are proportionate to the potential consequences of the decision. Risk-proportionate governance allows banks to preserve efficiency where the consequences of error are limited while ensuring that decisions with significant customer, financial, regulatory or systemic implications receive the level of scrutiny they require (Yeung, 2018; García-Llorente and Olmeda, 2026).
4.11 From Banking Use Cases to Governance Principles
The banking applications considered in this chapter demonstrate that the governance challenge is remarkably consistent across different AI use cases. Although the technical applications differ, the literature on AI in finance repeatedly identifies questions concerning explainability, fairness, accountability, human judgement and the organisational conditions under which AI outputs are used (Černevičienė and Kabašinskas, 2024; Sailer, 2026; Tuge and Msweli, 2026).
In credit assessment, the central issue is whether statistical prediction is appropriately translated into a fair and legitimate decision. Predictive performance alone does not resolve questions of fairness or legitimacy, particularly where algorithmic decisions affect access to financial resources and opportunities (Das et al., 2023; Lui et al., 2025).
In financial crime monitoring, the issue is whether pattern detection is appropriately distinguished from contextual investigation. AI can identify anomalous patterns and support the prioritisation of potentially suspicious activity, but an identified anomaly does not in itself establish that conduct is illicit or that a particular institutional response is justified (Černevičienė and Kabašinskas, 2024; Sailer, 2026). Meaningful human involvement is therefore required where contextual information and judgement are necessary to determine how an AI-generated alert should be interpreted and acted upon (Baum et al., 2022).
In customer-facing systems, the issue is whether technical explanations become meaningful information. Explainability is not simply a matter of exposing technical information about a model; explanations need to be appropriate to the recipient and useful for understanding and, where appropriate, questioning a decision (Adadi and Berrada, 2018; Barredo Arrieta et al., 2020; Tuge and Msweli, 2026). This becomes particularly important where AI contributes to consequential decisions affecting customers, including algorithmic credit decisions (Lui et al., 2025).
In model risk management, the issue is whether historical performance is mistaken for continuing validity. AI systems operate within changing data and organisational environments, meaning that satisfactory performance during development or validation cannot necessarily be treated as permanent evidence of appropriate performance after deployment (Barredo Arrieta et al., 2020; Tuge and Msweli, 2026). Continued monitoring is therefore necessary to identify changes in model behaviour, data conditions and the consequences of AI-assisted decisions.
Across all of these applications, the same underlying principle emerges:
AI should inform institutional judgement without silently acquiring institutional authority.
This principle provides the foundation for the governance framework developed in the following chapter. It reflects the distinction developed throughout this chapter between AI assistance and AI determination: AI may strengthen institutional decision-making by processing information, identifying patterns and generating recommendations, but the existence of an algorithmic recommendation should not by itself determine the institutional response. Meaningful human oversight requires sufficient information, competence, authority and organisational capacity to assess and, where necessary, challenge AI outputs (Baum et al., 2022; Alon-Barkat and Busuioc, 2023; Kupfer et al., 2023).
The framework must therefore address not only the technical properties of AI systems but also the organisational conditions necessary for meaningful judgement. In particular, banks require mechanisms for:
classifying AI applications according to risk;
documenting and reconstructing AI-assisted decisions;
enabling meaningful challenge and contestability;
providing genuine human override authority;
establishing independent review;
monitoring both technical and organisational outcomes; and
learning systematically from errors, exceptions and human interventions.
These mechanisms represent the operationalisation developed by this thesis from the preceding analysis rather than a governance checklist directly prescribed by any single source. Risk classification reflects risk-proportionate approaches to algorithmic governance (Yeung, 2018; García-Llorente and Olmeda, 2026). Documentation, traceability and accountability are particularly important where responsibility is distributed across technical and organisational actors (Lioliou et al., 2026). Challenge, override and meaningful review follow from the literature on human–AI interaction, explainability and automation bias (Baum et al., 2022; Alon-Barkat and Busuioc, 2023; Kupfer et al., 2023). Monitoring and organisational learning extend governance beyond initial model validation towards continued assessment of how AI performs and is used in practice (Barredo Arrieta et al., 2020; Tuge and Msweli, 2026).
These mechanisms transform the abstract principle of human oversight into operational governance. They also establish the basis for assessing whether human involvement is genuinely meaningful rather than merely procedural: the relevant question is not simply whether a human is present, but whether the organisation has created the conditions under which that person can understand, evaluate, challenge and, where necessary, override an AI-assisted recommendation.
4.12 Conclusion
AI has significant potential to improve banking by increasing analytical capacity, consistency and operational efficiency. However, its value depends on the institutional environment in which it is deployed.
The banking sector demonstrates particularly clearly why AI output cannot be treated as equivalent to decision quality. Credit models may predict repayment risk without resolving questions of fairness and context. Financial crime systems may identify anomalies without establishing wrongdoing. Explainability tools may provide technical information without creating meaningful understanding. Models may perform well under historical conditions while becoming less reliable as their operating environment changes.
These challenges do not imply that human decision-making should replace AI. Nor do they imply that AI systems should be distrusted by default. Instead, they demonstrate the need for governance arrangements that enable calibrated reliance: humans should benefit from AI's analytical capabilities while retaining the ability to question, contextualise and reject its recommendations.
The consequences of weak judgement are particularly significant in banking because AI systems can operate at scale, influence high-impact decisions and become deeply embedded within institutional processes. Governance failure can therefore be multiplied across large populations before it is detected.
The central lesson is that banking AI governance must govern not only models but the decisions in which those models participate.
This requires a shift from formal human involvement to meaningful human agency, from technical transparency to decision-relevant explanation, and from model governance to decision governance.
The next chapter develops a practical framework for this approach by proposing six core pillars of judgement-centred AI governance: contextual risk classification, traceability, contestability, meaningful override authority, independent challenge and organisational learning.
5. Towards a Judgement-Centred AI Governance Framework
The preceding chapters have established that responsible artificial intelligence (AI) governance in banking cannot be reduced to the technical performance of AI systems or to the formal presence of a human reviewer. The central governance challenge lies in the relationship between AI-generated outputs, human judgement and institutional decision-making. Research on explainable AI, human–AI interaction and algorithmic governance indicates that effective oversight depends not only on the characteristics of the AI system itself, but also on whether human decision-makers have the information, capacity and authority necessary to evaluate and act upon AI outputs (Barredo Arrieta et al., 2020; Baum et al., 2022; Yeung, 2018). This chapter develops a framework for addressing that challenge.
The proposed framework is termed Judgement-Centred AI Governance (JCAIG). Its underlying premise is that AI should augment institutional decision-making without acquiring decision authority by default. This premise rests on a distinction between AI-generated recommendations and the human exercise of institutional judgement and responsibility: the presence of an AI system should not, in itself, determine the institutional decision that follows from its output (Baum et al., 2022; Alon-Barkat and Busuioc, 2023; Kupfer et al., 2023; Torrecilla-Pinero, 2026). Torrecilla-Pinero (2026) develops a related distinction between merely retaining a human within the causal chain and preserving the human as the subject of the decision. From this perspective, meaningful oversight requires more than formal human involvement or attribution of responsibility; it requires the practical conditions under which a human can understand the relevant circumstances, evaluate the available evidence, question the AI output and, where necessary, exercise independent judgement over the resulting decision. The objective is not to minimise the use of AI or to privilege human judgement over algorithmic analysis. Rather, JCAIG seeks to establish governance arrangements in which the respective capabilities and limitations of humans and AI are recognised, documented and actively managed. This reflects the broader understanding of responsible AI as a socio-technical governance challenge in which technical performance is inseparable from human interaction, organisational context and accountability (Barredo Arrieta et al., 2020; Yeung, 2018).
The framework consists of six mutually reinforcing pillars:
contextual risk classification;
traceability and reconstructability;
contestability;
meaningful human override authority;
independent challenge; and
organisational learning.
These six pillars represent the original analytical structure developed in this thesis. They synthesise the governance requirements identified across the preceding analysis rather than reproducing an existing six-pillar framework from a single source. Their relationship to the literature is nevertheless clear: risk classification reflects risk-sensitive approaches to algorithmic governance (Yeung, 2018; García-Llorente and Olmeda, 2026); traceability and reconstructability respond to questions of accountability and organisational control around algorithmic systems (Lioliou et al., 2026); contestability and override reflect the requirements of meaningful human judgement and the risks of automation bias (Baum et al., 2022; Alon-Barkat and Busuioc, 2023; Kupfer et al., 2023); while independent challenge and organisational learning extend oversight beyond the individual decision to the wider governance system (Yeung, 2018; García-Llorente and Olmeda, 2026).
The proposed framework is termed Judgement-Centred AI Governance (JCAIG). Its underlying premise is that AI should augment institutional decision-making without acquiring decision authority by default. This premise rests on a distinction between AI-generated recommendations and the human exercise of institutional judgement and responsibility: the presence of an AI system should not, in itself, determine the institutional decision that follows from its output (Baum et al., 2022; Alon-Barkat and Busuioc, 2023; Kupfer et al., 2023; Torrecilla-Pinero, 2026). Torrecilla-Pinero (2026) develops a related distinction between merely retaining a human within the causal chain and preserving the human as the subject of the decision. From this perspective, meaningful oversight requires more than formal human involvement or attribution of responsibility; it requires the practical conditions under which a human can understand the relevant circumstances, evaluate the available evidence, question the AI output and, where necessary, exercise independent judgement over the resulting decision. The objective is not to minimise the use of AI or to privilege human judgement over algorithmic analysis. Rather, JCAIG seeks to establish governance arrangements in which the respective capabilities and limitations of humans and AI are recognised, documented and actively managed. This reflects the broader understanding of responsible AI as a socio-technical governance challenge in which technical performance is inseparable from human interaction, organisational context and accountability (Barredo Arrieta et al., 2020; Yeung, 2018).
5.1 The Rationale for a Judgement-Centred Framework
Existing approaches to AI governance typically emphasise principles such as transparency, fairness, accountability, robustness and explainability. These principles are necessary, but their practical effectiveness depends on how they are translated into organisational processes, decision-making practices and mechanisms of oversight (Barredo Arrieta et al., 2020; Yeung, 2018).
This issue is particularly apparent in the banking literature. Tuge and Msweli (2026), in their systematic review of explainable AI in banking, identify a persistent gap between the technical development of XAI and its effective implementation within banking organisations. Their analysis indicates that the implementation and effectiveness of XAI are influenced not only by technical characteristics but also by organisational culture, stakeholder requirements, available skills, interdisciplinary collaboration and governance structures (Tuge and Msweli, 2026). Similar concerns arise in the broader XAI literature, which emphasises that explanations must be designed around their users, purposes and decision contexts rather than treated as purely technical outputs (Adadi and Berrada, 2018; Barredo Arrieta et al., 2020).
This implementation gap is central to the present framework.
A bank may possess an accurate model, a sophisticated explainability tool and a comprehensive AI policy, yet still make poor decisions if employees do not understand the output, lack the authority to challenge it or operate in organisational conditions that discourage sufficiently independent scrutiny. The presence of formal human oversight does not necessarily ensure meaningful intervention, particularly where automation bias or selective adherence influences how decision-makers respond to algorithmic recommendations (Alon-Barkat and Busuioc, 2023; Kupfer et al., 2023). Similarly, explainability has limited governance value if the decision-maker does not receive information that enables them to understand and appropriately evaluate the AI recommendation (Baum et al., 2022).
The governance objective must therefore move beyond asking whether the appropriate controls exist. It must ask whether those controls work in practice. This requires attention to the conditions under which governance mechanisms are actually exercised: whether employees have sufficient information and competence, whether they have time to conduct meaningful review, whether they can challenge and override AI outputs, and whether organisational processes allow concerns to be escalated and acted upon (Baum et al., 2022; Kupfer et al., 2023).
The JCAIG framework is designed around this principle. Its purpose is therefore not simply to specify which governance controls a bank should possess, but to examine whether those controls create the organisational conditions necessary for meaningful judgement, intervention and accountability.
5.2 Pillar One: Contextual Risk Classification
The first pillar is contextual risk classification.
Not all AI applications create the same level of risk, and governance requirements should therefore be proportionate to the potential consequences of failure. Risk-sensitive approaches to algorithmic governance emphasise the importance of tailoring governance mechanisms to the potential consequences and context of algorithmic systems rather than applying identical controls to every application (Yeung, 2018). This principle is particularly relevant in banking, where AI may be deployed across activities with substantially different effects on customers, institutions and the wider financial system (García-Llorente and Olmeda, 2026). A generative AI system used to draft an internal summary does not require the same governance intensity as an AI system contributing to a credit decision or customer account restriction.
Risk classification should consider at least four dimensions:
impact: the potential consequences of an incorrect decision;
uncertainty: the degree to which system behaviour or outputs are uncertain;
scale: the number of decisions or customers potentially affected; and
reversibility: the extent to which an incorrect decision can be corrected.
These dimensions constitute a proposed operationalisation within the JCAIG framework, drawing on the broader principle that governance should reflect the context, risks and consequences of AI deployment (Yeung, 2018; García-Llorente and Olmeda, 2026). They provide a more useful basis for governance than technology type alone. The same underlying model architecture may present radically different governance requirements depending on what it is used to do, who is affected, how widely it is deployed and whether resulting decisions can be readily corrected. This is consistent with the broader literature on responsible and explainable AI, which emphasises that the appropriateness of an AI system cannot be evaluated independently of its application context and intended purpose (Barredo Arrieta et al., 2020; Černevičienė and Kabašinskas, 2024).
For example, an AI application may technically involve the same underlying model architecture in two different contexts while presenting radically different governance requirements. A language model used to summarise an internal meeting presents relatively limited risk. The same model used to generate recommendations affecting customer eligibility may require substantially greater controls because the consequences of an erroneous or inappropriate output are more significant. The distinction is therefore not simply between “high-risk” and “low-risk” technologies, but between different uses of technology within different institutional contexts.
The principle can therefore be expressed conceptually:
Governance intensity should increase as impact, uncertainty, scale and irreversibility increase.
This formulation should be understood as a conceptual governance heuristic rather than a quantitative risk model. Its purpose is to guide the allocation of oversight resources rather than to generate a numerical risk score. Greater risk may warrant stronger controls, including more extensive human review, enhanced documentation, independent assessment, greater challenge and override authority, and more frequent monitoring (Yeung, 2018; García-Llorente and Olmeda, 2026).
This approach also supports proportionality. Excessive human intervention in low-risk applications may reduce the efficiency benefits of AI without providing significant additional protection. Conversely, minimal human involvement in high-impact applications may leave substantial governance gaps, particularly where human reviewers lack the conditions necessary to meaningfully assess AI recommendations (Baum et al., 2022; Kupfer et al., 2023). Risk-proportionate governance therefore avoids both over-governance and under-governance by allocating human attention and institutional controls according to the potential consequences of AI-assisted decisions.
Risk classification should therefore occur before deployment and should be revisited when the purpose, data, model, operating environment or scale of an AI application changes. AI governance cannot rely solely on an initial assessment because changes in the system or its operating context may alter the nature or magnitude of its risks. Continued monitoring and reassessment are therefore necessary components of responsible AI governance (Barredo Arrieta et al., 2020; Tuge and Msweli, 2026).
5.3 Pillar Two: Traceability and Reconstructability
The second pillar is traceability.
A consequential AI-assisted decision should be reconstructable after it has occurred. This means that an institution should be able to determine not only what the AI system recommended but also what information informed the recommendation, which model or system version was used, who reviewed the output and why the final decision was reached. Traceability is particularly important where responsibility is distributed across technical and organisational actors, because reconstructing the decision process provides a basis for identifying how AI contributed to the eventual institutional outcome (Yeung, 2018; Lioliou et al., 2026).
Traceability is fundamental to accountability because decisions that cannot be reconstructed cannot be effectively audited or meaningfully reviewed. The ability to reconstruct an AI-assisted decision also supports human responsibility by making it possible to establish what information was available to the decision-maker, what the AI recommended and how the human decision-maker responded to that recommendation (Baum et al., 2022). In this sense, traceability connects technical records with the exercise of institutional judgement.
A robust traceability framework should capture, where proportionate:
the relevant data inputs;
model and system version;
date and time of the AI output;
AI recommendation or classification;
relevant uncertainty or confidence information;
explanation provided to the decision-maker;
human decision;
any modification or override;
reasons for disagreement;
escalation activity; and
subsequent outcome where available.
These elements constitute a proposed operationalisation within the JCAIG framework rather than a record set prescribed in its entirety by a single source. Their purpose is to connect the technical operation of an AI system with the human and organisational decisions that follow from its output. This is particularly important for consequential applications because accountability requires more than knowing that an AI system was used; it requires sufficient evidence to establish how the system contributed to the decision and how human judgement was exercised (Lioliou et al., 2026; Baum et al., 2022).
This record should not be understood simply as a compliance archive. It should form part of the institution's learning infrastructure. Patterns within the records can provide evidence about how an AI system performs in practice and how human reviewers interact with its recommendations. For example, if human reviewers consistently override a particular model recommendation in a particular type of case, this may indicate a potential limitation, calibration issue or contextual mismatch that warrants investigation. Conversely, if reviewers almost never override a system, the organisation should not automatically interpret this as evidence of model reliability. High levels of acceptance may also be consistent with automation bias or with human review becoming largely procedural (Alon-Barkat and Busuioc, 2023; Kupfer et al., 2023).
Traceability therefore creates the evidential basis for evaluating whether human oversight is actually meaningful. Without records of the AI recommendation, information available to the reviewer, human response, reasons for intervention and subsequent outcomes, it becomes difficult to determine whether human judgement genuinely influenced the decision or merely followed the AI output. Recording interventions and outcomes also allows the organisation to identify recurring patterns that may warrant changes to the model, review process or governance arrangements (Baum et al., 2022; Tuge and Msweli, 2026).
This is consistent with the broader accountability concerns identified by Yeung (2018) and with the emphasis on institutional accountability and control in banking AI governance literature (García-Llorente and Olmeda, 2026; Lioliou et al., 2026). Documentation is therefore valuable not simply because it creates an audit trail but because it makes responsibility, decision processes and human intervention sufficiently visible to be evaluated.
5.4 Pillar Three: Contestability
The third pillar is contestability.
An AI-assisted decision should be capable of being challenged by an appropriately authorised actor. Contestability is closely related to explainability but is conceptually broader. Explainability concerns whether the basis and behaviour of an AI system can be understood, whereas contestability concerns whether the resulting recommendation or decision can be questioned, challenged and, where appropriate, reconsidered (Adadi and Berrada, 2018; Barredo Arrieta et al., 2020).
Explainability asks:
“Can the decision be understood?”
Contestability asks:
“Can the decision be challenged?”
A system may provide a technically accurate explanation while still operating in a process where challenging the recommendation is difficult or discouraged. In that case, transparency may exist without producing meaningful accountability. An explanation is therefore most valuable from a governance perspective when it provides the decision-maker or affected stakeholder with sufficient information to assess whether the AI output should be accepted, questioned or rejected (Baum et al., 2022).
Tuge and Msweli (2026) demonstrate why this distinction is important. Their review identifies difficulties in translating XAI techniques into practical organisational value, particularly where stakeholders have different informational requirements and expectations. The implication is that explanations should be designed not merely to describe system behaviour but to support the decisions and challenges that stakeholders are expected to make (Tuge and Msweli, 2026). This is consistent with the broader XAI literature, which emphasises that explanations should be appropriate to their intended users, purposes and decision contexts rather than treated as purely technical outputs (Adadi and Berrada, 2018; Barredo Arrieta et al., 2020).
In banking, contestability should operate at multiple levels.
Internal contestability
Employees should be able to challenge AI outputs when they have reasonable grounds to believe that the recommendation is inappropriate. This requires not only access to relevant information but also sufficient authority and organisational support to depart from the AI recommendation. Such conditions are important because automation bias can lead decision-makers to accept algorithmic recommendations without sufficient independent verification (Alon-Barkat and Busuioc, 2023; Kupfer et al., 2023).
Functional contestability
Risk, compliance, model validation or internal audit functions should be capable of challenging the design or deployment of AI systems independently of commercial objectives. Contestability should therefore extend beyond individual decisions to the systems and organisational processes that generate those decisions. This reflects the broader understanding of algorithmic governance as involving institutional structures, oversight relationships and accountability mechanisms rather than merely technical model performance (Yeung, 2018; García-Llorente and Olmeda, 2026).
Customer contestability
Where AI contributes to a consequential customer decision, appropriate mechanisms should exist for customers to seek clarification, correction or review, subject to applicable legal and regulatory requirements. This is particularly important where algorithmic decisions affect access to credit or other significant financial services, since meaningful explanation and opportunities for review can be relevant to the ability of affected individuals to understand and challenge decisions (Lui et al., 2025). Customer-facing explanations should therefore be sufficiently meaningful to support the particular form of review or challenge available to the customer (Tuge and Msweli, 2026).
Supervisory contestability
Institutions should be capable of demonstrating to regulators how AI-assisted decisions are governed, reviewed and challenged. This requires sufficient documentation and traceability to allow supervisory or independent scrutiny of how AI contributes to institutional decisions and how governance controls operate in practice (Yeung, 2018; Lioliou et al., 2026).
Contestability therefore transforms explainability from a transparency objective into a mechanism for accountability. The important question is not simply whether an AI decision can be explained, but whether the explanation and surrounding governance arrangements enable an appropriately authorised actor to question, challenge and, where justified, change the resulting decision. Contestability consequently provides a practical link between explainability, human judgement and accountability.
5.5 Pillar Four: Meaningful Human Override Authority
The fourth pillar is meaningful human override authority.
Human oversight is ineffective if a reviewer cannot meaningfully change, reject or escalate an AI recommendation. The existence of an override mechanism is therefore an important but insufficient condition for effective human oversight. Meaningful oversight requires the reviewer to have both the authority and the practical capacity to exercise independent judgement (Baum et al., 2022; Alon-Barkat and Busuioc, 2023; Kupfer et al., 2023).
Meaningful override authority requires three elements.
First, formal authority.
The reviewer must have explicit permission to reject, modify or escalate an AI recommendation. This goes beyond merely assigning a human reviewer to the process: the reviewer must have a genuine capacity to exercise judgement and alter the outcome where appropriate (Baum et al., 2022).
Second, practical authority.
The organisation must provide the time, information and resources necessary to exercise that authority effectively. A reviewer cannot meaningfully challenge an AI recommendation without sufficient information about the recommendation, its basis and its limitations, or sufficient opportunity to assess the case (Adadi and Berrada, 2018; Baum et al., 2022). Research on automation bias also indicates that organisational and technological conditions can influence the extent to which decision-makers actively verify algorithmic recommendations (Alon-Barkat and Busuioc, 2023; Kupfer et al., 2023).
Third, cultural authority.
Employees must be able to disagree with AI systems without being implicitly discouraged from doing so. This dimension extends beyond formal rules to the organisational conditions under which human judgement is exercised. An employee may possess a formal right to override an AI recommendation while perceiving that challenging the system is undesirable, inefficient or inconsistent with organisational expectations. Such conditions can weaken the practical effectiveness of human oversight and increase the risk of automation bias (Alon-Barkat and Busuioc, 2023; Kupfer et al., 2023).
The third element is particularly important. An organisation can formally permit overrides while creating incentives that make them undesirable. If employees are expected to maximise throughput, minimise exceptions or follow standardised recommendations, then the formal existence of an override mechanism may have little practical value. This illustrates the distinction between procedural oversight and substantive oversight: the former establishes that a human is present and formally authorised, whereas the latter requires that the human can and does exercise informed judgement when circumstances warrant it (Baum et al., 2022; Kupfer et al., 2023).
The governance system should therefore monitor not only whether overrides occur but also why they occur and what happens afterwards. Override behaviour can provide evidence about the interaction between the AI system, human judgement and organisational processes. Traceability is therefore important because it allows organisations to reconstruct the recommendation, the information available to the reviewer, the intervention made and the resulting outcome (Lioliou et al., 2026; Baum et al., 2022).
Override data can become a valuable source of institutional intelligence. High override rates may indicate model weakness, inappropriate deployment or contextual mismatch and should therefore trigger investigation rather than automatically being interpreted as evidence of model failure. Very low override rates may indicate strong model performance, but they may also indicate automation bias, weak human engagement or procedural rather than substantive review (Alon-Barkat and Busuioc, 2023; Kupfer et al., 2023).
The correct interpretation therefore requires contextual analysis rather than a predetermined target. JCAIG does not treat a high or low override rate as inherently desirable. Instead, it treats justified override behaviour as evidence that can be used to assess whether human oversight is functioning meaningfully. The objective is not to maximise or minimise human intervention, but to ensure that authorised reviewers can exercise informed judgement when the circumstances require it.
5.6 Pillar Five: Independent Challenge
The fifth pillar is independent challenge.
AI governance should not rely entirely on the individuals or teams responsible for deploying and benefiting from the technology. Independent challenge provides an additional mechanism through which assumptions, limitations, risks and deployment decisions can be examined from a perspective that is not solely aligned with the objectives of the AI-owning business function (Yeung, 2018; García-Llorente and Olmeda, 2026).
The precise institutional arrangement will vary according to the size, complexity and governance structure of the bank, but potentially relevant functions include:
model risk management;
compliance;
operational risk;
internal audit;
data governance;
information security; and
senior risk governance committees.
The principle is that those responsible for evaluating AI risk should have sufficient functional independence, access to relevant information and authority to escalate concerns to challenge those whose performance objectives depend on successful AI deployment. Independence does not necessarily require complete organisational separation; rather, it requires sufficient distance from the incentives and decision-making responsibilities associated with deploying the system (Yeung, 2018; García-Llorente and Olmeda, 2026).
This is particularly important because AI adoption often involves competing organisational objectives. Business units may prioritise speed and customer experience, technology teams may prioritise functionality and innovation, while risk and compliance functions may prioritise resilience, control and regulatory protection. These objectives are not necessarily incompatible, but they can create different assessments of what constitutes an acceptable level of AI risk.
Effective governance therefore requires these perspectives to interact rather than allowing one to dominate. Independent challenge should involve more than reviewing documentation after deployment. It should provide an opportunity to question assumptions, examine evidence, identify limitations and, where necessary, require modification, escalation or further review before or during deployment. This reflects the broader principle that algorithmic governance is a socio-technical process in which technical performance must be considered alongside organisational, institutional and accountability considerations (Yeung, 2018; Barredo Arrieta et al., 2020).
Tuge and Msweli (2026) identify the relationship between technical and organisational stakeholders as an important consideration in the effective implementation of explainable AI. Their systematic review highlights the importance of skills, stakeholder requirements, interdisciplinary collaboration and governance structures in translating technical XAI capabilities into organisational practice. This supports the importance of governance arrangements capable of connecting technical expertise with business, risk, compliance and other relevant organisational perspectives (Tuge and Msweli, 2026).
Independent challenge should therefore be understood not as an obstacle to innovation but as a mechanism for making innovation sustainable. Its purpose is not to prevent AI deployment or to privilege risk functions over business and technology teams. Rather, it is to ensure that decisions about AI deployment are exposed to sufficiently independent scrutiny before technical capability, commercial objectives or implementation momentum become the dominant considerations. In this sense, independent challenge complements meaningful human override: override allows an authorised decision-maker to challenge an individual AI recommendation, while independent challenge allows the organisation to challenge the assumptions, design and deployment of the AI system itself.
5.7 Pillar Six: Organisational Learning
The sixth pillar is organisational learning.
AI governance cannot be treated as a static control framework because both AI systems and their operating environments change. Models may be updated, data distributions may shift, customer behaviour may change and new risks may emerge. The validity and appropriateness of an AI system must therefore be considered in relation to the context in which it operates rather than assumed to remain constant following initial validation (Barredo Arrieta et al., 2020; Tuge and Msweli, 2026).
Organisations must therefore learn from actual experience. A judgement-centred governance system should collect and analyse information relating to:
model errors;
false positives and false negatives;
human overrides;
customer complaints;
escalation events;
regulatory findings;
unexpected model behaviour;
changes in relevant populations or operating environments; and
instances where human intervention materially altered an AI-supported decision.
These categories should be understood as a JCAIG operationalisation of organisational learning, rather than as a monitoring checklist prescribed by the literature. Together, they provide evidence about not only whether a model is technically performing as expected, but also how it behaves in practice, how humans interact with it and whether its outputs remain appropriate for the populations and circumstances in which it is deployed (Das et al., 2023; Tuge and Msweli, 2026).
This information should feed back into model development, validation, training and governance. Organisational learning therefore requires mechanisms through which observed errors, exceptions and human interventions can result in changes to the system or to the way it is governed. This is consistent with the broader view of responsible AI as requiring continuing evaluation of technical performance, human interaction and organisational context rather than treating deployment as the endpoint of governance (Barredo Arrieta et al., 2020; Tuge and Msweli, 2026).
The resulting process can be conceptualised as a continuous feedback loop:
Deploy → Observe → Challenge → Learn → Adjust → Revalidate → Deploy
This feedback loop is a conceptual contribution of JCAIG, rather than an established model lifecycle from the cited literature. Its purpose is to emphasise that governance continues after deployment. This differs from a conventional conception of model validation in which assurance may be concentrated around development and pre-deployment assessment. Judgement-centred governance instead treats deployment as the beginning of an ongoing learning process in which real-world outcomes provide evidence for subsequent review, adjustment and revalidation.
Human judgement is particularly important within this feedback process. Human overrides, escalations, complaints and unexpected cases can reveal information that aggregate model-performance measures do not necessarily capture. At the same time, human intervention should not automatically be treated as evidence that the model is deficient. An intervention may reflect an appropriate exception, a limitation in the model, a change in circumstances or an issue in the way the system has been deployed. Organisational learning therefore requires interpretation of these events rather than simply counting them (Baum et al., 2022; Alon-Barkat and Busuioc, 2023; Kupfer et al., 2023).
This approach is particularly consistent with the findings of Tuge and Msweli (2026), who emphasise the need to embed XAI within broader organisational structures rather than treating it as a standalone technical intervention. Their review highlights organisational factors, stakeholder requirements, skills and interdisciplinary collaboration as important considerations in implementing explainable AI. Organisational learning extends this principle by requiring the organisation not only to establish governance structures but also to use experience from deployment to improve those structures over time (Tuge and Msweli, 2026).
Organisational learning therefore completes the JCAIG framework. The purpose of human oversight is not only to intervene in individual decisions, but also to generate knowledge that improves future decisions, models and governance arrangements. In this sense, meaningful human judgement becomes part of an institutional learning system rather than merely a final checkpoint before an AI-supported decision is implemented.
5.8 Integrating the Six Pillars
The six pillars should not operate independently. Their value lies in their interaction. Together, they provide a governance structure through which AI systems can be assessed, challenged, controlled and improved throughout their operational life. This reflects the broader view that responsible AI governance requires the interaction of technical, organisational and institutional mechanisms rather than reliance on a single control (Barredo Arrieta et al., 2020; Yeung, 2018).
Risk classification determines the level of governance required.
Traceability provides evidence about what occurred.
Contestability creates the ability to question the decision.
Override authority enables meaningful human intervention.
Independent challenge provides an additional perspective on the AI system and its deployment.
Organisational learning ensures that experience can inform future governance.
The framework can therefore be represented conceptually as a governance cycle:
Risk classification → AI-assisted decision → Human interpretation → Challenge/override → Independent review → Traceability → Learning → Reclassification/revalidation
This cycle is a conceptual representation of JCAIG, rather than a lifecycle model directly established by any single source. Its purpose is to demonstrate how the six pillars reinforce one another. Risk classification determines the appropriate level of scrutiny; traceability creates the evidence necessary for review; contestability and override authority create opportunities for intervention; independent challenge provides an organisational check; and organisational learning feeds experience back into future governance (Yeung, 2018; García-Llorente and Olmeda, 2026; Tuge and Msweli, 2026).
This cycle is important because it prevents governance from becoming a one-time approval process. An AI system should not be considered “governed” merely because it passed an initial review. Changes to models, data, populations, operating environments and organisational use can affect the appropriateness and performance of an AI system, requiring continuing assessment and governance (Barredo Arrieta et al., 2020; Tuge and Msweli, 2026).
The central implication is therefore that governance should continue throughout the operational life of the AI system. Deployment represents not the end of governance, but the point at which the organisation begins to accumulate evidence about how the system performs, how humans interact with it and whether its original assumptions remain appropriate.
5.9 Measuring Meaningful Oversight
A central implication of the judgement-centred governance framework is that human oversight should itself be measurable. It is not sufficient for a bank to demonstrate that a human was formally involved in a decision. Traditional governance metrics may record whether a case was reviewed or approved by an employee, but such measures provide limited insight into the quality or substance of that review. A reviewer who routinely accepts AI recommendations without adequately examining the underlying evidence may satisfy a formal requirement for human oversight while exercising little meaningful control over the decision (Baum et al., 2022; Alon-Barkat and Busuioc, 2023; Kupfer et al., 2023).
A more effective approach is therefore to assess whether the conditions necessary for meaningful judgement are present and whether human intervention has a substantive effect on the decision process. Relevant indicators may include the proportion of cases receiving substantive human review, the frequency with which AI outputs are questioned or challenged, and the rate and pattern of human overrides. Banks may also assess whether the explanations provided by an AI system actually support decision-making, whether reviewers demonstrate an appropriate understanding of the system's capabilities and limitations, and how frequently cases are escalated for additional consideration (Adadi and Berrada, 2018; Baum et al., 2022; Kupfer et al., 2023).
Other indicators should address the broader governance environment. These may include the proportion of consequential decisions that can be reconstructed after the event, the extent to which incidents or identified weaknesses result in changes to models or governance processes, and the degree to which AI systems are subject to review by independent control functions (Lioliou et al., 2026; García-Llorente and Olmeda, 2026). Ultimately, however, the most important consideration is whether human intervention improves or protects decision quality. This requires examining outcomes and subsequent consequences rather than relying exclusively on process-based indicators.
These measures should not be interpreted mechanically. A high override rate, for example, is not necessarily evidence of effective oversight. It may indicate that a model is poorly calibrated for its deployment context, that reviewers are unnecessarily rejecting appropriate recommendations, or that the system is being used outside its intended context. Similarly, a low override rate may indicate either that the AI system performs effectively or that reviewers are excessively relying on its recommendations. Research on automation bias demonstrates why low levels of human intervention cannot automatically be interpreted as evidence of effective AI performance (Alon-Barkat and Busuioc, 2023; Kupfer et al., 2023).
The purpose of measurement should consequently be to identify patterns of judgement, rather than to reward a particular level of human intervention. Effective governance should seek evidence that reviewers are engaging with AI outputs critically, using relevant evidence, challenging recommendations when appropriate and exercising their authority when circumstances require it.
The distinction is therefore fundamental: governance metrics should measure the quality and effectiveness of human judgement, not simply the quantity of human involvement.
The proposed indicators in this section constitute a JCAIG measurement framework, rather than an established measurement standard in the literature. Their purpose is to operationalise the broader principle of meaningful human oversight and provide evidence that can be examined by management, risk functions, auditors and regulators.
5.10 The Relationship Between AI and Human Judgement
The framework does not assume that human judgement is inherently superior to algorithmic decision-making.
Humans are subject to cognitive biases, inconsistency, fatigue and incomplete information. AI systems can process large quantities of data consistently and identify patterns that humans may overlook. In many contexts, combining the capabilities of humans and AI may therefore provide advantages over relying exclusively on either (Černevičienė and Kabašinskas, 2024; Kou et al., 2026).
The objective of judgement-centred governance is consequently complementarity.
AI should perform tasks for which computational analysis provides genuine advantages. Humans should retain responsibility for contextual interpretation, normative judgement and decisions involving significant uncertainty or consequence. This does not imply that every element of a decision must be performed manually. Rather, it requires the organisation to determine which elements are appropriately delegated to AI and which require continuing human judgement (Baum et al., 2022; Kou et al., 2026).
This creates a division of labour rather than a hierarchy.
The appropriate question is not:
“Should humans or AI make the decision?”
It is:
“Which aspects of the decision should be supported by AI, which require human judgement, and how should responsibility be allocated between them?”
This question is particularly important as AI systems become increasingly capable. Greater technical capability should not automatically lead to greater decision authority. The ability of an AI system to generate persuasive, accurate or highly sophisticated recommendations does not by itself determine whether the system should have authority to make the resulting institutional decision (Baum et al., 2022; Alon-Barkat and Busuioc, 2023).
Indeed, as AI systems become more persuasive and capable, the ability of humans to critically evaluate and challenge their outputs may become increasingly important. Research on automation bias suggests that decision-makers can place excessive weight on algorithmic recommendations, making the preservation of meaningful human judgement an organisational rather than merely technical requirement (Alon-Barkat and Busuioc, 2023; Kupfer et al., 2023).
5.11 Governance as Institutional Capability
The six-pillar framework ultimately treats AI governance as an institutional capability rather than a collection of controls.
This distinction is important. A policy can state that AI outputs must be reviewed. A procedure can require an explanation. A system can provide an override button. A governance committee can approve deployment.
None of these mechanisms guarantees meaningful governance. As the preceding pillars demonstrate, their effectiveness depends on whether the organisation has the information, competence, authority, independence and organisational processes necessary to use them effectively (Baum et al., 2022; Tuge and Msweli, 2026).
Institutional capability exists when the organisation can repeatedly and reliably:
understand where AI is being used;
identify the consequences of failure;
evaluate AI outputs critically;
challenge inappropriate recommendations;
intervene when necessary;
reconstruct consequential decisions;
identify emerging weaknesses; and
learn from experience.
This eight-part capability model is a JCAIG operationalisation of the six pillars rather than a framework directly reproduced from the literature. Its purpose is to shift attention from whether individual governance controls exist to whether the organisation can repeatedly perform the activities required for meaningful oversight.
This capability must be distributed throughout the organisation. It cannot reside exclusively with a central AI governance team. AI governance involves different forms of expertise and responsibility, and effective implementation requires interaction between technical, organisational and governance stakeholders (Tuge and Msweli, 2026; Yeung, 2018).
Senior management establishes risk appetite and accountability. Technical teams understand model behaviour. Risk and compliance functions provide challenge. Business teams understand operational context. Front-line employees interact with the systems in practice. Internal audit provides independent assurance.
Judgement-centred governance therefore requires distributed but coordinated responsibility. Distribution does not mean that accountability becomes diffuse or that responsibility is impossible to identify. Rather, it recognises that meaningful AI governance depends on coordinated contributions from actors with different forms of knowledge, authority and organisational responsibility (Yeung, 2018; García-Llorente and Olmeda, 2026).
5.12 A Proposed Governance Maturity Model
The framework can be further developed through a five-stage maturity model. The model is proposed as an analytical extension of JCAIG, rather than as an empirically validated maturity framework.
Level 1: Automated
AI outputs are used with minimal human involvement. Governance is predominantly technical and reactive.
Level 2: Supervised
Humans formally review AI outputs, but review is primarily procedural. Human involvement exists, but there is limited evidence that reviewers consistently exercise independent judgement.
Level 3: Challenged
Human reviewers have explicit authority and mechanisms for questioning AI recommendations. Human intervention is recognised as a substantive part of the decision process rather than merely a procedural approval step.
Level 4: Judgement-Centred
Human oversight is supported by competence, information, independence, traceability and systematic escalation. Reviewers are able to understand, assess and challenge AI outputs, and the organisation can reconstruct consequential decisions and interventions.
Level 5: Learning Governance
The organisation continuously evaluates AI–human interactions and uses the results to improve models, processes, training and governance. Experience from deployment becomes an input into subsequent validation, adjustment and governance decisions.
The objective should not necessarily be to reach Level 5 for every AI application. Governance maturity should be proportionate to risk. A low-impact application may appropriately operate with a lower level of human involvement, while high-impact banking applications should move beyond procedural supervision towards genuine judgement-centred governance (Yeung, 2018; García-Llorente and Olmeda, 2026).
The maturity model therefore complements the risk-classification pillar. Maturity should be assessed in relation to the consequences, uncertainty, scale and reversibility of the AI application, rather than treated as a universal organisational target.
This maturity model also provides a practical means of evaluating progress. Rather than asking simply whether an institution has an AI governance framework, regulators and boards can ask how mature that framework is in practice. The relevant question becomes whether governance arrangements provide increasing levels of human capability, challenge, accountability and organisational learning as the potential consequences of AI use increase.
5.13 Implications for Boards and Senior Management
The proposed framework has important implications for senior management and boards.
First, AI governance should be treated as a business and institutional risk issue rather than exclusively as a technology issue. AI systems can influence credit, customer access, financial crime detection, operational processes and risk management, meaning that their governance can have consequences extending beyond the technical performance of the underlying model (Černevičienė and Kabašinskas, 2024; García-Llorente and Olmeda, 2026).
Second, boards should seek evidence that high-impact AI systems are subject to meaningful challenge and that responsibility for consequential decisions remains clear. This requires attention not only to whether appropriate policies exist but also to whether decision processes can be reconstructed and challenged when necessary (Lioliou et al., 2026; García-Llorente and Olmeda, 2026).
Third, management should ensure that employees responsible for AI-assisted decisions have adequate training, information, resources and authority. Human oversight cannot be meaningful where reviewers lack the capability or organisational conditions necessary to exercise independent judgement (Baum et al., 2022; Kupfer et al., 2023).
Fourth, organisations should avoid creating performance incentives that unintentionally reward acceptance of AI recommendations over critical review. Automation bias and organisational conditions can influence whether decision-makers actively verify algorithmic recommendations, making the design of work processes and incentives relevant to the effectiveness of human oversight (Alon-Barkat and Busuioc, 2023; Kupfer et al., 2023).
Finally, senior management should ensure that AI governance information reaches the appropriate level of organisational oversight. Boards should be able to understand not merely how many AI systems are deployed but where those systems influence consequential decisions, what incidents have occurred, how often significant challenges or overrides occur, and whether human intervention is functioning as intended.
These recommendations constitute the governance implications of the JCAIG framework. They translate the six pillars into questions of organisational responsibility and board oversight rather than suggesting that the cited literature provides a specific board-level checklist.
5.14 Implications for Regulators
The framework also has implications for regulatory supervision.
Regulatory assessment could move beyond checking whether institutions possess AI policies and documented human oversight towards examining evidence of effective governance in operation. This is consistent with the broader shift from formal governance arrangements towards accountability-oriented and risk-sensitive oversight (Yeung, 2018; García-Llorente and Olmeda, 2026).
Relevant supervisory questions might include:
Can the bank identify all high-impact AI applications?
Can consequential AI-assisted decisions be reconstructed?
Are human reviewers actually challenging AI outputs?
What evidence demonstrates that reviewers understand model limitations?
How are overrides analysed?
How are customers able to contest consequential decisions?
How does independent challenge operate?
What changes have been made in response to AI incidents?
These questions are proposed as a JCAIG supervisory application, rather than as a list directly established by the literature. Their purpose is to translate the six pillars into observable evidence of institutional capability.
Such questions would move regulatory attention towards outcomes and organisational capability. Rather than asking solely whether a bank has established a formal human-in-the-loop procedure, supervision could examine whether the institution can demonstrate that humans have the information, competence and authority necessary to exercise meaningful judgement and whether the organisation learns from the results of that judgement (Baum et al., 2022; Tuge and Msweli, 2026).
This is consistent with the broader argument that governance principles become meaningful only when translated into institutional practice. García-Llorente and Olmeda (2026) emphasise the relationship between governance principles, risk and accountability-oriented oversight in banking, while Yeung (2018) highlights the importance of examining how algorithmic governance operates within broader institutional structures.
5.15 Conclusion
This chapter has proposed a judgement-centred AI governance framework designed to address the limitations of purely technical and procedural approaches to AI governance in banking.
The framework consists of six pillars: contextual risk classification, traceability and reconstructability, contestability, meaningful human override authority, independent challenge and organisational learning.
The six pillars are intended to work as an integrated system. Risk classification determines where stronger safeguards are necessary; traceability makes decisions reconstructable and auditable; contestability enables challenge; override authority makes meaningful human intervention possible; independent challenge reduces organisational blind spots; and organisational learning ensures that governance evolves in response to experience (Yeung, 2018; Baum et al., 2022; Tuge and Msweli, 2026; García-Llorente and Olmeda, 2026).
The framework also establishes an important conceptual distinction between human presence and human agency. A human reviewer is not necessarily exercising meaningful oversight simply because they approve an AI recommendation. Meaningful oversight exists when the reviewer has the competence, information, authority, independence and organisational support necessary to exercise genuine judgement (Baum et al., 2022; Alon-Barkat and Busuioc, 2023; Kupfer et al., 2023).
The framework therefore shifts the central question of AI governance from:
“Was a human involved?”
to:
“Could the human meaningfully disagree?”
This is the core of judgement-centred governance.
The objective is not to make humans the default decision-makers in every AI-assisted process. Nor is it to restrict the use of AI. Rather, the objective is to ensure that technological capability does not silently become institutional authority. AI should augment institutional decision-making while remaining subject to governance arrangements that preserve appropriate human judgement, accountability and organisational control (Baum et al., 2022; Alon-Barkat and Busuioc, 2023).
The ultimate test of a banking AI governance framework should therefore be whether it enables an institution to recognise when an AI output is reliable, when it requires contextual interpretation and when it should not be accepted.
In this sense, effective AI governance is not principally about controlling machines. It is about preserving the institutional capacity for responsible judgement in an increasingly machine-assisted decision environment.
6. Operationalising Judgement-Centred AI Governance
The governance framework developed in Chapter 5 provides a conceptual structure for preserving meaningful human judgement in AI-assisted banking decisions. However, a governance framework has limited value if it remains at the level of principles. The central implementation challenge is therefore to translate the six pillars of judgement-centred AI governance into operational practices that can function within the constraints of real banking organisations (Barredo Arrieta et al., 2020; Tuge and Msweli, 2026).
This chapter addresses that challenge. It argues that effective AI governance requires more than policies, committees and approval processes. Banks must design workflows in which human judgement is practically possible, measure whether it is actually being exercised, and create organisational conditions that support appropriate challenge rather than passive acceptance (Alon-Barkat and Busuioc, 2023; Kupfer et al., 2023).
The operationalisation of judgement-centred AI governance can be understood through five interconnected activities:
identifying and classifying AI-assisted decisions;
designing human–AI workflows around risk;
equipping reviewers with appropriate information and authority;
testing governance arrangements through controlled experimentation; and
monitoring outcomes and continuously adapting the governance system.
These five activities constitute a proposed JCAIG implementation model. They translate the six conceptual pillars developed in Chapter 5 into organisational activities that can be observed, tested and improved.
The resulting approach treats governance as an ongoing operational capability rather than a one-time approval exercise.
6.1 From Governance Principles to Operational Controls
The first implementation challenge is to translate broad governance principles into specific and observable organisational controls. Concepts such as human oversight, transparency, accountability and explainability are central to responsible AI governance, but they can be interpreted in different ways. Without clear operational definitions, these concepts risk becoming statements of compliance rather than safeguards that meaningfully influence how decisions are made (Barredo Arrieta et al., 2020; Yeung, 2018; Tuge and Msweli, 2026).
For example, a requirement that “human oversight must be maintained” provides little practical guidance unless the organisation specifies what that oversight entails. It must be clear who is responsible for reviewing the AI output, what information that person should have access to, how much time should be available for the review and what level of expertise is required. The organisation must also determine which decisions require mandatory human review, whether the reviewer has genuine authority to reject or modify an AI recommendation, how disagreements should be documented and who ultimately remains accountable for the decision (Baum et al., 2022; Kupfer et al., 2023).
Judgement-centred governance therefore requires each governance principle to be translated into an observable organisational practice. Contextual risk classification should determine how consequential a decision is and, consequently, what level of human review is appropriate. Traceability should ensure that consequential decisions can be reconstructed after they have been made. Contestability should establish who is able to challenge an AI output and the grounds on which that challenge can be made. Override authority should ensure that a human reviewer has the practical ability to change an AI-supported outcome where appropriate. Independent challenge should establish which functions can assess the system without being directly responsible for its business performance. Finally, organisational learning should ensure that errors, overrides and unexpected outcomes are systematically considered in improvements to models, workflows, training and governance arrangements (Yeung, 2018; Baum et al., 2022; García-Llorente and Olmeda, 2026; Tuge and Msweli, 2026).
This translation from principle to practice is essential because the effectiveness of governance depends not only on whether controls formally exist, but also on whether they can be used effectively in the circumstances in which they are required. A theoretically robust control may provide little protection if employees do not have sufficient time, information, expertise or authority to operate it under normal working conditions (Baum et al., 2022; Kupfer et al., 2023).
Consequently, the effectiveness of an AI governance framework should be assessed not simply by asking whether the appropriate controls are documented, but by asking whether employees can realistically apply those controls when making decisions. Governance becomes meaningful only when its principles are embedded in the everyday processes through which AI-assisted decisions are produced, reviewed, challenged and ultimately acted upon.
6.2 Designing Risk-Proportionate Human–AI Workflows
The second implementation requirement is workflow design.
AI should not simply be inserted into an existing banking process and surrounded by a generic human approval step. The introduction of AI can change the distribution of attention, information and responsibility within a workflow. Human–AI interaction is therefore partly a question of organisational and process design rather than solely a question of model performance (Alon-Barkat and Busuioc, 2023; Kupfer et al., 2023; Kou et al., 2026).
For low-risk activities, AI may legitimately automate substantial parts of the process. For higher-risk decisions, however, the workflow should deliberately preserve opportunities for human interpretation and challenge. This reflects the principle of risk-proportionate governance established in Chapter 5 (Yeung, 2018; García-Llorente and Olmeda, 2026).
A useful distinction is between three broad modes of AI involvement.
AI as drafting support
The AI prepares, summarises or structures information while the human remains responsible for the substantive assessment.
This model is appropriate for many administrative and documentation tasks, subject to appropriate controls concerning accuracy, confidentiality and the consequences of erroneous outputs.
AI as decision support
The AI provides a recommendation, classification or prioritisation that informs a human decision.
This model requires stronger controls because the AI output can influence the direction of human judgement. Automation bias research demonstrates that decision-makers may rely on algorithmic recommendations in ways that make the design of the surrounding workflow particularly important (Alon-Barkat and Busuioc, 2023; Kupfer et al., 2023).
AI as decision execution
The AI output directly triggers an action with limited human intervention.
This model requires the highest level of governance where the action is consequential, difficult to reverse or capable of producing significant harm. The appropriate governance intensity should depend on the consequences and context of the action rather than on the technical sophistication of the AI system alone (Yeung, 2018; García-Llorente and Olmeda, 2026).
These categories should not be treated as fixed technological classifications. The same AI system may fall into different categories depending on how it is deployed, what authority is assigned to its output and who is affected by the resulting decision.
This reinforces the importance of contextual risk classification established in Chapter 5.
6.3 Designing the Reviewer Interface for Critical Judgement
An important but frequently overlooked aspect of AI governance is the design of the interface through which humans encounter AI outputs.
The reviewer does not interact with an abstract model. They interact with a screen, recommendation, alert, summary or explanation. Consequently, interface design can influence whether a reviewer challenges or accepts the AI output, particularly where outputs are persuasive or appear authoritative (Alon-Barkat and Busuioc, 2023; Kupfer et al., 2023; Romeo and Conti, 2026).
A governance-oriented interface should therefore provide more than a final recommendation. Where appropriate, it should expose:
relevant supporting evidence;
important missing information;
uncertainty or confidence indicators;
material assumptions;
relevant model limitations;
contradictory evidence;
unusual characteristics of the case; and
available escalation or override mechanisms.
This is particularly important in situations where an AI system produces fluent or highly persuasive outputs. The existence of an explanation does not by itself ensure meaningful human judgement; the information must be relevant to the task and usable by the person responsible for the decision (Adadi and Berrada, 2018; Baum et al., 2022).
The purpose of the interface should not be to overwhelm the reviewer with technical information. Rather, it should provide the information necessary to exercise professional judgement.
This reflects the distinction between technical explainability and decision-useful explanation discussed earlier. Tuge and Msweli (2026) emphasise that XAI must account for different stakeholder requirements and be embedded within organisational processes. An explanation that is useful to a model developer may be largely irrelevant to a relationship manager or compliance reviewer. Explainability should therefore be designed around the purpose, user and decision context (Adadi and Berrada, 2018; Barredo Arrieta et al., 2020).
The appropriate design question is therefore:
“What information does this reviewer need in order to make a defensible decision?”
This question operationalises the broader JCAIG principle that explainability should function as a resource for judgement rather than as a transparency objective in isolation.
6.4 Training for AI-Assisted Judgement
Meaningful oversight requires more than general AI literacy.
Employees involved in consequential AI-assisted decisions should understand at least:
what the system is designed to do;
what it is not designed to do;
the types of information it uses;
relevant limitations and failure modes;
indicators of unreliable output;
how uncertainty should be interpreted;
when human judgement should take precedence;
how to challenge or override an output; and
how to document the rationale for disagreement.
This list represents a JCAIG operationalisation of reviewer competence rather than a training curriculum established by a single source. Its purpose is to translate the competence and information requirements identified in Chapter 5 into practical training objectives.
This does not mean that every employee must become a machine-learning specialist.
The required level of competence should be proportional to the decision and the employee's role. A front-line employee using an AI system to structure customer notes requires different knowledge from a model validator responsible for assessing model performance. A compliance specialist reviewing AI-supported financial-crime alerts requires a different understanding from a board member overseeing the institution's overall AI risk appetite.
Training should therefore be role-specific.
Tuge and Msweli (2026) identify organisational and cultural barriers, as well as gaps between technical and business communities, as important challenges for XAI implementation. Training can help address these gaps, but only when it is integrated into the workflow rather than delivered as a one-off compliance exercise. The objective should be to develop the practical capability to interpret, question and appropriately use AI outputs rather than simply to demonstrate awareness of AI terminology (Tuge and Msweli, 2026; Barredo Arrieta et al., 2020).
6.5 Controlled Experimentation as a Governance Capability
One of the most important implementation principles is controlled experimentation.
Banks cannot fully understand the practical effects of AI governance before employees begin using AI in real workflows. At the same time, unrestricted experimentation in high-impact decision processes creates unacceptable risk.
The solution is controlled experimentation.
A controlled experiment should establish:
a clearly defined use case;
a limited population or scope;
explicit objectives;
predefined risk boundaries;
authorised participants;
monitoring requirements;
escalation procedures;
documentation requirements; and
criteria for continuation, modification or termination.
This constitutes a proposed JCAIG approach to testing governance arrangements. It draws on the broader principle of risk-proportionate and controlled algorithmic governance (Yeung, 2018), while extending that principle to the empirical testing of human–AI interaction.
The purpose is not simply to test whether the AI produces technically useful outputs. It is to examine how people interact with those outputs.
For example, a bank introducing an AI tool for customer-contact documentation could examine:
whether employees detect missing information;
whether AI-generated text is accepted without verification;
how frequently users correct the output;
which types of errors are most common;
whether users understand system limitations;
whether workload changes the depth of review; and
whether the tool changes the way employees interact with customers.
These observations provide governance information that cannot be obtained from model accuracy metrics alone. A technically strong system can still create governance problems if users misunderstand its limitations, rely on it excessively or fail to identify particular types of error (Baum et al., 2022; Alon-Barkat and Busuioc, 2023; Kupfer et al., 2023).
Controlled experimentation should therefore be regarded as a learning mechanism.
6.6 The Importance of Psychological and Organisational Conditions
Human judgement is influenced by organisational conditions.
Even a well-trained employee may fail to challenge an AI recommendation if the surrounding organisation communicates that speed, standardisation or productivity are the dominant objectives. Research on automation bias indicates that human interaction with algorithmic advice is influenced by contextual and organisational conditions rather than being determined solely by the existence of an algorithm (Alon-Barkat and Busuioc, 2023; Kupfer et al., 2023).
This creates a potential conflict between efficiency incentives and judgement incentives.
If an employee receives positive performance recognition for processing a large number of cases but receives little recognition for identifying inappropriate AI outputs, the organisational environment may implicitly encourage acceptance. The precise relationship between incentives and acceptance will vary by organisational context, so this should be treated as a governance risk to be examined rather than as an assumption that employees will necessarily behave in this way.
The governance system should therefore recognise challenge as productive work.
Employees should not be measured exclusively on throughput. Relevant performance considerations may include:
quality of review;
appropriate escalation;
identification of system limitations;
justified overrides;
detection of missing information; and
contribution to process improvement.
These indicators are proposed JCAIG measures rather than a validated performance framework. They are intended to shift attention from the quantity of human involvement towards the quality and appropriateness of human judgement.
This does not imply that employees should be rewarded simply for disagreeing with AI. Unnecessary overrides can also create risk.
The objective is evidence-based challenge.
The organisational culture should communicate that accepting an AI recommendation and rejecting it are both legitimate outcomes when supported by appropriate reasoning. This is consistent with the distinction between automation bias and meaningful human judgement: the objective is not maximum disagreement but appropriate critical engagement with the AI output (Alon-Barkat and Busuioc, 2023; Kupfer et al., 2023).
6.7 Preventing Oversight Compression
As discussed in Chapter 3, increased AI efficiency can create a paradox.
AI may reduce the time required to generate or analyse information, but this does not necessarily mean that the time required for responsible decision-making has fallen by the same amount. AI may increase processing capacity while leaving the need for contextual interpretation, verification and challenge largely intact (Baum et al., 2022; Kupfer et al., 2023).
If an organisation responds to AI productivity gains by increasing the number of cases assigned to each reviewer, the amount of attention available for each consequential decision may decline.
This phenomenon can be described as oversight compression.
The workflow becomes more efficient in terms of processing volume while potentially becoming weaker in terms of individual scrutiny. This is a JCAIG analytical concept, intended to describe a governance risk arising when increases in technological throughput are not matched by sufficient human attention.
The risk may be particularly significant where AI produces polished and apparently complete outputs. Reviewers may process large numbers of cases because the information appears easier to assess, even though detecting subtle errors may require careful attention. Automation bias and the conditions surrounding human verification make this relationship an important governance consideration (Alon-Barkat and Busuioc, 2023; Kupfer et al., 2023; Romeo and Conti, 2026).
Banks should therefore consider workload and review time as governance variables.
For high-impact applications, monitoring should examine whether employees have sufficient time to perform the level of review required by the risk classification. Human oversight should therefore be treated partly as a resource-allocation problem: the organisation must allocate sufficient human attention to the decisions for which human judgement is expected to provide meaningful control.
6.8 Governance Across the AI Lifecycle
Judgement-centred governance should operate throughout the AI lifecycle rather than beginning only after deployment.
Stage 1: Use-case identification
The organisation identifies the intended purpose of the AI system and determines whether the use case is appropriate.
Stage 2: Risk classification
The potential impact, uncertainty, scale and reversibility of decisions are assessed.
Stage 3: Development and selection
Relevant data, model characteristics, performance criteria and limitations are assessed.
Stage 4: Validation
Technical performance and relevant governance risks are independently evaluated.
Stage 5: Workflow design
Human roles, review requirements, escalation mechanisms and override authority are established.
Stage 6: Controlled deployment
The system is introduced within defined operational boundaries.
Stage 7: Monitoring
Technical performance and human–AI interaction are monitored.
Stage 8: Review and learning
Incidents, overrides, complaints and unexpected outcomes are analysed.
Stage 9: Revalidation or retirement
The system is modified, revalidated, restricted or withdrawn where evidence indicates that its governance or performance is inadequate.
This nine-stage lifecycle is a proposed operationalisation of JCAIG. It draws together the risk, human oversight, traceability, challenge and learning requirements developed in Chapter 5 and is consistent with the literature's emphasis on context, organisational implementation and continuing evaluation (Barredo Arrieta et al., 2020; Tuge and Msweli, 2026).
This lifecycle reinforces a central principle of this paper:
AI governance does not end when a model is approved for production.
The operational environment itself generates evidence about whether governance assumptions remain valid. Changes in models, data, populations, customer behaviour or organisational use may require renewed assessment rather than reliance on the conditions that existed at initial approval (Das et al., 2023; Tuge and Msweli, 2026).
6.9 Allocating Accountability Across Organisational Roles
Judgement-centred governance also requires clear allocation of responsibility.
AI decisions often involve multiple actors, including technology teams, data scientists, model validators, business owners, front-line employees, risk functions and external vendors.
This creates a risk of accountability diffusion. Responsibility for AI-supported decisions can extend across technical and organisational networks, making it important to establish clearly who holds which responsibilities rather than assuming that the person closest to the final decision bears all responsibility (Yeung, 2018; Baum et al., 2022; Lioliou et al., 2026).
The governance framework should therefore distinguish between different forms of responsibility.
Board and senior management
Responsible for strategic direction, risk appetite, organisational resources and the overall governance environment.
Business owner
Responsible for the appropriateness of the AI application within its operational context.
Model or system developers
Responsible for technical design, documentation and identified limitations.
Model validation and risk functions
Responsible for independent assessment of model and system risks.
Front-line decision-makers
Responsible for applying professional judgement within the defined workflow and escalating cases where necessary.
Compliance and legal functions
Responsible for assessing relevant regulatory, legal and conduct requirements.
Internal audit
Responsible for independent assurance regarding the effectiveness of governance and controls.
External technology providers
Responsible for meeting contractual, technical and transparency requirements established by the bank, without displacing the bank's institutional responsibility.
This allocation is a proposed JCAIG responsibility structure, rather than a claim that every bank should use precisely these roles. Institutional arrangements will differ according to organisational structure, regulatory requirements and the nature of the AI application.
The final point is particularly important. Outsourcing an AI system does not outsource accountability for the banking decision in which that system is used. External providers may bear contractual and technical responsibilities, but the bank remains responsible for determining whether and how the technology is appropriately deployed within its institutional decision-making processes (Yeung, 2018; García-Llorente and Olmeda, 2026).
6.10 Monitoring Human–AI Interaction
Traditional AI monitoring tends to focus on model performance.
Judgement-centred governance requires an additional monitoring layer: human–AI interaction.
Relevant measures include:
acceptance rates;
override rates;
escalation rates;
frequency of reviewer corrections;
time spent reviewing outputs;
differences between AI recommendations and final decisions;
patterns in reviewer behaviour;
complaints and appeals;
error types detected by humans; and
outcomes following human intervention.
These indicators constitute a proposed JCAIG human–AI interaction monitoring layer. They can reveal governance problems that conventional model monitoring may not capture.
For example, a model may maintain stable statistical performance while reviewers increasingly accept its recommendations without meaningful examination. The model itself may not have degraded, but the governance surrounding it may have weakened. Conversely, an increase in overrides may indicate model deterioration, changes in the operating environment or improved human challenge.
Monitoring should therefore analyse patterns and causes, rather than relying on individual metrics. This is particularly important given the ambiguity of override rates: both unusually high and unusually low intervention may have multiple explanations (Alon-Barkat and Busuioc, 2023; Kupfer et al., 2023).
The objective is therefore not to establish a predetermined optimal level of human intervention, but to determine whether human–AI interaction is consistent with the risk and purpose of the system.
6.11 Evaluating Whether Human Oversight Works
The central test of judgement-centred governance is whether human oversight improves or protects decision quality.
This is difficult to establish through policy documentation alone. A documented review requirement demonstrates that a control exists, but it does not demonstrate that the reviewer understands the AI output, evaluates relevant evidence or exercises independent judgement (Baum et al., 2022; Kupfer et al., 2023).
Banks should therefore seek empirical evidence.
Possible evaluation methods include:
comparing outcomes before and after AI deployment;
comparing AI-assisted and human-only workflows;
analysing override outcomes;
reviewing samples of accepted and rejected recommendations;
conducting structured reviewer assessments;
examining incidents and near misses;
measuring customer outcomes; and
testing whether reviewers identify deliberately introduced system errors in controlled environments.
These methods are proposed JCAIG evaluation techniques rather than an established empirical standard. Their purpose is to test whether governance arrangements function as intended rather than assuming that formal controls necessarily produce effective oversight.
The final method is particularly useful for training and governance validation.
A controlled exercise could present reviewers with a mixture of correct and intentionally flawed AI recommendations and assess whether they identify the problems. This would provide evidence about actual oversight capability rather than assumed oversight capability. Such testing should be proportionate to risk and designed to avoid creating unnecessary operational exposure.
The broader principle is that meaningful human oversight should be tested behaviourally where practicable, rather than inferred solely from policies, training records or the existence of an override function.
6.12 Addressing Automation Bias Through Governance Design
The literature on human–AI interaction suggests that the relationship between people and algorithmic recommendations is more complex than the assumption that humans will always blindly follow machines.
Alon-Barkat and Busuioc (2023), for example, found no general pattern of automation bias across their experimental settings, while identifying circumstances in which people were more likely to follow algorithmic recommendations when those recommendations aligned with existing stereotypes. This finding is important because it suggests that governance should not be based on simplistic assumptions about human behaviour.
Instead, banks should identify the conditions under which reliance becomes problematic.
Romeo and Conti (2026) similarly emphasise the importance of factors including trust calibration, cognitive load and the timing of human intervention in human–AI collaboration. These findings reinforce the need to design governance interventions around the conditions of use rather than assuming that one generic human-review procedure will be effective across all applications.
Possible mechanisms include:
requiring reviewers to assess key evidence before displaying the AI recommendation in particularly sensitive workflows;
displaying contradictory evidence alongside the recommendation;
requiring a rationale for high-impact overrides or approvals;
introducing escalation triggers for unusual cases;
periodically testing reviewer ability to detect system errors; and
redesigning workflows where evidence indicates excessive reliance.
These mechanisms constitute proposed JCAIG interventions rather than universally validated controls. Their suitability should depend on the application, risk level and evidence concerning actual human–AI interaction.
The objective is not to make human review unnecessarily difficult. It is to prevent convenience from becoming a substitute for judgement.
6.13 Managing the Trade-Off Between Control and Innovation
Strong governance can itself create risks if implemented without proportionality.
Excessive approval requirements may slow innovation, discourage experimentation and cause employees to avoid potentially valuable AI applications. Governance that imposes identical controls on every AI use case may therefore be inefficient and may fail to reflect differences in potential consequence (Yeung, 2018).
The solution is not weak governance but risk-proportionate governance.
Low-risk applications should generally be subject to lighter controls, while high-impact applications require greater scrutiny. The precise controls should reflect factors such as potential impact, uncertainty, scale and reversibility, as established in the contextual risk-classification pillar.
This approach enables organisations to experiment safely without treating every AI application as though it carried identical consequences.
Controlled experimentation is particularly valuable in this context because it creates a middle ground between unrestricted deployment and complete prohibition. The organisation can learn about AI systems while keeping their scope, users and potential consequences within defined boundaries.
The objective is therefore to reconcile innovation with governance rather than treating them as inherently opposing objectives. Risk-proportionate controls can allow experimentation while preserving stronger safeguards where the consequences of failure are greater (Yeung, 2018; García-Llorente and Olmeda, 2026).
6.14 The Role of Governance Committees
Governance committees can provide coordination across business, technical and control functions, but their role should not be limited to approving new AI systems.
An effective AI governance forum should also examine:
significant incidents;
emerging risks;
changes in AI use cases;
material model updates;
override patterns;
customer complaints;
regulatory developments;
evidence of automation bias or oversight compression;
results from controlled experiments; and
lessons incorporated into policies and training.
This is a proposed JCAIG conception of the governance committee as a learning and challenge mechanism, rather than simply an approval gate. It reflects the broader importance of interdisciplinary organisational structures in implementing responsible and explainable AI (Tuge and Msweli, 2026; Yeung, 2018).
This is particularly important as AI becomes embedded across multiple business functions. Without central coordination, different business units may develop inconsistent approaches to risk classification, human review and accountability. Governance committees can therefore provide a mechanism for identifying cross-organisational patterns and ensuring that lessons from one AI application inform governance elsewhere.
6.15 A Practical Implementation Roadmap
A bank seeking to implement judgement-centred AI governance could proceed through five phases.
Phase 1: Map
Identify where AI currently influences decisions, including informal or employee-led uses that may not be captured by formal inventories.
Phase 2: Classify
Assess applications according to impact, uncertainty, scale and reversibility.
Phase 3: Design
Define the appropriate human role, information requirements, challenge mechanisms, override authority and accountability structure.
Phase 4: Test
Use controlled experimentation, reviewer testing and workflow monitoring to determine whether the governance design functions as intended.
Phase 5: Learn
Use incidents, overrides, complaints, performance data and reviewer feedback to refine models, workflows, training and governance requirements.
This five-phase roadmap is a proposed implementation sequence derived from JCAIG, rather than an established banking implementation standard. Its purpose is to provide a practical bridge between the conceptual framework developed in Chapter 5 and organisational implementation.
The roadmap reflects the central philosophy of the framework: governance should evolve through evidence.
6.16 Implementation Challenges
Several challenges may complicate implementation.
First, measurement is difficult. Decision quality is often multidimensional and may not be observable immediately. Some consequences of AI-assisted decisions may emerge only over time, making it difficult to establish whether an intervention genuinely improved the outcome.
Second, human judgement is itself variable. Greater human involvement does not automatically guarantee better outcomes. Humans may introduce their own biases, inconsistencies and errors, meaning that the objective is not to maximise human intervention but to make human judgement appropriately informed and accountable (Alon-Barkat and Busuioc, 2023; Baum et al., 2022).
Third, governance creates costs. Additional review, documentation and independent challenge require organisational resources. This reinforces the importance of risk-proportionate governance rather than uniform controls across all applications (Yeung, 2018).
Fourth, AI systems change rapidly. Governance processes designed for one generation of AI may become inadequate as capabilities, use cases and deployment environments evolve. Continuing review and organisational learning are therefore necessary (Tuge and Msweli, 2026).
Fifth, vendor dependence can create information asymmetry. Banks may not possess complete visibility into externally developed AI systems. This creates challenges for explainability, validation, traceability and accountability, particularly where technical information is necessary for meaningful organisational oversight (Adadi and Berrada, 2018; Lioliou et al., 2026).
Sixth, organisational incentives can undermine formal controls. A bank may possess a strong governance framework while employees are still rewarded primarily for speed and volume. Automation-bias research reinforces the importance of examining the organisational conditions under which human review actually occurs (Alon-Barkat and Busuioc, 2023; Kupfer et al., 2023).
These challenges do not invalidate the framework. They reinforce the need to treat governance as a continuous capability that must itself be tested and improved.
6.17 From Compliance to Capability
The most important implication of operationalising judgement-centred AI governance is that governance should not be understood solely as a compliance obligation.
Compliance asks whether specified requirements have been met.
Capability asks whether the organisation can reliably respond when circumstances change.
This distinction becomes increasingly important as AI systems become more capable and widespread. A bank may satisfy formal requirements while lacking the organisational capacity to recognise when an AI system is being used outside its intended context, when human reviewers are relying on it excessively or when changing conditions have undermined the assumptions on which governance was originally based (Tuge and Msweli, 2026; García-Llorente and Olmeda, 2026).
A bank with strong governance capability should be able to answer questions such as:
Where is AI influencing consequential decisions?
What happens when the AI is wrong?
Who notices?
Who can intervene?
Who can challenge the decision?
Who is accountable?
What evidence demonstrates that the control works?
What has the organisation learned from previous failures?
These questions provide a more meaningful test of institutional readiness than the existence of policies alone.
The distinction between compliance and capability is therefore central to JCAIG. The framework does not reject policies, procedures or formal controls. Rather, it asks whether those mechanisms produce an organisation capable of exercising responsible judgement under real operating conditions.
6.18 Conclusion
This chapter has translated the judgement-centred AI governance framework into an operational model for banking organisations.
The analysis shows that meaningful governance depends on the interaction between technology, workflow design, human capability and organisational incentives. Human oversight cannot be added to an AI system as a final approval step. It must be designed into the process from the beginning (Baum et al., 2022; Alon-Barkat and Busuioc, 2023; Tuge and Msweli, 2026).
Five operational priorities are particularly important.
First, AI use cases should be classified according to their consequences rather than their technological characteristics alone.
Second, human–AI workflows should be designed to provide reviewers with the information, time, competence and authority necessary for meaningful judgement.
Third, banks should monitor human–AI interaction alongside conventional model performance.
Fourth, controlled experimentation should be used to test whether governance arrangements work in practice and to develop organisational understanding of AI limitations.
Fifth, governance should operate as a learning cycle in which incidents, overrides, complaints and other evidence lead to changes in models, workflows, training and controls.
These five priorities constitute the operational expression of the six JCAIG pillars developed in Chapter 5. They translate the framework from a set of governance principles into organisational practices that can be designed, tested and evaluated.
The result is a shift from governance as permission to governance as capability.
The critical question is therefore no longer whether a bank has a human in the loop. It is whether the institution has created the conditions under which that human can genuinely understand, question and, where necessary, reject an AI-generated recommendation (Baum et al., 2022; Alon-Barkat and Busuioc, 2023).
This distinction provides the foundation for the discussion of the broader implications, limitations and future development of judgement-centred AI governance in the following chapter.
7. Implications, Limitations and Future Directions
The preceding chapters have developed a judgement-centred approach to AI governance in banking. The central argument is that effective governance cannot be reduced to model performance, technical explainability or the formal presence of a human reviewer. Instead, governance must preserve the institutional capacity to understand, question, challenge and, where appropriate, reject AI-generated recommendations (Baum et al., 2022; Barredo Arrieta et al., 2020; Alon-Barkat and Busuioc, 2023).
The framework developed in Chapter 5 and operationalised in Chapter 6 therefore has implications beyond the design of individual AI systems. It affects how banks conceptualise accountability, how regulators assess effective oversight, how employees interact with increasingly capable AI systems and how organisations should evaluate the success of AI adoption. These implications are consistent with the broader view that AI governance is not solely a technical problem but involves organisational, institutional and regulatory arrangements surrounding the use of algorithmic systems (Yeung, 2018; Tuge and Msweli, 2026; Lioliou et al., 2026).
At the same time, the framework has important limitations. Human judgement is not inherently superior to algorithmic judgement, governance controls can create costs and unintended consequences, and the effectiveness of human oversight is difficult to measure. These limitations provide an important basis for further research.
This chapter therefore considers the broader implications of judgement-centred AI governance, identifies its limitations and proposes directions for future research.
7.1 Theoretical Implication: From Model Governance to Decision Governance
The first theoretical implication is a shift in the unit of analysis. Much of the AI governance literature focuses on the AI system itself: its data, architecture, performance, fairness, robustness, interpretability and technical reliability. These dimensions remain essential. Explainability research, for example, has developed extensive approaches for understanding the behaviour of otherwise opaque models (Adadi and Berrada, 2018; Guidotti et al., 2018; Barredo Arrieta et al., 2020). Research in finance similarly identifies explainability, fairness, transparency and model reliability as important dimensions of responsible AI deployment (Černevičienė and Kabašinskas, 2024; Tuge and Msweli, 2026). However, these dimensions do not fully explain whether an AI-assisted banking decision is legitimate or appropriate.
The framework developed in this thesis therefore proposes that the relevant unit of analysis should often be the AI-assisted decision process. This process includes:
the AI system;
the information supplied to it;
the output it produces;
the human who interprets the output;
the organisational context in which the decision occurs;
the authority assigned to each actor;
the opportunity to challenge the recommendation; and
the consequences of the final decision.
This perspective complements existing work on algorithmic regulation and accountability. Yeung (2018) demonstrates that algorithmic systems can alter how governance and institutional decision-making are structured, while García-Llorente and Olmeda (2026) emphasise the importance of translating governance principles into institutional arrangements. Lioliou et al. (2026) similarly highlight the organisational construction of accountability and control around algorithmic systems in financial firms. The implication is that AI governance research should examine not only whether an AI system behaves appropriately, but also how its outputs become institutional decisions.
The distinction between AI output and institutional decision is therefore fundamental. An AI system may identify a pattern, estimate a probability, rank alternatives or recommend an action, but these computational functions do not by themselves determine who exercises institutional judgement. Torrecilla-Pinero (2026) develops this distinction in philosophical terms by arguing that AI can assist human deliberation without replacing the human subject who exercises judgement. From this perspective, the relevant governance question is not simply whether a human remains somewhere within the decision process, but whether the institutional arrangement preserves the human capacity to understand, evaluate and contest the AI output and to assume responsibility for the decision that ultimately follows. JCAIG develops a complementary organisational interpretation of this problem by examining the conditions under which human judgement remains practically possible, consequential and accountable within AI-assisted decision processes.
This represents a conceptual distinction between model governance and decision governance. Model governance asks whether the system is appropriately developed, validated, monitored and controlled. Decision governance additionally asks how the system's output is interpreted, challenged and converted into institutional action, and whether the actors responsible for that action retain sufficient authority and capability to exercise independent judgement. The latter is particularly important where AI outputs influence consequential banking decisions. The distinction does not imply that model governance is insufficient or unnecessary. Rather, it suggests that technical governance is one component of a broader socio-technical decision process. A technically reliable model can still be used inappropriately, while an appropriately designed human review process can potentially identify contextual circumstances that are not captured by the model itself (Baum et al., 2022; Barredo Arrieta et al., 2020).
7.2 Theoretical Implication: Human Oversight as a Socio-Technical Capability
A second implication concerns the meaning of human oversight.
The phrase “human in the loop” can imply that the presence of a human is itself a meaningful safeguard. This thesis challenges that assumption. The literature on human–AI interaction demonstrates that humans may accept or rely on algorithmic recommendations in ways that are affected by automation bias, trust, cognitive load and interaction conditions (Alon-Barkat and Busuioc, 2023; Kupfer et al., 2023; Romeo and Conti, 2026).
Human oversight should instead be understood as a socio-technical capability consisting of several interdependent conditions:
competence;
relevant information;
sufficient time and attention;
authority to intervene;
independence to challenge; and
accountability for the resulting decision.
These six conditions constitute the analytical synthesis developed by this thesis rather than a framework taken directly from a single source. They draw on the literature concerning explainability, human–AI interaction, automation bias, accountability and algorithmic governance (Adadi and Berrada, 2018; Baum et al., 2022; Alon-Barkat and Busuioc, 2023; Yeung, 2018). Torrecilla-Pinero (2026) makes a related distinction between merely retaining a human within the causal chain and preserving the human as the subject of the decision. From this perspective, meaningful oversight requires more than attribution: it requires the practical conditions under which a human can perceive relevant circumstances, deliberate about the available evidence, question the AI output and assume responsibility for the resulting decision.
Removing any one of these conditions can weaken the effectiveness of oversight. For example, a reviewer may possess formal authority to override an AI recommendation but lack sufficient information to understand why the recommendation was produced. Conversely, a reviewer may understand the recommendation but lack sufficient time or organisational authority to investigate and challenge it. Baum et al. (2022) are particularly relevant here because they argue that meaningful human responsibility requires sufficient epistemic access to the basis of an AI-supported recommendation rather than merely formal human presence.
This perspective is also consistent with the broader literature on explainability and human–AI interaction. Tuge and Msweli (2026) identify organisational, cultural and skills-related challenges in implementing explainable AI in banking, while Romeo and Conti (2026) emphasise the importance of human reliance, trust and interaction conditions. Kupfer et al. (2023) further demonstrate that organisational and technological conditions can affect the extent to which people verify AI-supported recommendations.
Human oversight should therefore be treated as an organisational capability that requires investment, measurement and continuous improvement. This reframes oversight from a procedural requirement into a capability that an institution must actively maintain.
7.3 Implication for Banking Strategy
The framework also has strategic implications for banks.
AI adoption is often justified through improvements in efficiency, scalability, consistency and cost reduction. These benefits remain important, and financial-sector research identifies applications of AI across areas including risk assessment, fraud detection, customer services and other data-intensive activities (Černevičienė and Kabašinskas, 2024; Sailer, 2026; Tuge and Msweli, 2026). However, competitive advantage may increasingly depend on whether institutions can combine AI capabilities with strong decision governance.
A bank that adopts AI rapidly but cannot reliably detect or correct inappropriate outputs may accumulate operational, compliance, customer and reputational risks. The potential consequences of algorithmic error are particularly important in financial services because automated systems may operate at substantial scale and can influence access to financial resources and services (Das et al., 2023; García-Llorente and Olmeda, 2026).
Conversely, a bank capable of deploying AI while preserving high-quality institutional judgement may obtain a more sustainable advantage. This proposition should be understood as a strategic implication of the framework rather than as an empirically established causal relationship.
This suggests that AI governance should not be positioned solely as a defensive function.
Effective governance can support innovation by creating controlled conditions under which employees can experiment, learn and scale successful applications. Risk-sensitive approaches to algorithmic governance similarly suggest that governance arrangements should be proportionate to the potential consequences of the system's use rather than uniformly applied across all applications (Yeung, 2018; García-Llorente and Olmeda, 2026).
Governance therefore has both a risk-reduction function and an innovation-enabling function.
The strategic challenge is consequently not to maximise either automation or human intervention. It is to establish conditions under which AI can be deployed at an appropriate level of autonomy while preserving sufficient institutional capacity to identify and respond to failure.
7.4 Implication for the Role of Employees
The increasing use of AI also changes the nature of professional work.
If AI takes over routine analytical and documentation tasks, human employees may spend a greater proportion of their time evaluating exceptions, interpreting ambiguous evidence and making consequential judgements. This is a plausible implication of increasing AI adoption rather than a universal prediction, since the effect will depend on how individual organisations redesign workflows.
This creates an important paradox.
As AI becomes better at routine tasks, the remaining human tasks may become more difficult rather than less important.
Employees may increasingly encounter cases where information is incomplete, evidence conflicts or circumstances fall outside the assumptions under which the AI system was developed or validated. Explainability and human–AI interaction research suggests that understanding model limitations and contextualising outputs are important components of responsible use (Barredo Arrieta et al., 2020; Tuge and Msweli, 2026).
The required skill profile may therefore shift from simple information processing toward:
critical evaluation;
contextual reasoning;
exception handling;
communication;
ethical judgement;
challenge and escalation; and
understanding of AI limitations.
These capabilities are consistent with the thesis's conception of meaningful human oversight, but the specific effects on professional roles require empirical investigation.
This has implications for recruitment, training and career development within banking. Training should not focus solely on how employees operate AI tools. It should also develop the ability to recognise uncertainty, identify limitations, assess whether outputs are appropriate to the case and determine when escalation or intervention is required (Baum et al., 2022; Kupfer et al., 2023).
The objective should not be to create employees who compete with AI on computational efficiency. Rather, organisations should develop employees who can recognise where computational outputs require contextual interpretation.
7.5 Implication for Professional Expertise
The framework also challenges a simplistic distinction between technical expertise and business expertise.
AI governance requires interaction between both.
Technical specialists may understand model behaviour but lack detailed knowledge of the operational context in which an output is used. Business specialists may understand customers and processes but lack knowledge of model limitations. This creates a potential organisational boundary problem in which no single actor possesses all the knowledge required to evaluate the AI-assisted decision.
Effective governance therefore requires boundary-spanning expertise.
This can be supported through multidisciplinary governance teams involving:
business specialists;
data scientists;
model-risk professionals;
compliance experts;
legal specialists;
operational-risk teams;
information-security professionals; and
internal audit.
The precise composition will depend on the application and its risk profile. The underlying principle is that technical and contextual knowledge should be connected rather than treated as separate governance domains.
Tuge and Msweli (2026) identify organisational and interdisciplinary challenges as important barriers to effective XAI implementation. Their findings support the broader argument that explainability cannot be separated from organisational context and stakeholder requirements. Judgement-centred governance responds to this problem by treating cross-functional collaboration as a core governance requirement rather than merely an optional coordination mechanism.
This also reinforces the distinction between model-level expertise and decision-level expertise. The former concerns how an AI system behaves; the latter concerns whether its output is appropriate for the particular institutional decision being made.
7.6 Implication for Boards and Senior Management
The framework also changes what senior management and boards should expect from AI governance.
A board should not be satisfied simply because management reports that AI systems have been approved, validated or subjected to human oversight. Such evidence demonstrates the existence of governance processes but does not necessarily establish their effectiveness.
More meaningful questions include:
Which AI systems influence consequential decisions?
What level of human judgement is required?
How do we know reviewers understand the outputs?
How often are AI recommendations challenged?
What happens when reviewers disagree with the system?
Are employees genuinely empowered to override it?
Which AI-related incidents have occurred?
What has changed as a result?
Can consequential decisions be reconstructed?
Which risks are increasing as AI usage scales?
These questions shift governance discussions from documentation to evidence.
This is consistent with the accountability-oriented perspective advanced by García-Llorente and Olmeda (2026), which emphasises the importance of translating governance principles into institutional practices. It is also consistent with the accountability and responsibility concerns identified by Baum et al. (2022) and Lioliou et al. (2026).
The proposed questions should therefore be understood as a practical extension of the JCAIG framework rather than as an established board-level checklist from the literature. Their purpose is to shift senior oversight from asking whether controls exist to asking whether those controls produce evidence of effective institutional control.
7.7 Implication for Regulators and Supervisors
The framework has significant implications for financial regulators and supervisors.
Traditional supervisory approaches may focus on whether an institution has appropriate policies, documentation, model validation and risk controls. These remain important but may not reveal whether human oversight functions effectively in practice.
Supervisors could increasingly assess evidence of actual governance performance. Such an approach would be consistent with broader movement toward risk-sensitive and accountability-oriented forms of algorithmic governance (Yeung, 2018; García-Llorente and Olmeda, 2026).
Relevant supervisory indicators might include:
decision reconstruction capability;
reviewer competence;
override patterns;
escalation behaviour;
independent challenge;
customer contestability;
incident response;
model and workflow changes following failures; and
evidence that governance requirements are adapted to the risk of the application.
These indicators are proposed by this thesis as possible operational measures of effective governance. They should not be interpreted as a regulatory standard already established by the cited literature.
This approach would move supervision closer to outcome-oriented governance. Rather than asking only whether an institution has formally complied with a governance requirement, supervisors could examine evidence concerning whether the requirement functions as intended.
It would also recognise that the same AI technology can present very different risks depending on how it is integrated into an institution's processes. A model used to summarise internal documents should not necessarily receive the same governance treatment as a system influencing credit access, financial-crime decisions or customer eligibility. This is consistent with the context-sensitive approach to algorithmic governance developed in Yeung (2018) and the banking-specific risk perspective of García-Llorente and Olmeda (2026).
7.8 Implication for Explainability
The framework also contributes to the debate concerning explainable AI.
Explainability is often presented as a governance solution because explanations can make algorithmic decisions more transparent. The XAI literature demonstrates the importance of explanation for understanding complex models and supporting different users of AI systems (Adadi and Berrada, 2018; Guidotti et al., 2018; Barredo Arrieta et al., 2020).
However, this thesis argues that the governance value of an explanation depends partly on what the recipient is expected to do with it.
An explanation that merely describes why a model produced an output may have limited governance value if the reviewer cannot determine:
whether the output is appropriate;
whether relevant evidence is missing;
whether the case falls outside normal conditions;
whether the model may be unreliable; or
whether the decision should be challenged.
Accordingly, explainability should increasingly be evaluated according to decision usefulness.
This proposition builds on, rather than rejects, the existing XAI literature. Adadi and Berrada (2018) and Barredo Arrieta et al. (2020) emphasise that explanation methods need to be considered in relation to their users and purposes. Baum et al. (2022) further provide a particularly relevant basis for the argument that explanations can contribute to meaningful human responsibility when they provide the epistemic access required for evaluation.
This does not diminish the importance of technical XAI. Rather, it places technical explanations within a broader governance process.
As Tuge and Msweli (2026) argue, the implementation of XAI requires attention to organisational context, stakeholder needs and the practical integration of explanations into decision processes. The implication is therefore a progression from technical explainability, to meaningful explanation, to decision-relevant explanation.
The latter is an analytical concept developed in this thesis: an explanation should provide the information necessary for the relevant actor to understand the AI contribution sufficiently to determine whether it should be accepted, questioned, escalated or rejected.
7.9 Implication for Accountability
A central implication of the framework is that accountability should follow decision authority rather than technological complexity.
The fact that an AI system generated a recommendation does not make the system accountable for the institutional decision. AI systems are technological artefacts and decision-support mechanisms; institutional responsibility remains distributed among the organisations and people who develop, deploy, supervise and act upon them (Yeung, 2018; Baum et al., 2022; Lioliou et al., 2026).
The relevant governance question is therefore:
Who had the authority, information and responsibility to act when the AI recommendation became a decision?
This question helps prevent accountability from becoming diffused across developers, vendors, model validators and front-line users. Responsibility for AI-supported decisions may be distributed across an organisational network, but distribution should not result in the absence of identifiable responsibility (Baum et al., 2022; Lioliou et al., 2026).
It also highlights the importance of aligning three elements:
authority + capability + responsibility.
This formulation is an analytical synthesis of the thesis rather than a three-part framework established by one source.
If an employee is responsible for a decision but lacks the authority to override the AI system, accountability is structurally weakened. Similarly, if an employee has override authority but lacks the information or competence necessary to exercise it, formal authority does not constitute effective control (Baum et al., 2022). This reinforces the distinction between formal accountability and effective accountability. Formal accountability assigns responsibility. Effective accountability requires that the responsible actor possesses sufficient capacity and authority to exercise that responsibility meaningfully.
Meaningful override authority therefore requires more than the technical ability to disagree with an AI system. The authorised decision-maker must have the practical capacity to exercise independent judgement and to depart from the recommendation when the evidence, circumstances or consequences warrant doing so. This requires sufficient information, competence, time and organisational support, but also an institutional arrangement in which human judgement remains consequential. Otherwise, an override function may exist formally while the AI recommendation remains the effective decision. This concern is consistent with the subject-preserving approach proposed by Torrecilla-Pinero (2026), under which meaningful human control requires the human to remain the subject who judges rather than merely the person who formally endorses an algorithmic output.
7.10 Limitations of the Judgement-Centred Approach
Despite its advantages, the framework has several limitations.
7.10.1 Human judgement is not inherently reliable
The framework should not be interpreted as an argument that human decisions are automatically better than AI decisions.
Humans are subject to cognitive biases, inconsistency, fatigue and incomplete information. Algorithmic systems can provide valuable consistency and identify patterns that humans may overlook. Research on algorithm appreciation and automation bias also demonstrates that human preferences and reliance on algorithmic recommendations are complex rather than uniformly negative (Logg et al., 2019; Alon-Barkat and Busuioc, 2023; Kupfer et al., 2023).
The objective is therefore not to replace algorithmic judgement with human judgement.
It is to determine where human judgement adds necessary contextual, ethical or institutional value.
This is an important limitation of a judgement-centred approach: the mere creation of opportunities for human intervention does not guarantee better outcomes. Human intervention itself requires governance.
7.10.2 Human oversight creates costs
Meaningful review requires time, expertise and organisational resources.
If every AI output requires extensive human examination, some of the efficiency benefits of AI may disappear. This creates an important tension between automation and oversight. The solution cannot simply be to maximise human involvement because governance itself consumes organisational resources.
Risk-proportionate governance is therefore essential (Yeung, 2018; García-Llorente and Olmeda, 2026).
The purpose is not maximum human intervention but appropriate human intervention.
This is one reason why the contextual risk classification pillar is central to JCAIG. Low-consequence applications may require relatively limited intervention, while applications capable of producing significant or difficult-to-reverse consequences may justify stronger review, challenge and escalation arrangements.
7.10.3 Oversight can itself become ritualised
Even well-designed governance mechanisms can become procedural.
Employees may learn to complete review fields without genuinely examining the evidence. Committees may approve systems because approval has become routine. Human oversight may consequently exist formally while providing little substantive challenge. Research on automation bias and AI-supported decision-making provides a basis for concern that the presence of decision-support technology can affect the extent to which individuals critically evaluate recommendations (Alon-Barkat and Busuioc, 2023; Kupfer et al., 2023).
Consequently, the framework must continuously evaluate whether governance controls are functioning substantively rather than merely existing formally.
This limitation is particularly important because it creates a potential governance paradox: adding more formal controls does not necessarily produce more effective governance. Controls may themselves become routine, creating the appearance of scrutiny without the substance of scrutiny.
7.10.4 Decision quality is difficult to measure
There is no single universal metric for decision quality.
An increase in override rates could indicate stronger human challenge, but it could also indicate poor model performance or unnecessary human intervention. Similarly, a low override rate could indicate a highly accurate system or excessive automation bias.
Governance metrics must therefore be interpreted in context.
This is an important limitation for the empirical development of JCAIG. Indicators such as override frequency, review time and acceptance rates are potentially informative, but none should be interpreted in isolation. A governance system should examine patterns and outcomes rather than assume that a particular numerical direction is inherently desirable.
For example, the objective should not be to maximise overrides. The relevant question is whether interventions are appropriately justified and produce better decisions where intervention is warranted.
7.10.5 AI capabilities are evolving
The framework is developed in a period of rapid technological change.
As AI systems become more autonomous, multimodal and capable of performing complex sequences of tasks, the boundary between recommendation and execution may become less clear. Recent work on human–AI hybrid finance similarly points toward a movement from AI as an isolated tool toward more integrated decision systems (Kou et al., 2026).
Governance mechanisms will therefore need to evolve alongside technical capabilities.
The six-pillar structure should consequently be understood as adaptable rather than technologically fixed. The underlying principles—risk sensitivity, traceability, contestability, meaningful intervention, independent challenge and organisational learning—may remain relevant even as the technical architecture of AI changes.
7.11 Future Research: Measuring Meaningful Human Oversight
One of the most important areas for future research is the empirical measurement of human oversight.
Current governance frameworks frequently specify that humans should review AI outputs without establishing how the quality of that review should be evaluated. The distinction developed in this thesis between formal human presence and meaningful human oversight therefore creates a clear empirical research opportunity.
Future studies could develop validated measures of meaningful oversight based on:
reviewer competence;
information availability;
review time;
challenge frequency;
override quality;
escalation behaviour;
decision outcomes; and
post-decision learning.
These measures should not simply count whether intervention occurred. They should investigate whether the reviewer understood the recommendation, considered relevant evidence, identified limitations and exercised intervention appropriately.
Experimental research could compare different workflow designs to determine which conditions produce more effective human challenge. For example, studies could examine whether providing uncertainty information, alternative evidence, structured challenge prompts or additional contextual information affects the quality of human review (Adadi and Berrada, 2018; Baum et al., 2022; Kupfer et al., 2023).
This would help move the literature from normative claims about “human oversight” toward empirically testable governance mechanisms.
A particularly important research question is whether the six conditions proposed in this thesis—competence, information, time and attention, authority, independence and accountability—predict meaningful intervention when considered jointly rather than individually.
7.12 Future Research: The Effect of AI on Professional Expertise
A second research direction concerns how AI changes professional expertise.
If AI performs more routine analytical work, professionals may become less experienced in performing those tasks independently.
This raises a potential deskilling problem.
At the same time, AI may enable employees to focus on more complex cases and develop new forms of expertise. The net effect is therefore uncertain and should not be assumed to be either positive or negative.
Future research should investigate whether AI adoption:
strengthens professional judgement;
weakens independent expertise;
shifts expertise toward exception handling;
changes training requirements; or
creates new forms of human–AI professional competence.
Longitudinal studies would be particularly valuable because these effects may not become visible immediately after implementation.
Such research could also investigate whether professionals retain sufficient underlying knowledge to recognise when an AI system is behaving unexpectedly. This is particularly important for judgement-centred governance because effective challenge depends partly on the ability of employees to recognise that intervention is necessary.
The issue is therefore not simply whether AI changes what employees do, but whether it changes the knowledge base on which meaningful human oversight depends.
7.13 Future Research: Governance Under Increasing AI Autonomy
The framework developed in this thesis assumes that meaningful human judgement remains possible.
Future AI systems may challenge this assumption.
As systems become capable of planning, executing multi-step processes and adapting dynamically, the traditional distinction between “AI recommendation” and “human decision” may become increasingly difficult to maintain (Kou et al., 2026).
Future research should therefore investigate governance models for systems where:
AI selects among multiple actions;
AI executes decisions automatically;
human intervention occurs only at exception points;
multiple AI systems interact; or
AI agents initiate actions across organisational systems.
This may require moving from human-in-the-loop governance toward human-on-the-loop, human-over-the-loop or other forms of supervisory architecture. However, terminology alone does not establish effective governance. A nominal human supervisor may still lack the information, time, competence or authority required to intervene meaningfully.
The key issue will remain whether humans retain meaningful institutional control rather than merely nominal responsibility.
This is particularly important for the JCAIG framework because increasing autonomy may alter the point at which the AI output becomes an institutional action. Future research should therefore examine how traceability, contestability and override authority operate when decisions are executed continuously rather than presented as discrete recommendations.
7.14 Future Research: Organisational Culture and AI Governance
A further research opportunity concerns organisational culture.
Formal governance mechanisms may function differently depending on whether an organisation encourages questioning, transparency and learning. The literature on automation bias and organisational implementation already indicates that human interaction with AI is influenced by organisational and contextual conditions rather than technical design alone (Alon-Barkat and Busuioc, 2023; Kupfer et al., 2023; Tuge and Msweli, 2026).
Future research could examine relationships between:
psychological safety;
performance incentives;
management attitudes;
employee willingness to challenge AI;
incident reporting;
override behaviour; and
governance effectiveness.
This would extend AI governance research beyond technological and regulatory dimensions into organisational behaviour.
Such research is particularly relevant because an employee may possess formal authority to challenge an AI system while simultaneously perceiving that doing so will negatively affect performance evaluations or relationships with management. The existence of authority therefore does not necessarily establish that employees will exercise it.
Empirical research could examine whether organisational environments that reward evidence-based challenge produce different AI governance outcomes from environments in which speed, throughput or acceptance of automated recommendations is prioritised.
7.15 Future Research: Governance Metrics and Early-Warning Indicators
Another important research direction is the development of governance-specific early-warning indicators.
Traditional model monitoring may detect statistical deterioration, but it may not detect declining human engagement. Model performance can remain within acceptable statistical thresholds while the surrounding human–AI decision process changes in ways that reduce effective oversight.
Potential early-warning indicators include:
declining review times;
unusually stable acceptance rates;
declining rates of documented challenge;
concentration of overrides among a small number of employees;
increasing reliance on default recommendations;
repeated unexplained exceptions;
growing customer complaints; and
increasing divergence between formal policy and actual workflow behaviour.
These indicators are proposed research variables rather than validated measures of governance failure. Their value will depend on empirical testing.
Research should examine which indicators reliably predict governance failure and under what conditions. It should also investigate interactions between indicators rather than treating them independently.
For example, a declining override rate might be positive if model performance improves and review remains substantive. The same decline could be concerning if review times are simultaneously falling and acceptance rates are becoming unusually uniform. The objective is therefore to identify patterns of weakening oversight, rather than prescribe universal thresholds.
This could allow banks to identify weakening oversight before significant customer, financial or regulatory harm occurs.
7.16 Future Research: Comparative Regulatory Studies
The regulatory environment for AI in banking is developing rapidly and differs across jurisdictions.
Future research should compare how different regulatory systems operationalise concepts such as:
human oversight;
explainability;
accountability;
contestability;
model risk;
documentation; and
independent challenge.
Comparative research could determine whether principle-based approaches, prescriptive requirements or risk-based frameworks are more effective in creating meaningful governance.
The comparative approach of García-Llorente and Olmeda (2026) provides a useful starting point by examining how different jurisdictions structure algorithmic governance in banking. Future work could extend this approach by comparing not only formal regulatory requirements but also actual supervisory practices and institutional outcomes.
This distinction is important because regulatory similarity at the level of formal rules does not necessarily imply similarity in implementation. Two jurisdictions may impose comparable requirements while producing different supervisory expectations or institutional behaviours.
Comparative research could therefore examine the relationship between:
regulatory requirement → supervisory interpretation → institutional implementation → governance outcome.
Such research would provide a stronger empirical basis for determining which regulatory approaches most effectively support meaningful human oversight.
7.17 Future Research: Evaluating the Six-Pillar Framework
The six-pillar framework proposed in Chapter 5 remains conceptual and requires empirical validation.
Future studies should test whether the proposed dimensions—
contextual risk classification;
traceability;
contestability;
meaningful override authority;
independent challenge; and
organisational learning
are sufficient to explain differences in AI governance effectiveness.
These dimensions represent the original analytical structure of JCAIG. The existing literature provides support for their individual importance, but no existing study in the reviewed literature establishes that these six pillars collectively form a sufficient governance framework.
Research could examine whether organisations with stronger performance across these dimensions experience:
fewer significant AI-related incidents;
faster detection of system failures;
better customer outcomes;
stronger regulatory compliance;
more appropriate AI adoption; and
greater organisational learning.
Future research could operationalise each pillar through measurable indicators and test whether the resulting framework predicts governance outcomes across different banking use cases.
The framework could subsequently be refined using empirical evidence.
Importantly, validation should test not only whether the pillars matter individually but also whether they interact. For example, traceability may have greater governance value where contestability and override authority are also present, because recorded information is more useful when an authorised actor can act upon it. Similarly, risk classification may determine the appropriate intensity of the other pillars.
The framework could therefore eventually be tested as an interconnected governance system rather than as six independent controls.
7.18 Towards an Evidence-Based Theory of AI Governance
Taken together, these research directions point toward a broader objective: developing an evidence-based theory of AI governance.
Such a theory would move beyond asking whether a governance principle sounds desirable and instead investigate whether specific governance mechanisms produce better outcomes under particular organisational and technological conditions.
This requires interdisciplinary research combining:
information systems;
organisational behaviour;
financial economics;
law and regulation;
risk management;
human–computer interaction;
machine learning; and
behavioural science.
The complexity of AI governance makes such interdisciplinarity essential.
AI systems are simultaneously technical artefacts, organisational tools, decision-support mechanisms and sources of institutional power. No single disciplinary perspective is sufficient to capture all of these dimensions. The governance problem therefore arises partly from the interaction between technical capability and institutional arrangements (Yeung, 2018; Barredo Arrieta et al., 2020).
An evidence-based theory would ideally explain not simply whether a governance mechanism exists, but when, why and under what conditions it works.
This would also enable research to move beyond binary distinctions such as “human oversight present” versus “human oversight absent”. A more useful empirical framework would examine the quality, intensity and consequences of human involvement in relation to the risk and context of the AI application.
7.19 Final Synthesis
The central argument of this thesis can now be expressed more precisely.
The fundamental governance challenge created by AI in banking is not simply that machines may produce incorrect outputs.
It is that organisations may gradually lose the capacity to recognise when an output is incorrect, inappropriate, incomplete or inconsistent with the broader context of a decision.
This is a fundamentally different risk.
An inaccurate model can potentially be detected through technical validation, performance monitoring and model-risk controls. A weakening of institutional judgement may be much harder to observe because the organisation may continue to meet formal governance requirements while employees increasingly defer to AI-generated recommendations. Research on automation bias provides evidence that human reliance on algorithmic recommendations can occur under particular conditions, although such effects are not universal (Alon-Barkat and Busuioc, 2023; Kupfer et al., 2023).
The response is not to reject AI or to assume that humans should always make better decisions.
Instead, banks should design AI systems and organisational processes so that the strengths of machines and humans are complementary.
AI can provide speed, scale, pattern recognition, consistency and information synthesis. These capabilities are among the reasons AI has significant potential across banking and financial services (Černevičienė and Kabašinskas, 2024; Sailer, 2026).
Humans can provide contextual interpretation, responsibility, ethical reasoning, exception handling and institutional judgement. The value of human involvement, however, depends on whether the organisational conditions necessary for meaningful intervention are preserved (Baum et al., 2022).
Effective governance must therefore determine where these capabilities should interact and ensure that neither becomes an unexamined substitute for the other.
The resulting principle is not human versus AI, but human judgement with AI capability under institutional control.
7.20 Conclusion
This thesis has argued that the future of AI governance in banking should be understood as a problem of decision quality and institutional judgement, rather than solely as a problem of model quality.
The distinction between AI output and institutional decision is therefore fundamental.
A technically accurate, explainable and validated AI system can still produce governance failures if its output is accepted without sufficient scrutiny, applied outside its intended context or treated as possessing authority that it does not legitimately hold (Baum et al., 2022; Alon-Barkat and Busuioc, 2023; Tuge and Msweli, 2026).
The judgement-centred framework proposed in this thesis responds to this challenge through six interconnected pillars: contextual risk classification, traceability and reconstructability, contestability, meaningful human override authority, independent challenge and organisational learning.
Chapter 6 demonstrated how these principles can be operationalised through risk-proportionate workflows, role-specific training, controlled experimentation, human–AI interaction monitoring and lifecycle governance.
The broader implications are significant.
For banks, AI governance should become an organisational capability rather than a compliance checklist. For employees, the increasing automation of routine tasks may make critical judgement and exception handling more important. For boards, governance should focus increasingly on evidence that human oversight actually works. For regulators, effective supervision may require greater attention to governance outcomes and institutional behaviour rather than policies alone (Yeung, 2018; García-Llorente and Olmeda, 2026).
At the same time, the framework should not be interpreted as a rejection of automation or an assumption that human judgement is inherently superior. Its purpose is to preserve the conditions under which humans and AI can contribute their respective strengths while maintaining clear institutional responsibility.
The ultimate test of AI governance is therefore not whether a bank can demonstrate that a human was present somewhere in the decision process.
It is whether the institution can demonstrate that meaningful judgement remained possible, was actually exercised where required, and could change the outcome when the evidence demanded it.
In this sense, the most important question for AI governance is not simply whether the system works.
It is whether the organisation remains capable of knowing when it should not trust the system.
8. Conclusion
The increasing adoption of artificial intelligence in banking is transforming how financial institutions process information, assess risk and make decisions. AI can provide substantial benefits in speed, scale, consistency, pattern recognition and information synthesis across a wide range of banking activities. Yet these capabilities do not remove the need for institutional judgement. As AI becomes more deeply embedded in consequential decisions, the ability of organisations to recognise, evaluate and respond appropriately to problematic outputs becomes increasingly important.
This paper has argued that the central governance problem is therefore broader than AI accuracy. A technically accurate, robust, fair or explainable system does not automatically produce a high-quality institutional decision. The quality and legitimacy of the resulting decision also depend on how AI outputs are interpreted, contextualised, challenged and incorporated into organisational processes. The distinction between AI output and institutional decision is consequently fundamental.
This distinction exposes a limitation in the conventional concept of the “human in the loop”. Human presence does not necessarily constitute meaningful oversight. A reviewer who lacks relevant information, sufficient competence, adequate time, independence or genuine authority to disagree with an AI recommendation may provide formal supervision without exercising substantive control. Conversely, meaningful oversight requires organisational conditions in which challenge is possible, supported and capable of influencing the final decision. The governance question is therefore not simply whether a human is present, but whether meaningful human judgement remains possible.
The paper has consequently proposed a shift from output-centred AI governance to judgement-centred AI governance. Rather than treating governance primarily as a set of controls surrounding an AI model, the proposed Judgement-Centred AI Governance (JCAIG) framework treats the AI-assisted decision process as the central object of governance. This shifts attention from the isolated question of whether an AI system works to the broader question of how its outputs acquire institutional significance and authority.
The framework comprises six interconnected pillars. Contextual risk classification determines the intensity and nature of governance according to the consequences, uncertainty, scale and reversibility associated with an AI application. Traceability and reconstructability ensure that consequential AI-assisted decisions can be understood and examined after the event. Contestability provides mechanisms through which AI outputs can be questioned and reconsidered. Meaningful human override authority ensures that human oversight represents genuine decision capacity rather than procedural approval. Independent challenge provides a counterweight to commercial, operational or technological pressures that might otherwise weaken scrutiny. Finally, organisational learning ensures that incidents, overrides, complaints, exceptions and unexpected outcomes can inform subsequent improvements in systems, workflows and governance.
The six pillars are mutually reinforcing rather than independent controls. Risk classification determines how much governance a particular application requires; traceability provides the evidence needed to understand and challenge decisions; contestability creates an opportunity for challenge; override authority determines whether challenge can alter an outcome; independent challenge reduces the risk of unchecked organisational or commercial influence; and organisational learning ensures that experience is translated into improved governance. Together, these mechanisms transform human oversight from a procedural requirement into a substantive institutional capability.
The operational analysis further demonstrated that these principles cannot remain abstract. They must be incorporated into workflow design, role-specific training, decision interfaces, escalation procedures, monitoring arrangements and organisational incentives. Controlled experimentation provides a means of testing how employees actually interact with AI rather than assuming that formal oversight will operate as intended. Similarly, monitoring human–AI interaction alongside technical model performance can provide evidence of governance weaknesses that conventional model monitoring may not reveal. The objective is therefore not simply to establish controls, but to determine whether those controls work in practice.
This has implications for how banks, boards and regulators understand AI governance. For banks, governance should be treated as an organisational capability rather than a compliance checklist. For employees, increasing automation of routine tasks may make contextual reasoning, exception handling and critical evaluation more important rather than less. For boards and senior management, the relevant evidence is not merely that systems have been approved or reviewers assigned, but that consequential decisions can be reconstructed, recommendations can be challenged and employees are genuinely able to intervene. For regulators and supervisors, this suggests greater attention to evidence of effective oversight, contestability, independent challenge and organisational learning alongside formal policies and controls.
At the same time, the paper does not argue that human judgement is inherently superior to algorithmic judgement. Humans are themselves subject to bias, inconsistency, fatigue and incomplete information, while AI systems can provide valuable analytical consistency and identify patterns that humans may overlook. The purpose of JCAIG is therefore not to replace machine judgement with human judgement. It is to establish an institutional relationship in which the respective strengths and limitations of humans and AI are recognised and governed appropriately.
This also explains why the framework is explicitly risk-proportionate. Meaningful governance does not require maximum human intervention in every AI-assisted activity. Excessive intervention can undermine the efficiency and scalability that make AI valuable, while insufficient intervention can leave consequential decisions vulnerable to automation bias, contextual error or inappropriate reliance. The appropriate objective is therefore not maximum human involvement, but sufficient human judgement in proportion to the potential consequences of error.
The principal contribution of this paper is consequently conceptual and operational. Conceptually, it reframes AI governance as a problem of decision governance, extending attention beyond model properties to the institutional process through which AI outputs become decisions. Operationally, it provides a six-pillar framework for preserving meaningful human judgement through risk classification, traceability, contestability, override authority, independent challenge and organisational learning. The framework also identifies an empirical research agenda: its effectiveness must ultimately be tested rather than assumed.
The limitations of the approach reinforce this need for further research. Human oversight is costly and difficult to measure; human judgement is not inherently reliable; governance controls can themselves become ritualised; and changing AI capabilities may increasingly blur the distinction between recommendation and execution. Future research should therefore examine how meaningful oversight can be measured, how AI affects professional expertise, how organisational culture influences willingness to challenge AI, how governance can operate under increasing autonomy, and whether the six JCAIG pillars predict better institutional and customer outcomes across different banking contexts.
The broader argument can therefore be stated simply.
The fundamental risk of AI in banking is not only that an algorithm may produce an incorrect output. It is that an organisation may gradually lose the capacity to recognise when that output is incorrect, inappropriate, incomplete or inconsistent with the wider context of the decision. Technical validation can identify many forms of model failure. It is less capable of identifying a gradual erosion of institutional judgement in which employees continue to comply with formal governance requirements while increasingly deferring to AI recommendations.
Responsible AI governance must therefore preserve both technical capability and institutional judgement.
AI can provide speed, scale, consistency, pattern recognition and analytical power. Humans can provide contextual interpretation, responsibility, ethical reasoning, exception handling and institutional judgement. Effective governance must determine where these capabilities should interact, how much authority should be delegated to AI, and when human intervention must remain capable of changing the outcome.
The ultimate test of AI governance is therefore not whether a bank can demonstrate that a human was present somewhere in the decision process. It is whether the institution can demonstrate that meaningful judgement remained possible, was actually exercised where required, and could change the outcome when the evidence demanded it.
In this sense, the central principle of judgement-centred AI governance is both simple and demanding: AI may inform a decision, but it should not silently acquire the authority to make that decision. This conclusion is consistent with the emerging philosophical argument that AI may legitimately extend human deliberation while the exercise of judgement should remain with a responsible human subject (Torrecilla-Pinero, 2026). JCAIG does not, however, assume that human judgement is inherently superior to algorithmic judgement. Its concern is institutional: where decisions carry meaningful consequences, organisations should preserve the capacity for a responsible human actor to understand the basis of an AI recommendation, assess its appropriateness in context, contest it where necessary and change the resulting decision when warranted. The objective is therefore not to prevent automation, but to prevent the substitution of optimisation for institutional judgement where human responsibility remains necessary.
References
Adadi, A. and Berrada, M. (2018) ‘Peeking inside the black-box: A survey on explainable artificial intelligence (XAI)’, IEEE Access, 6, pp. 52138–52160. doi: 10.1109/ACCESS.2018.2870052.
Alon-Barkat, S. and Busuioc, M. (2023) ‘Human–AI interactions in public sector decision making: “Automation bias” and “selective adherence” to algorithmic advice’, Journal of Public Administration Research and Theory, 33(1), pp. 153–169. doi: 10.1093/jopart/muac007.
Barredo Arrieta, A., Díaz-Rodríguez, N., Del Ser, J., Bennetot, A., Tabik, S., Barbado, A., Garcia, S., Gil-Lopez, S., Molina, D., Benjamins, R., Chatila, R. and Herrera, F. (2020) ‘Explainable artificial intelligence (XAI): Concepts, taxonomies, opportunities and challenges toward responsible AI’, Information Fusion, 58, pp. 82–115. doi: 10.1016/j.inffus.2019.12.012.
Baum, K., Mantel, S., Schmidt, E. and Speith, T. (2022) ‘From responsibility to reason-giving explainable artificial intelligence’, Philosophy & Technology, 35, Article 12.
Černevičienė, J. and Kabašinskas, A. (2024) ‘Explainable artificial intelligence (XAI) in finance: A systematic literature review’, Artificial Intelligence Review, 57(8), Article 216. doi: 10.1007/s10462-024-10854-8.
Chen, C., Ibekwe-SanJuan, F. and Hou, J. (2010) ‘The structure and dynamics of cocitation clusters: A multiple-perspective cocitation analysis’, Journal of the American Society for Information Science and Technology, 61(7), pp. 1386–1409. doi: 10.1002/asi.21309.
Das, S., Stanton, R. and Wallace, N. (2023) ‘Algorithmic fairness’, Annual Review of Financial Economics, 15, pp. 565–593. doi: 10.1146/annurev-financial-110921-125930.
García-Llorente, C. and Olmeda, I. (2026) ‘Algorithmic governance in banking: A comparative analysis of risk-based and accountability-oriented oversight’, Journal of Banking Regulation, 27, Article 19. doi: 10.1057/s41261-026-00320-6.
Guidotti, R., Monreale, A., Ruggieri, S., Turini, F., Giannotti, F. and Pedreschi, D. (2018) ‘A survey of methods for explaining black box models’, ACM Computing Surveys, 51(5), Article 93. doi: 10.1145/3236009.
Kamm, A. (2026) ‘AI Governance in Banking: The Most Dangerous Sentence Is “Looks Good to Me”’, Digital Age, 1 September.
Kou, G., Li, Y., Wang, H. et al. (2026) ‘Human–AI hybrid finance: From AI tools to decision systems’, Financial Innovation, 12, Article 125. doi: 10.1186/s40854-026-00941-w.
Kupfer, C.K., Prassl, R., Fleiß, J., Malin, C., Thalmann, S. and Kubicek, B. (2023) ‘Check the box! How to deal with automation bias in AI-based personnel selection’, Frontiers in Psychology, 14, Article 1118723. doi: 10.3389/fpsyg.2023.1118723.
Lioliou, E., Seitanidi, M.M. and Stadtler, L. (2026) ‘Governance under algorithmic opacity: How financial firms construct accountability and control around AI in risk disclosures’, Information Systems Frontiers. doi: 10.1007/s10796-026-10762-y.
Logg, J.M., Minson, J.A. and Moore, D.A. (2019) ‘Algorithm appreciation: People prefer algorithmic to human judgment’, Organizational Behavior and Human Decision Processes, 151, pp. 90–103. doi: 10.1016/j.obhdp.2018.12.005.
Lui, A.T., Lamb, G. and Durodola, L. (2025) ‘A right to explanation for algorithmic credit decisions in the UK’, Law, Innovation and Technology, 17(1), pp. 289–317. doi: 10.1080/17579961.2025.2469352.
Nagpal, G.K. and Cotte, J. (2026) ‘Human in the loop, or perceived oversight? The psychological inference that drives AI credibility’, Psychology & Marketing, 43(9), pp. 2305–2316. doi: 10.1002/mar.70174.
Nair, R. P. (2026). “From Human-in-the-Loop to Human-on-the-Hook: Rethinking Organisational Accountability for Agentic AI.”
Page, M.J., McKenzie, J.E., Bossuyt, P.M., Boutron, I., Hoffmann, T.C., Mulrow, C.D., Shamseer, L., Tetzlaff, J.M., Akl, E.A., Brennan, S.E., Chou, R., Glanville, J., Grimshaw, J.M., Hróbjartsson, A., Lalu, M.M., Li, T., Loder, E.W., Mayo-Wilson, E., McDonald, S., McGuinness, L.A., Stewart, L.A., Thomas, J., Tricco, A.C., Welch, V.A., Whiting, P. and Moher, D. (2021) ‘The PRISMA 2020 statement: An updated guideline for reporting systematic reviews’, BMJ, 372, Article n71. doi: 10.1136/bmj.n71.
Romeo, G. and Conti, D. (2026) ‘Exploring automation bias in human–AI collaboration: A review and implications for explainable AI’, AI & Society, 41, pp. 259–278. doi: 10.1007/s00146-025-02422-7.
Rudin, C. (2019) ‘Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead’, Nature Machine Intelligence, 1(5), pp. 206–215. doi: 10.1038/s42256-019-0048-x.
Sailer, L. (2026) ‘Explainable artificial intelligence in finance: Mapping business problems and technological solutions’, Finance Research Open, 2(2), Article 100128. doi: 10.1016/j.finr.2026.100128.
Torrecilla-Pinero, J.A. (2026) ‘Judgment Cannot Be Delegated: A Subject-Preserving Framework for AI Governance’, Philosophy & Technology, 39, Article 135. doi: 10.1007/s13347-026-01140-2.
Tuge, N.B. and Msweli, N.T. (2026) ‘Explainable artificial intelligence in the banking sector: A systematic literature review’, Applied Artificial Intelligence, 40(1), Article e2676353. doi: 10.1080/08839514.2026.2676353.
Yeung, K. (2018) ‘Algorithmic regulation: A critical interrogation’, Regulation & Governance, 12(4), pp. 505–523. doi: 10.1111/rego.12158.
Contact
Reach out via email for inquiries.
Subscribe to newsletter
info@grcadvisory.ch
© 2025. All rights reserved.