Next Article in Journal
Exploring a Potential Capability–Adoption Gap in Digital Financial Adoption Intentions: Evidence from a Fragile Post-Crisis Economy
Previous Article in Journal
The Fairness Illusion? A Cross-Dataset Audit of Accuracy and Demographic Bias in Credit Scoring Based on Machine Learning
Previous Article in Special Issue
Implementation Risk in Accrual Accounting Reform: Internal Audit Readiness and the Risk of Institutional Decoupling in the Saudi Public Sector
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Professional Judgment and AI Disclosure Governance in Audit and Sustainability Assurance: Public Evidence from the UK Big Four

by
Radosveta Krasteva-Hristova
Department of Accounting, Tsenov Academy of Economics, 5250 Svishtov, Bulgaria
J. Risk Financ. Manag. 2026, 19(9), 675; https://doi.org/10.3390/jrfm19090675
Submission received: 22 July 2026 / Revised: 22 August 2026 / Accepted: 31 August 2026 / Published: 3 September 2026
(This article belongs to the Special Issue Accounting and Auditing in the Age of Sustainability and AI)

Abstract

Artificial intelligence (AI) is entering audit workflows while sustainability reporting expands the evidence subject to professional evaluation. This exploratory study examines how the UK Big Four publicly describe safeguards that keep AI-assisted work human-led, reviewable and accountable. The complete 2024 transparency-report cross-section was coded against seven pre-specified dimensions and summarized in an AI–Judgment Governance Disclosure Index (AI-JGDI). The index measures AI disclosure governance—the completeness of public accountability commitments—not internal control effectiveness. Firm evidence is reported with page-level passages and a decision-level coding log; the single-coder design remains a substantive limitation. The results are compared only as external context with recent inspection outcomes published by the UK Financial Reporting Council. All four firms disclose deployed AI capabilities and retained human responsibility; disclosure is most complete for oversight, accountability and learning, and least complete for AI-specific validation and engagement-level traceability. Sensitivity analysis supports the broad cross-firm pattern but not a precise ranking. The audit evidence is empirical; the sustainability-assurance extension is analytical. Across the four reports, AI disclosure governance and sustainability-assurance disclosures remain largely parallel, with no explicit engagement-level methodological link identified.

1. Introduction

Financial reporting and assurance are becoming simultaneously more data-intensive and more judgment-intensive. Artificial intelligence can search contracts, identify anomalous transactions, score populations of journal entries, retrieve technical guidance, draft working papers and direct attention toward unusual patterns. Sustainability reporting adds heterogeneous operational, environmental and value-chain data, forward-looking assumptions, double-materiality assessments and information that may not have passed through financial-reporting control systems. These developments enlarge the evidence available to accountants and auditors, but they also increase the need to decide which evidence is relevant, reliable and sufficiently persuasive.
Professional judgment is therefore not a residual activity left after automation. It is the mechanism through which standards, evidence, expertise and ethical responsibility are converted into a defensible conclusion. Accounting judgment research has long shown that performance depends on task knowledge, experience, motivation and the decision environment (Bonner, 1999; Libby & Luft, 1993). Audit research adds professional skepticism: a questioning mind, alertness to contradictory evidence and a critical assessment of audit evidence (Hurtt, 2010; Nelson, 2009). AI changes this environment because a system can structure the evidence presented to the auditor, rank risks, recommend an action or create a first draft that becomes a powerful anchor. The system can improve consistency and coverage, but an apparently precise output can also conceal data limitations, model error, bias or a management-defined objective.
The behavioral evidence is deliberately cautionary. People may avoid algorithms after observing an error (Dietvorst et al., 2015), yet they may also prefer algorithmic advice under other conditions (Logg et al., 2019). In audit settings, reliance depends on task complexity, perceived control and the opportunity to provide input (Commerford et al., 2022, 2024). AI can consequently produce either under-reliance or automation bias. The governance question is how to calibrate reliance: the professional should use the system where it is competent, challenge it where its limits matter, and remain accountable for the conclusion.
Public incidents make the risk concrete, but they also reveal an evidentiary boundary. Deloitte Australia’s Targeted Compliance Framework independent assurance review was republished to correct identified errors after the firm disclosed the use of generative AI in preparing the report (Department of Employment and Workplace Relations [DEWR], 2026); this was an assurance review, not a statutory financial-statement audit. In the closest public audit analogue, the PCAOB found that five EY audits inspected in 2019 lacked sufficient testing of the accuracy and completeness of data and reports used in substantive procedures (PCAOB, 2020). The PCAOB did not attribute those deficiencies to AI. As of August 2026, no public FRC or PCAOB enforcement decision identified by this study establishes AI misjudgment as the cause of a statutory audit failure. Instead, the record documents two failure mechanisms—unverified generated content and reliance on insufficiently tested electronic information—that AI can magnify. This is why the PCAOB amended AS 1105 and AS 2301 to reduce the risk that an auditor using technology-assisted analysis might issue an opinion without sufficient appropriate audit evidence (PCAOB, 2024).
The normative environment is also moving quickly. The IESBA technology-related revisions preserve the fundamental principles of integrity, objectivity, professional competence and due care, confidentiality and professional behavior when technology is used (IESBA, 2023, 2024). ISA 315 (Revised 2019) requires understanding of the entity and its information system and strengthens the exercise of professional skepticism in risk assessment (IAASB, 2019). ISA 540 (Revised) is directly relevant to AI-supported estimates because it emphasizes estimation uncertainty, management bias and skeptical evaluation (IAASB, 2018). The EU Artificial Intelligence Act establishes a risk-based regulatory architecture and, for relevant systems, requirements related to governance, technical documentation, logging, transparency, human oversight, accuracy, robustness and cybersecurity (European Parliament and Council, 2024). The NIST AI Risk Management Framework organizes voluntary risk management around Govern, Map, Measure and Manage (NIST, 2023). None of these instruments transfers an auditor’s responsibility to an algorithm.
At the same time, sustainability reporting and assurance have moved from voluntary practice toward formal requirements. The Corporate Sustainability Reporting Directive and European Sustainability Reporting Standards require extensive sustainability information and governance disclosures in the European Union (European Commission, 2023; European Parliament and Council, 2022). IFRS S1 and IFRS S2 establish an investor-focused global baseline organized around governance, strategy, risk management, and metrics and targets (ISSB, 2023a, 2023b). ISSA 5000 supplies a global, framework-neutral standard for sustainability assurance engagements (IAASB, 2024). AI may help process these data, but sustainability evidence often contains measurement uncertainty, estimates and value-chain information for which explainability and provenance are indispensable.
Existing research offers important insights but leaves three gaps. First, many studies discuss potential benefits and risks without examining what large audit firms publicly disclose about the governance of AI-assisted judgment. Second, corporate reports often combine technology, people, quality and risk narratives; a reproducible coding framework is needed to distinguish a deployed tool from a governance safeguard. Third, the relationship between AI governance and sustainability assurance is frequently asserted but rarely traced in comparable public documents. Recent work calls for research that examines specific AI configurations and the institutional arrangements surrounding them rather than treating AI adoption as a single binary variable (Kokina et al., 2025; Lehner et al., 2022; Stratopoulos & Wang, 2025).
This study asks three research questions:
RQ1. How do the UK Big Four publicly describe the role of AI in audit work and professional judgment?
RQ2. Which safeguards for validation, documentation, data governance, accountability and learning are disclosed, and how complete are these disclosures across firms?
RQ3. To what extent do the transparency reports connect AI disclosure governance evidence with sustainability reporting and assurance?
The empirical setting is the United Kingdom. The Big Four operate under a common regulatory environment, publish transparency reports under comparable requirements and are subject to public inspection by the Financial Reporting Council (FRC). The corpus comprises the complete 2024 transparency-report cross-section for Deloitte, EY, KPMG and PwC UK. A structured content analysis codes seven governance dimensions and constructs an AI–Judgment Governance Disclosure Index (AI-JGDI). The FRC’s 2024 inspection results are used only as external context; the study does not infer that a disclosure score causes an inspection outcome.
The paper makes four contributions. It provides public empirical evidence about how leading audit firms frame human–AI responsibility. It introduces a transparent index whose items can be replicated or extended to other jurisdictions and years. It separates disclosure completeness from actual governance effectiveness and from regulatory audit quality. Finally, it develops a practical human–AI judgment protocol relevant to both financial-statement audit and sustainability assurance. To avoid overstating scope, the empirical evidence concerns AI governance in a financial-statement audit, whereas the treatment of sustainability assurance is conceptual: the protocol is extended analytically to ISSA 5000 engagements rather than tested on sustainability-assurance evidence. Stated precisely, the contribution is to measure the completeness of firms’ public accountability commitments, not to assess the effectiveness of their internal controls.
The theoretical contribution is a refinement rather than a new theory. The study extends the judgment-performance perspective in accounting (Bonner, 1999; Libby & Luft, 1993) and the professional-skepticism literature (Hurtt, 2010; Nelson, 2009) by adding an AI-mediated evidence layer to the judgment process. In this account the professional makes judgments not only about the reporting issue and the underlying evidence, but also about the technological system that structures that evidence—its inputs, assumptions and limits. The AI-mediated layer changes the object of skepticism: calibrated reliance and the reviewability of an AI-influenced conclusion become properties that professional judgment must actively manage. The paper operationalizes this refinement as a measurable disclosure construct (the AI-JGDI) and illustrates it empirically; it is therefore a conceptual refinement combined with a measurement contribution, not merely an empirical illustration.
A distinction is maintained throughout between AI governance and AI disclosure governance. AI governance refers to the internal structures, controls, responsibilities and practices through which a firm manages AI. AI disclosure governance refers to the subset of those arrangements described in public reports. Because the empirical evidence consists exclusively of public documents, the AI-JGDI measures AI disclosure governance rather than the existence or operating effectiveness of internal AI governance arrangements. Empirical claims therefore concern disclosed mechanisms; AI governance is reserved for the broader conceptual and normative framework.

2. Literature Review and Analytical Framework

2.1. Professional Judgment, Discretion and Audit Quality

Professional judgment is a reasoned choice among alternatives under conditions in which rules and evidence do not determine a single answer. It differs from preference because the conclusion must be defensible by reference to applicable requirements, facts, expertise and the objective of the engagement. Judgment is especially important for materiality, risk assessment, accounting estimates, going concern, control evaluation, fraud risk and the sufficiency of evidence. In sustainability assurance, the same logic applies to materiality processes, greenhouse-gas estimates, value-chain information, scenario assumptions and qualitative claims.
Accounting standards create a bounded judgment space. Principles allow economic substance and entity-specific conditions to be reflected, while detailed rules can improve consistency. Neither approach eliminates judgment, and standard precision interacts with preparer incentives, audit-committee strength and auditor challenge (Agoglia et al., 2011; Backof et al., 2016; Bennett et al., 2006; Nelson, 2003). Managerial discretion can communicate private information or can be used opportunistically to select assumptions, timing and disclosures that favor a target (Fields et al., 2001; Healy & Wahlen, 1999). Audit quality therefore depends not only on whether a selected number lies within a plausible range but also on whether evidence was searched neutrally, alternatives were considered and contradictory information was challenged.
Professional skepticism is the behavioral safeguard at this interface. Nelson (2009) distinguishes skeptical judgment from skeptical action: an auditor may recognize risk but still fail to obtain additional evidence or challenge management. AI can support skeptical action by analyzing complete populations and surfacing anomalies. It can also weaken it if the system’s output becomes a substitute for independent evaluation. Judgment quality consequently contains two dimensions: technical defensibility and process integrity. A technically plausible conclusion produced by a biased or opaque process remains vulnerable. These two dimensions are affected asymmetrically by AI. By analyzing complete populations and surfacing anomalies, AI can strengthen skeptical judgment; but the same fluency can weaken skeptical action, because an authoritative-looking output can reduce an auditor’s willingness to perform additional procedures or to challenge a favorable result. A governance design must therefore protect skeptical action specifically, not assume that better inputs to judgment translate into more rigorous follow-up.

2.2. AI-Assisted Audit: Augmentation, Reliance and Accountability

AI in audit includes machine learning, natural-language processing, anomaly detection, intelligent search and generative systems. Early research anticipated that automation would shift work from routine procedure execution toward exception analysis and judgment (Abdullah & Almaqtari, 2024; Kokina & Davenport, 2017; Sutton et al., 2016). More recent field evidence shows wider use but also implementation challenges involving data access, integration with methodology, regulation, skills and accountability (Kokina et al., 2025). The potential contribution is strongest when AI and human capabilities are complementary: machines scale search and pattern recognition, whereas professionals frame the problem, assess context, challenge incentives and accept responsibility.
Reliance is not automatically calibrated. Algorithm aversion can lead users to reject a useful model after observing an error (Dietvorst et al., 2015). Algorithm appreciation can produce the opposite tendency (Logg et al., 2019). In complex audit-estimation tasks, the perceived source and competence of an AI system affect reliance (Commerford et al., 2022). Allowing auditors to provide input can increase reliance and perceived control, but greater confidence is beneficial only if it reflects better understanding of system limitations (Commerford et al., 2024). A governance design should therefore support contestability rather than merely encourage adoption.
Ethical risks arise when responsibility is displaced. Training data can reproduce historical bias; model objectives can privilege efficiency over audit quality; generated text can hallucinate; and proprietary systems can be difficult to explain to audit committees or regulators (Lehner et al., 2022; Murikah et al., 2024). Management or engagement teams may also use technological opacity strategically, accepting outputs that support a preferred conclusion and overriding those that do not. Algorithmic management research shows that algorithmic systems can redistribute control and reshape pre-existing organizational power relations (Kellogg et al., 2020; Jarrahi et al., 2021). In audit, contestable assumptions embedded in an apparently neutral system can shift the locus of judgment away from the accountable professional. This algorithmically mediated discretion is a form of technocratic authority: technical expertise and system opacity can make challenge more difficult and may erode auditor independence unless decision rights, review and override are explicit.
The challenge to judgment also depends on the type of AI. Predictive systems—anomaly scores, risk-assessment and estimation models—raise problems of opacity and bias: their training data and objectives can encode historical distortions that are hard to inspect, so the principal safeguards are validation, explainability and neutral evidence search. Generative systems—drafting and summarization assistants—raise different problems of fabrication (hallucination) and anchoring: a fluent first draft can become a powerful reference point that narrows the alternatives an auditor considers, so the principal safeguards are source traceability, corroboration and substantive review before any generated content enters the working papers. The seven governance dimensions below are intended to hold for both classes, but their emphasis shifts with the use case.
Accountable augmentation requires at least seven linked safeguards. First, the use case must be approved for a defined purpose. Second, the system must be tested or validated for that purpose and population. Third, data provenance, confidentiality and permitted use must be controlled. Fourth, users need an explanation adequate for understanding why the output is relevant and what uncertainty remains. Fifth, professional review and override must be substantive. Sixth, responsibility and escalation routes must be assigned. Seventh, users require both accounting/audit competence and AI literacy. These safeguards form the coding dimensions used in this study. Together, this literature creates two analytical levels. The first concerns AI governance as the internal system of validation, oversight, documentation, data control, accountability and competence. The second concerns AI disclosure governance: the extent to which those arrangements are made visible and assessable in public reporting. The empirical analysis examines the second level and does not infer the effectiveness of the first.

2.3. Normative Requirements for AI-Assisted Judgment

No single standard currently governs every use of AI in accounting and audit. The applicable architecture is layered (Table 1). At the engagement level, ISA 200 retains the auditor’s responsibility to obtain reasonable assurance and exercise professional judgment and skepticism (IAASB, 2009). ISA 315 (Revised 2019) addresses risk identification, information systems and automated controls. ISA 540 (Revised) requires robust work on accounting estimates and management bias. ISA 220 (Revised) and ISQM 1 allocate engagement and firm-level quality-management responsibilities (IAASB, 2020a, 2020b). These requirements apply regardless of whether evidence or analysis is generated manually or technologically.
Ethical requirements add objectivity, competence, confidentiality and appropriate professional behavior. The IESBA technology-related revisions explicitly recognize that technology can create threats to compliance with the fundamental principles and that professional accountants must remain alert to information that may be incomplete, biased or misleading (IESBA, 2023). An organization’s approval of a tool does not eliminate the individual professional’s responsibility to use it competently and question its output.
General AI governance instruments supply more specific control concepts. The NIST AI RMF links governance with contextual mapping, measurement and management of risk (NIST, 2023). The EU AI Act uses a risk classification rather than declaring every accounting application high risk. Nevertheless, its control vocabulary—risk management, data governance, documentation, logging, transparency, human oversight, accuracy, robustness and cybersecurity—provides a useful benchmark for systems that influence consequential professional decisions (European Parliament and Council, 2024). The benchmark is used analytically in this paper, so as not to claim that every audit tool falls within the Act’s high-risk category.
Sustainability standards intensify the need for these controls. CSRD and ESRS broaden the information boundary to impacts, risks and opportunities and require extensive governance and value-chain information (European Commission, 2023; European Parliament and Council, 2022). IFRS S1 and S2 emphasize connected information and decision-useful disclosures (ISSB, 2023a, 2023b). ISSA 5000 applies across sustainability topics, reporting frameworks and both limited- and reasonable-assurance engagements (IAASB, 2024). When AI is used to classify narratives, estimate emissions, screen evidence or draft assurance documentation, provenance and explainability become part of the credibility of the engagement.

2.4. Analytical Model

The study treats AI governance as a conversion system (Figure 1) between technological capability and professional output. AI capability enters an audit task through data, models and interfaces. A professional interprets the output within standards and methodology. Governance controls determine whether that interpretation is informed, reviewable and challengeable. The final judgment then affects audit evidence, documentation and communication. Feedback from review, inspection and incidents should update the system and the organization’s approved-use boundary.
The model predicts neither that more technology necessarily improves quality nor that more disclosure proves effective implementation. It instead identifies observable governance commitments. A public report can provide evidence that a firm recognizes a safeguard and describes a mechanism. It cannot demonstrate that the mechanism operated effectively on every engagement. This distinction defines the empirical claims made below.

3. Materials and Methods

3.1. Research Design and Public-Document Corpus

The study uses comparative qualitative content analysis with a structured ordinal coding instrument. Public organizational documents are appropriate for examining how firms construct accountability, identify risks and communicate governance to investors, audit committees, regulators and other stakeholders. They are not neutral descriptions: firms select what to disclose and may use favorable language. For that reason, the unit measured is disclosure completeness, not latent governance quality. Because transparency reports are self-reported and outward-facing, they are also subject to social-desirability bias: firms have an incentive to present a more favorable picture than internal practice may warrant. This does not negate the study’s value, but it means that a high disclosure score should be read as evidence of a stated commitment rather than of realized control, and the index is interpreted accordingly throughout.
The population was defined as the four largest UK audit firms subject to the FRC’s operational-separation principles and commonly described as the Big Four. The sampling frame was each firm’s report labeled Transparency Report for the reporting period ending in 2024. All four available reports were included; no observation was sampled within the population. The documents are comparable in jurisdiction, broad regulatory purpose and reporting cycle, although fiscal year-ends and report structures differ. The corpus contains 669 PDF pages in total (Table 2).
The reports were downloaded from the firms’ official websites. Searchable text was extracted from the PDF files, and each complete report—not only its technology sections—was searched for AI, artificial intelligence, generative AI, GenAI, machine learning, algorithm, analytics, automation, professional judgment/judgement, professional skepticism/scepticism, validation, testing, pilot, documentation, explainability, security, privacy, data governance, accountability, training, sustainability, ESG, climate, assurance, CSRD, ESRS, emissions, greenhouse gas, value chain and materiality. Each located passage was read in full-page and section context. For RQ3, sustainability-assurance disclosure and AI disclosure governance evidence were coded separately. An integration link was recorded only when a report explicitly connected an AI tool or AI-specific safeguard with a sustainability-assurance procedure, evidence source or judgment; separate coverage of both themes was insufficient. Section 4 presents the principal page-level evidence and short supporting passages, while Appendix A provides the complete decision-level evidence trail.
The FRC’s Annual Review of Audit Quality 2024 was added as a public regulatory source. The FRC reported the percentage of inspected audits assessed as good or requiring no more than limited improvements: Deloitte 94%, EY 76%, KPMG 89% and PwC 76% (FRC, 2024). The inspected audits were a risk-based sample. These percentages are therefore descriptive context and cannot be extrapolated to every audit or treated as a dependent variable in a four-observation causal model.

3.2. Coding Instrument

Seven dimensions were derived deductively from the professional-judgment literature and normative frameworks, then applied consistently to all reports (Table 3). Each dimension received a score of 0, 1 or 2. A score of 0 means that no relevant disclosure was located. A score of 1 means that the report contains a general commitment, adjacent control or limited mechanism. A score of 2 requires an explicit AI- or audit-specific mechanism. The stricter threshold avoids awarding full credit for generic statements about technology or firm-wide security.
The dimensions are intentionally process-oriented. Audit-specific AI integration identifies whether the report moves beyond aspiration to deployed or scaled audit use. Human oversight identifies whether professional judgment remains substantive. Validation and testing identify pre-release or use-case evaluation. Explainability and documentation identify whether an external reviewer can trace how AI affected work. Data governance and security identify controls over data and permitted use. Accountability identifies assigned owners, committees or approval routes. AI-specific learning identifies structured competence development rather than broad technology awareness. Consistent with the distinction established in the Introduction, the AI-JGDI measures AI disclosure governance, not the operating effectiveness of internal AI governance.
Content validity for the seven dimensions rests on their direct correspondence to established normative requirements. AI use and human oversight follow ISA 315 (Revised 2019) on understanding the information system and the IESBA and EU AI Act requirements for meaningful human oversight (IAASB, 2019; IESBA, 2023, 2024; European Parliament and Council, 2024). Validation and testing map to ISA 540 (Revised) on estimation uncertainty and to the accuracy and robustness expectations of the EU AI Act and the Measure function of the NIST AI Risk Management Framework (IAASB, 2018; NIST, 2023). Traceability corresponds to audit-documentation requirements and to the technical-documentation and logging provisions of the EU AI Act. Data governance and security reflect IESBA confidentiality and the data-governance expectations of the same instruments. Accountability follows ISQM 1 and ISA 220 (Revised) on assigned responsibility for quality, and AI-specific learning follows the professional-competence and due-care principle of the IESBA Code. The seven dimensions are thus deduced from, and individually anchored in, the standards the paper interprets rather than assembled ad hoc.

3.3. AI–Judgment Governance Disclosure Index

For firm i, the AI–Judgment Governance Disclosure Index is defined in Equation (1):
A I - J G D I i = 100 14 k = 1 7 s i k ,
where sik is the score (0, 1 or 2) for dimension k. Equal weighting is retained as the primary specification because no validated empirical or normative basis exists for assigning differential importance to the seven dimensions. The individual scores are ordered categories; their normalized sum is used only as a descriptive composite of disclosure completeness, not as a cardinal measure of governance effectiveness. A numerical difference in a given size should therefore not be interpreted as an equivalent substantive difference in internal control quality. Three conceptually motivated alternatives test sensitivity: control-risk emphasis doubles validation, traceability and data-governance weights; human-accountability emphasis doubles human oversight and accountability; and competence emphasis doubles AI learning. The recalculated disclosure profiles are reported with the results.
Coding was conducted by one researcher in two passes. The first pass extracted every potentially relevant passage and recorded the report page. The second applied the pre-specified 0–2 rules and revisited the full page and surrounding section whenever a passage was ambiguous. No second coder was available; therefore, the study does not report inter-coder reliability, and this remains a substantive limitation of the central empirical contribution. The study does not present intra-coder checking as a substitute for independent agreement. Transparency is instead strengthened through the complete 28-decision log in Appendix A, which reports a short source passage, page, score and decision rationale for every firm-by-dimension observation. Independent double-coding remains necessary before the index is used for inference or ranking.
Adjacent-score boundaries were applied consistently across all seven dimensions. AI use receives 2 only for a deployed or scaled audit-specific application and 1 for experimentation or aspiration. Human oversight receives 2 only for explicit review, challenge or override of AI output. Validation/testing receives 2 for AI-specific testing, certification or release criteria and 1 for piloting, approval or general evaluation. Traceability receives 2 for an explicit AI-use documentation, explanation or audit-trail requirement and 1 for general documentation or transparency. Data governance/security receives 2 for AI-specific controls over permitted data, privacy, security or model training and 1 for general technology controls. Accountability receives 2 for a named owner, committee, approval body or escalation route and 1 for general governance. AI learning receives 2 for structured or mandatory AI-specific learning for auditors and 1 for optional or general technology learning. Silence or the absence of a qualifying disclosure receives 0. Appendix A applies these rules to all 28 decisions.

3.4. Analysis and Limitations of Inference

The analysis combines within-case reading, cross-case comparison and descriptive scoring. No inferential statistics are used because the population contains four firms and the item scores are ordered, disclosure-based categories. Dimension-level patterns are reported as frequency counts rather than means and standard deviations. The weighting analysis is a sensitivity exercise, not a validated weighting model. The FRC percentages are shown only as external context; no correlation, causal inference or governance ranking is made. Differences may reflect disclosure strategy, report architecture and communication choices as well as underlying practice.
All empirical inputs are publicly accessible. No interviews, confidential engagement files, personal data or proprietary datasets were used. The study did not evaluate the source code, model performance or engagement-level operation of any named tool. Assertions about systems are therefore attributed to the firms’ public reports.

4. Results

4.1. AI Is Disclosed as Deployed Augmentation, Not Autonomous Judgment

All four reports move beyond a generic expectation that AI may be used in the future. Deloitte describes PairD, AI and machine-learning functions in Omnia and audit pilots for research, document retrieval, first review and document creation (Deloitte, 2024, pp. 8, 90–92). EY reports globally scaled AI integrated with EY Canvas and explicit testing and certification controls (EY, 2024, pp. 36–37). KPMG reports Clara AI chat for all UK auditors and transaction scoring deployed to nearly 900 UK audits (KPMG, 2025, pp. 36–37). PwC reports an Audit GenAI Hub, ChatPwC and approved engagement use cases (PwC, 2024, pp. 111–112).
The common operational model is augmentation. Deloitte states that expertise, professional skepticism and judgment are used to challenge and assure the reliability of output. EY says that AI-enabled technology supports procedures but does not replace the professional’s experience and judgment. KPMG combines smart technology with curious and inquisitive minds and professional skepticism. PwC describes a human-led, technology-powered audit and requires skeptical review of GenAI outputs. These disclosures support a consistent answer to RQ1: public accountability remains attached to the auditor, even where AI is scaled across the practice.
This framing is important because tools perform different functions. An anomaly score reallocates attention; a technical chatbot retrieves or synthesizes guidance; a drafting assistant creates an initial artifact; and a transaction-scoring model analyses a population. None of these functions establishes by itself whether evidence is sufficient or whether an accounting estimate is reasonable. The reports generally recognize this boundary.

4.2. Validation and Traceability Are the Least Consistently Disclosed Safeguards

Validation disclosure varies more than adoption disclosure. EY provides the clearest lifecycle description: technology concepts pass through a global committee; testing with end users, piloting, feedback and certification are prerequisites for release. PwC describes prompt engineering and validation practices in the Audit GenAI Hub. Deloitte reports pilots, use-case approval, risk thresholding and a clearing house process, which demonstrate gatekeeping but provides less detail about performance validation. KPMG reports extensive deployment and responsible user challenge, but the 2024 report does not describe an AI-specific validation or certification mechanism at the same level of detail. Under the strict codebook, this produces scores of 2 for EY and PwC, 1 for Deloitte and 0 for KPMG on the validation dimension.
Explainability is also incompletely disclosed. The reports discuss transparent audit services, documentation and the ability of AI tools to support or improve working papers. However, only PwC explicitly states that clear documentation is required where approved GenAI use cases have been used. None of the reports supplies public model-level information such as performance thresholds, error rates, explainability methods, override frequency or post-deployment drift monitoring. Such details may exist internally and may be inappropriate to disclose fully for security or proprietary reasons. Yet an external reader cannot determine from most reports what minimum explanation must be retained in an engagement file when an AI output materially influences a judgment.
The result identifies a disclosure boundary rather than proving a control deficiency. Public transparency reports are designed for multiple regulatory and stakeholder purposes, not as model cards. Nevertheless, validation and traceability are the dimensions where the difference between a statement of responsible intent and an externally assessable control is greatest.

4.3. Data Governance and Accountability Are More Visible

Deloitte discloses a safe and secure environment for PairD and a firm-wide GenAI risk response involving use-case approval, data use and management, ethical use, cyber risk, a Global Data Council, risk thresholding and a Trustworthy AI framework (Deloitte, 2024, p. 108). EY links responsible technology use with standardized development protocols, a global evaluation committee, information-security policies and privacy-impact assessments for new technology (EY, 2024, pp. 36–37, 60–61, 129–130). KPMG reports Risk Committee deep dives on AI and data risk, secure interaction within Clara, and general information-security governance, but the AI section contains less explicit detail about AI-specific data lineage or permitted-data rules. PwC describes a secure environment, states that ChatPwC does not use prompts or responses to train the underlying model, limits allowable tools and use cases, and requires clear documentation (PwC, 2024, pp. 111–112).
Accountability structures are explicit across the corpus. Deloitte identifies use-case governance and program ownership. EY identifies a global committee involving Professional Practice, the Assurance Quality Network and Technology. KPMG identifies central technology teams and board/Risk Committee oversight. PwC identifies an Audit GenAI Hub with audit subject-matter experts, data scientists and innovation managers. All four therefore receive the maximum accountability score. The reports differ less on whether someone owns AI governance than on what public evidence is supplied about validation outputs and engagement-level traceability.

4.4. Learning Is Treated as a Condition of Responsible Use

Three firms disclose broad AI-specific learning at a level meeting the maximum criterion. Deloitte’s principal-risk response includes an AI-fluency workstream. KPMG reports that all auditors were trained in prompt engineering at its 2024 Audit University so that they could engage with and challenge Clara AI chat responsibly. PwC makes training on GenAI fundamentals and audit business rules mandatory before granting access to ChatPwC. EY discloses AI badges and technology learning, but the 2024 Transparency Report is less explicit about a mandatory audit-wide AI curriculum; it therefore receives a score of 1 under the strict rule.
The emphasis on learning supports a dual-competence model. An auditor needs domain competence to recognize an implausible output and AI literacy to understand data, uncertainty, permitted use and limitations. Prompt skill alone is not professional competence. Conversely, a technically expert accountant who cannot evaluate model limitations may either reject useful evidence or accept output ceremonially. The reports generally treat technology training as complementary to professional skepticism rather than as its replacement.

4.5. Disclosure Scores and External Inspection Context

The cross-firm comparison reveals a common baseline of deployed audit-specific AI, retained human oversight and identifiable accountability structures. The principal differences concern the specificity of validation and testing mechanisms, engagement-level traceability, AI-related data controls and learning arrangements. Table 4 presents the dimension-level coding and the resulting AI–Judgment Governance Disclosure Index (AI-JGDI).
To make the evidence directly visible alongside the scores, Table 5 identifies the named tools and governance mechanisms supporting the most complete disclosures and the specific limitations responsible for scores of 0 or 1.
Because the 0–1–2 items are ordered categories, frequency patterns—not means or standard deviations—are used at the dimension level. All four firms score 2 for AI use, human oversight and accountability. Validation/testing contains one 0, one 1 and two 2s; traceability contains three 1s and one 2; data governance and AI learning each contain one 1 and three 2s. Table 6 tests whether the firm-level disclosure profiles are sensitive to alternative weighting assumptions.
The broad endpoints are stable across the four specifications: PwC retains the most complete public disclosure profile and KPMG the least complete. The relative position of Deloitte and EY is not invariant: EY is higher under control-risk emphasis, whereas Deloitte is higher under competence emphasis. The analysis therefore supports a broad cross-firm disclosure pattern but not a precise ranking. Variation arises primarily from the specificity of validation, traceability, AI-related data controls and structured learning.
The AI-JGDI should therefore be interpreted as a profile of public disclosure completeness rather than as a ranking of actual governance effectiveness or audit quality. A maximum score indicates that explicit public evidence was identified for every coded dimension; it does not demonstrate that the disclosed controls operated effectively across all engagements. Correspondingly, a lower score indicates that the specified mechanism was not located in the transparency report and does not establish that the firm lacks an equivalent internal control. To avoid implying a governance ranking, the firm-level values are read as disclosure profiles rather than league-table positions, and the paper deliberately avoids describing any firm as having stronger or weaker governance on the basis of the index.
Several explanations for the observed differences are plausible, although none can be confirmed with public data and each is offered only to orient future research. First, the differences may reflect strategic disclosure choices: a firm may adopt a more expansive AI-disclosure posture to project market leadership, which would raise its completeness score without necessarily implying superior internal practice. Second, they may reflect report-architecture effects: where relevant material is dispersed across governance, technology and quality sections rather than consolidated, the strict coding rules are less likely to locate an explicit mechanism, which may partly explain a lower score such as KPMG’s on validation and data governance. Third, they may reflect differences in risk culture and communication style, which shape how much operational detail a firm is willing to place in a public document. Distinguishing these explanations would require interviews or engagement-level access and is left to future work.
The FRC’s risk-based inspection results provide an external context for interpreting the disclosure index. As shown in Table 7, the ordering of the inspection outcomes does not correspond to the AI-JGDI ordering. This descriptive mismatch confirms that public AI disclosure governance and regulatory audit-quality inspection capture different aspects of accountability and should not be treated as interchangeable measures.
The comparison does not support a firm-level causal inference. The AI-JGDI measures the completeness of governance mechanisms disclosed in public transparency reports, whereas the FRC percentages reflect the outcomes of risk-based inspections of selected audit engagements. The two indicators differ in their objects of measurement, evidence bases and sampling conditions. Consequently, a more complete AI disclosure governance profile cannot be interpreted as evidence of higher audit quality, just as a lower disclosure score cannot be interpreted as evidence of weaker internal practice.

4.6. AI Disclosure Governance and Sustainability Assurance Remain Parallel Rather than Integrated Narratives

RQ3 was assessed using three pre-specified indicators. Sustainability assurance was coded present when a report discussed methodology, services, competence or engagement responsibilities. AI disclosure governance was coded present when a report disclosed an AI-specific governance mechanism. Integration required an explicit connection between an AI tool or safeguard and a sustainability-assurance procedure, evidence source or judgment.
The evidence shows concrete separation. Deloitte describes AI in Omnia and PairD (pp. 8, 90–92) and separately reports an “enhanced methodology for undertaking sustainability assurance engagements” (p. 94). EY describes AI integrated with Canvas (p. 36) and separately presents the EY Sustainability Assurance Methodology (p. 41). KPMG describes Clara AI chat and transaction scoring (pp. 36–37) and separately a “dedicated ESG Assurance team” (p. 46). PwC describes the Audit GenAI Hub and ChatPwC (pp. 111–112) and separately an “integrated financial and non-financial assurance team” (p. 116). None of these passages applies an AI tool, validation rule, provenance control or documentation requirement to sustainability-assurance evidence. Table 8 summarizes this firm-by-firm evidence and the resulting integration assessment.
The systematic answer to RQ3 is therefore that all four firms disclose sustainability-assurance activity and separate AI disclosure governance evidence, but none provides an explicit public methodological link between them. This matters because sustainability evidence is especially exposed to inconsistent definitions, estimation uncertainty, missing value-chain data and narrative bias. Any future AI-assisted sustainability procedure would require explicit criteria for source provenance, validation, documentation and retained professional judgment.

5. Discussion

5.1. AI Redistributes Professional Judgment

The evidence supports a redistribution thesis. AI does not remove judgment; it moves judgment across the workflow. Before use, people decide the purpose, training or reference data, permitted population and performance threshold. During use, the auditor decides whether an output is relevant, whether contradictory evidence exists and whether further procedures are necessary. After use, reviewers decide whether documentation supports the conclusion and whether incidents require remediation. Professional judgment therefore operates both on the accounting or assurance issue and on the technology used to analyze it.
This result extends the accounting judgment literature. The traditional bounded judgment space is created by standards, transactions and uncertainty. AI adds an algorithmic evidence layer that determines what is salient and how alternatives are presented. A high anomaly score, generated summary or suggested conclusion can become an anchor. The professional must evaluate both the underlying economic question and the reliability of the mediation. This is why human oversight should be defined as effective intervention, not final approval after the system has framed the answer.

5.2. From Human-in-the-Loop to Accountable Human Control

The phrase human-in-the-loop is too weak if the human role is ceremonial. The public reports use stronger language—human-led, professional judgment, challenge and skepticism—but external accountability also requires operational evidence. An accountable human-control design should record: the approved purpose; the tool and version; the source and permissible use of data; relevant validation; the material output used; contradictory evidence; the professional’s evaluation; any override; the reviewer; and the final conclusion.
This design aligns engagement quality management with AI risk management (Table 9). ISA 220 (Revised) assigns responsibility for managing and achieving quality at engagement level, while ISQM 1 requires a risk-based system of quality management (IAASB, 2020a, 2020b). NIST’s Govern–Map–Measure–Manage sequence provides a compatible technology lens (NIST, 2023). The EU AI Act’s concepts of documentation, logging, human oversight, robustness and cybersecurity provide additional design prompts even where a specific audit tool is not legally classified as high risk (European Parliament and Council, 2024).
The table is not intended as a universal checklist. Controls should be proportionate to the influence of the system. A search assistant that retrieves paragraphs from an approved standards library may require different validation from a model that scores journal entries or generates a valuation range. The decisive factor is not whether the tool is labeled AI, but how its output can affect the nature, timing or extent of procedures and the resulting professional judgment. It should be emphasized that this protocol (Table 9) is an analytical synthesis derived from the governance framework and the disclosure findings; it is a normative proposition rather than an artifact tested on engagement data in this study, and its effect on calibrated reliance and reviewability should be evaluated in future experimental and field research.
A brief illustration shows how the protocol is applied differentially. For a predictive journal-entry risk-scoring model, the approve-purpose stage records that the tool ranks items for attention but does not conclude on misstatement; validate records the test population, the known false-positive and false-negative behavior and the approved version; govern data records the ledger scope and permissions; evaluate output asks what a high score does not establish and what corroboration is required before a conclusion; and override records the auditor’s disposition of flagged and unflagged items. For a generative drafting assistant that prepares a first version of a working-paper narrative, the same stages are populated differently: validation centers on source grounding and hallucination checks; evaluate output asks which assertions are unsupported by cited evidence; and the override stage requires substantive rewriting and sign-off so that the generated text does not anchor the conclusion. The retained-evidence fields are identical in form; what changes is the risk each stage is guarding against.

5.3. Disclosure Completeness Is a Governance Outcome in Its Own Right

Transparency reports serve several functions: regulatory compliance, stakeholder communication and reputation. Measuring them cannot reveal every internal control. Yet disclosure completeness matters because audit committees, investors and regulators need a basis for informed dialog. A firm can reasonably protect proprietary model details while still explaining governance roles, testing categories, permitted-use boundaries, monitoring and the documentation expected when AI influences an engagement.
The AI-JGDI is therefore best understood as a conversation and research instrument. It shows where a public report supplies enough information to identify a mechanism and where it supplies only a general commitment. Equal weighting is the transparent primary specification; Table 6 shows which broad conclusions survive alternative assumptions and which relative positions change. Future research can validate weights through regulator, audit-committee and practitioner elicitation and can test the index against engagement-level evidence.

5.4. Implications for Sustainability Assurance

Sustainability assurance creates a test of whether AI governance is genuinely integrated. CSRD/ESRS and IFRS S1/S2 require connected, decision-useful information, while ISSA 5000 requires appropriate evidence across diverse sustainability matters. AI can support document comparison, evidence classification, anomaly detection and consistency checks across narrative and quantitative information. It can also scale weak source data or produce fluent but unsupported explanations.
Firms should therefore define sustainability-specific AI use cases and evidence boundaries. A model used to classify value-chain evidence should retain links to source documents. A system used to compare disclosures with ESRS should not be treated as determining materiality. An emissions-estimation model should be evaluated for methodology, data completeness, uncertainty and sensitivity. A generative tool used to draft assurance documentation should not be allowed to convert absence of evidence into confident prose. These controls connect AI governance directly to the professional judgments required by ISSA 5000.

5.5. Implications for Regulators, Firms, Audit Committees and Education

Regulators can improve comparability by developing non-prescriptive disclosure expectations for material AI use in audit. Useful categories include use-case governance, validation, data controls, human oversight, documentation, incident monitoring and competence. This would avoid demanding proprietary source code while enabling stakeholders to distinguish aspiration from an operating governance process.
Audit firms can map approved AI use cases to their system of quality management. The map should identify the quality objective affected, the risk created, the control response, the owner, monitoring evidence and remediation route. Model or tool changes should trigger reassessment. Engagement teams should document material use in the same way they document specialists, data analytics or other sources of evidence.
Audit committees should ask focused questions: Which AI tools affected the audit? What data were used? How were the tools tested for the relevant purpose? What outputs were challenged or overridden? What remains a human judgment? Were any limitations communicated? For sustainability assurance, the committee should ask how source provenance and double-materiality judgments were protected from automated simplification.
Education should integrate accounting judgment and AI literacy rather than teach them separately. Case-based learning can require students to evaluate a model output against accounting standards, identify missing evidence, document an override and explain the conclusion to governance bodies. This approach preserves the professional identity of the accountant while preparing graduates for AI-mediated work.

6. Conclusions

This study examined how the UK Big Four publicly describe AI-assisted professional judgment using a complete 2024 cross-section of transparency reports and only public empirical data. All four firms disclosed audit-specific AI use and retained human professional responsibility. Governance ownership and learning were prominent. Validation, model-level explanation and engagement-file traceability were less consistently described. The AI-JGDI ranged from 71.4 to 100.0, but it measures public disclosure completeness and should not be interpreted as actual audit quality. FRC inspection results were reported separately and did not mirror the disclosure ordering.
The main theoretical conclusion is that AI redistributes rather than eliminates judgment. Accountants and auditors make judgments about the reporting issue, the evidence and the technological system mediating that evidence. The main practical conclusion is that human oversight must be evidenced through purpose approval, validation, data controls, documentation, review, override and assigned accountability. A professional signature alone does not demonstrate meaningful control.
The study also identifies a strategic gap for the Special Issue theme: AI and sustainability assurance are prominent but largely parallel disclosures. Public reports provide limited detail about how AI governance is adapted to sustainability evidence, materiality processes, emissions estimates or value-chain information. Extending the human–AI judgment protocol to ISSA 5000 engagements is therefore an immediate research and practice priority.
The limitations are material. The sample contains four firms in one jurisdiction and one reporting cycle. Transparency reports are self-reported, outward-facing documents that may be affected by social-desirability bias and report architecture. The ordered codebook uses equal primary weights and one coder; no inter-coder reliability statistic is reported. The complete decision log improves transparency and contestability but does not replace independent coding. Public disclosures cannot establish engagement-level operation or model performance, and FRC inspection results measure a different construct.
Future research should create a multi-year, multi-jurisdiction panel; use independent coders; validate the index with audit committees and regulators; and examine engagement-level evidence under confidentiality protections. Experiments can test whether the proposed documentation improves calibrated reliance. Field studies can compare financial audit with sustainability assurance. Research should also examine override direction: whether professionals challenge both unfavorable and favorable AI outputs with equal rigor. The public-document method and appendix supplied here provide a reproducible starting point. Extending the index to a multi-year, multi-jurisdiction panel would also allow disclosure trends to be tracked; a reasonable expectation is that AI-specific validation, explainability and data-governance disclosures will become more detailed as the EU AI Act takes effect and sustainability-assurance mandates mature, which the present single-year baseline is designed to measure against.

Funding

This research was partially funded by the Institute for Scientific Research of D. A. Tsenov Academy of Economics, Svishtov, Bulgaria: 1-2026.

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

All empirical data are public. The four firm transparency reports and the FRC publication are available at the URLs listed in the References; selected evidence is reported in Table 5 and Table 8, and the complete 28-decision coding log is documented in Appendix A (Table A1).

Acknowledgments

During the preparation of this manuscript, the author used generative AI tools for language and drafting support. The author has reviewed and edited the output and takes full responsibility for the content of this publication.

Conflicts of Interest

The author declares no conflicts of interest.

Appendix A

Evidence Trail for the AI-JGDI Coding
Appendix A provides a decision-level evidence trail for all 28 firm-by-dimension observations. Each row reports the final single-coder score, a short exact passage or documented non-location, the printed report page and the rationale under Table 3.
Table A1. Decision-level evidence trail supporting AI-JGDI scores.
Table A1. Decision-level evidence trail supporting AI-JGDI scores.
Firm/ScoreDimensionExact Public-Report Evidence and PageDecision Rationale
Deloitte/2AI use“They integrate cognitive technologies, artificial intelligence, customised workflows, and advanced data analytics” (p. 90).Deployed audit-specific platform capability; meets score 2.
Deloitte/2Human oversight“We use our expertise, professional scepticism and judgement to challenge and assure the reliability of any output” (p. 8).Explicit human challenge of output; score 2.
Deloitte/1Validation/testing“continued to pilot GenAI technologies into our Omnia platform” (p. 91); “risk thresholding, triage and clearing house process” (p. 108).Piloting and approval are disclosed, but not performance-testing or release criteria; score 1.
Deloitte/1Traceability“First reviews—a digital review that provides suggestions and prompts to the preparer of audit documentation” (p. 91).Documentation support is described, but not a requirement to record AI influence; score 1.
Deloitte/2Data governance/security“particular focus on use-case approval, data use and management” and UK representation on the “Deloitte Global Data Council” (p. 108).AI-specific data-use governance and named oversight; score 2.
Deloitte/2Accountability“GenAI use-case governance including risk thresholding, triage and clearing house process” (p. 108).Explicit approval and governance route; score 2.
Deloitte/2AI learning“Business change workstream to deliver against the AI fluency and adoption priorities” (p. 108).Structured AI-specific learning/adoption program; score 2.
EY/2AI use“These AI-enabled capabilities… are directly, seamlessly integrated with EY Canvas” (p. 36).Scaled audit-specific AI integrated into the audit platform; score 2.
EY/2Human oversightAI-enabled technology is intended to “support… but not replace the important role of the professional in applying their experience and judgement” (p. 37).Explicit retained professional judgment; score 2.
EY/2Validation/testing“Robust testing throughout the development cycle… is a prerequisite for the release of any audit technology”; release follows testing, piloting, feedback and certification (p. 37).Explicit testing and release conditions; score 2.
EY/1TraceabilityEY Canvas supports “effective coordination, consistent documentation and easier collaboration” (p. 36).General documentation, but no AI-specific engagement trace; score 1.
EY/2Data governance/securityProcedures encourage responsible use of “AI-enabled technologies” (p. 37); certification addresses legal requirements including “data privacy” (p. 41).AI-specific responsible-use control linked with privacy review; score 2.
EY/2Accountability“New Assurance technology concepts are presented to a global committee… for evaluation” (p. 37).Named multidisciplinary approval body; score 2.
EY/1AI learning“EY Badges is a self-directed learning initiative” covering capabilities “such as AI” (p. 44).AI learning is available but self-directed rather than mandatory audit-wide learning; score 1.
KPMG/2AI use“In 2024 we launched KPMG Clara AI chat… to all UK auditors” (p. 36); transaction scoring was “deployed to nearly 900 audits” (p. 37).Deployed audit-specific AI at scale; score 2.
KPMG/2Human oversightKPMG links smart technology with “curious and inquisitive minds and professional scepticism” (p. 36) and trains auditors to “challenge” Clara AI chat (p. 37).Explicit skeptical human challenge; score 2.
KPMG/0Validation/testingNo qualifying AI-specific testing, certification or release-criteria passage was located; deployment and training are described on pp. 36–37.The score-1 threshold requires at least piloting, approval or general evaluation; none was located.
KPMG/1TraceabilityClara AI is expected to help teams “track and report audit progress… [and] improve documentation” (p. 37).General documentation benefit without an AI-use record requirement; score 1.
KPMG/1Data governance/securityKPMG Clara supports “real-time, secure interaction” (p. 36); the Risk Committee reviews “data risks, artificial intelligence and the impact of Copilot” (p. 133).Security and risk oversight are adjacent/general rather than a specific permitted-data control; score 1.
KPMG/2AccountabilityThe “central audit technology team” develops and deploys technology (p. 37); the Risk Committee reviews AI and Copilot risks (p. 133).Named team and committee oversight; score 2.
KPMG/2AI learning“we trained all of our auditors in prompt engineering… [to] challenge and get the most out of KPMG Clara AI chat in a responsible way” (p. 37).Structured AI-specific training for all auditors; score 2.
PwC/2AI use“One GenAI tool that we have rolled out is ChatPwC” and auditors can use approved GenAI use cases on engagements (p. 112).Deployed audit-specific tool and approved engagement uses; score 2.
PwC/2Human oversight“we maintain professional scepticism over the outputs and perform all necessary quality reviews” (p. 112).Explicit skeptical review of output; score 2.
PwC/2Validation/testingThe GenAI Hub proof of concept focuses on “prompt engineering and validation practices” (p. 111).AI-specific validation mechanism; score 2.
PwC/2TraceabilityApproved use cases require “the clear documentation needed where it has been used” (p. 112).Explicit AI-use documentation requirement; score 2.
PwC/2Data governance/securityChatPwC operates in a “secure PwC environment” and “does not use prompts or responses to train the underlying model” (p. 112).AI-specific security and training-data control; score 2.
PwC/2AccountabilityThe “Audit GenAI Hub” is a dedicated team of audit experts, data scientists and innovation managers (p. 112).Named AI governance owner/hub; score 2.
PwC/2AI learning“all Audit staff must complete specific training” on GenAI fundamentals, business rules and prompting (p. 112).Mandatory AI-specific training; score 2.
Note: Coding was performed by one researcher; Table A1 is a transparency and reproducibility device, not an inter-coder reliability test. For score 0, no qualifying passage exists to quote, so the nearest relevant pages and the failed threshold are documented. References to Copilot reproduce the terminology used in KPMG’s transparency report; no software version is specified in the source.
Page numbers refer to the pagination printed in the reports. Short quotations are limited to the text necessary to identify the coded mechanism. A lower score indicates that the specified public disclosure was not located; it does not establish that the firm lacks an undisclosed internal control.

References

  1. Abdullah, A. A. H., & Almaqtari, F. A. (2024). The impact of artificial intelligence and Industry 4.0 on transforming accounting and auditing practices. Journal of Open Innovation: Technology, Market, and Complexity, 10(1), 100218. [Google Scholar] [CrossRef] [Scilit]
  2. Agoglia, C. P., Doupnik, T. S., & Tsakumis, G. T. (2011). Principles-based versus rules-based accounting standards: The influence of standard precision and audit committee strength on financial reporting decisions. The Accounting Review, 86(3), 747–767. [Google Scholar] [CrossRef] [Scilit]
  3. Backof, A. G., Bamber, E. M., & Carpenter, T. D. (2016). Do auditor judgment frameworks help in constraining aggressive reporting? Evidence under more precise and less precise accounting standards. Accounting, Organizations and Society, 51, 1–11. [Google Scholar] [CrossRef] [Scilit]
  4. Bennett, B., Bradbury, M., & Prangnell, H. (2006). Rules, principles and judgments in accounting standards. Abacus, 42(2), 189–204. [Google Scholar] [CrossRef] [Scilit]
  5. Bonner, S. E. (1999). Judgment and decision-making research in accounting. Accounting Horizons, 13(4), 385–398. [Google Scholar] [CrossRef] [Scilit]
  6. Commerford, B. P., Dennis, S. A., Joe, J. R., & Ulla, J. W. (2022). Man versus machine: Complex estimates and auditor reliance on artificial intelligence. Journal of Accounting Research, 60(1), 171–201. [Google Scholar] [CrossRef] [Scilit]
  7. Commerford, B. P., Eilifsen, A., Hatfield, R. C., Holmstrom, K. M., & Kinserdal, F. (2024). Control issues: How providing input affects auditors’ reliance on artificial intelligence. Contemporary Accounting Research, 41(4), 2134–2162. [Google Scholar] [CrossRef] [Scilit]
  8. Deloitte. (2024). Deloitte LLP and Deloitte Limited 2024 transparency report. Deloitte LLP. Available online: https://www.deloitte.com/content/dam/assets-zone2/uk/en/docs/about/2024/deloitte-uk-annual-review-2024-audit-transparency-report.pdf (accessed on 12 July 2026).
  9. Department of Employment and Workplace Relations [DEWR]. (2026). Targeted compliance framework assurance review—Final report. Australian Government. Available online: https://www.dewr.gov.au/assuring-integrity-targeted-compliance-framework/resources/targeted-compliance-framework-assurance-review-final-report (accessed on 22 August 2026).
  10. Dietvorst, B. J., Simmons, J. P., & Massey, C. (2015). Algorithm aversion: People erroneously avoid algorithms after seeing them err. Journal of Experimental Psychology: General, 144(1), 114–126. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  11. European Commission. (2023). Commission delegated regulation (EU) 2023/2772 of 31 July 2023 supplementing directive 2013/34/EU as regards sustainability reporting standards. Official Journal of the European Union, L 2023/2772. Available online: https://eur-lex.europa.eu/eli/reg_del/2023/2772/oj/eng (accessed on 12 July 2026).
  12. European Parliament and Council. (2022). Directive (EU) 2022/2464 of 14 December 2022 as regards corporate sustainability reporting. Official Journal of the European Union, L 322, 15–80. Available online: https://eur-lex.europa.eu/eli/dir/2022/2464/oj/eng (accessed on 12 July 2026).
  13. European Parliament and Council. (2024). Regulation (EU) 2024/1689 of 13 June 2024 laying down harmonised rules on artificial intelligence (artificial intelligence act). Official Journal of the European Union, L 2024/1689. Available online: https://eur-lex.europa.eu/eli/reg/2024/1689/oj/eng (accessed on 12 July 2026).
  14. EY. (2024). EY UK 2024 transparency report. Ernst & Young LLP. Available online: https://www.ey.com/content/dam/ey-unified-site/ey-com/en-uk/about-us/documents/ey-uk-2024-transparency-report.pdf (accessed on 12 July 2026).
  15. Fields, T. D., Lys, T. Z., & Vincent, L. (2001). Empirical research on accounting choice. Journal of Accounting and Economics, 31(1–3), 255–307. [Google Scholar] [CrossRef] [Scilit]
  16. FRC. (2024). FRC publishes annual tier 1 audit firm inspection results. Financial Reporting Council. Available online: https://www.frc.org.uk/news-and-events/news/2024/07/frc-publishes-annual-tier-1-audit-firm-inspection-results/ (accessed on 12 July 2026).
  17. Healy, P. M., & Wahlen, J. M. (1999). A review of the earnings management literature and its implications for standard setting. Accounting Horizons, 13(4), 365–383. [Google Scholar] [CrossRef] [Scilit]
  18. Hurtt, R. K. (2010). Development of a scale to measure professional skepticism. Auditing: A Journal of Practice & Theory, 29(1), 149–171. [Google Scholar] [CrossRef] [Scilit]
  19. IAASB. (2009). International Standard on Auditing 200: Overall objectives of the independent auditor and the conduct of an audit in accordance with International Standards on Auditing. International Auditing and Assurance Standards Board. Available online: https://www.iaasb.org/ (accessed on 12 July 2026).
  20. IAASB. (2018). International Standard on Auditing 540 (revised): Auditing accounting estimates and related disclosures. International Auditing and Assurance Standards Board. Available online: https://www.iaasb.org/focus-areas/embedding-professional-skepticism (accessed on 12 July 2026).
  21. IAASB. (2019). International Standard on Auditing 315 (revised 2019): Identifying and assessing the risks of material misstatement. International Auditing and Assurance Standards Board. Available online: https://www.iaasb.org/consultations-projects/isa-315-revised (accessed on 12 July 2026).
  22. IAASB. (2020a). International Standard on Auditing 220 (revised): Quality management for an audit of financial statements. International Auditing and Assurance Standards Board. Available online: https://www.iaasb.org/publications/international-standard-auditing-220-revised-quality-management-audit-financial-statements (accessed on 12 July 2026).
  23. IAASB. (2020b). International Standard on Quality Management 1: Quality management for firms that perform audits or reviews of financial statements, or other assurance or related services engagements. International Auditing and Assurance Standards Board. Available online: https://www.iaasb.org/publications/international-standard-quality-management-isqm-1-quality-management-firms-perform-audits-or-reviews (accessed on 12 July 2026).
  24. IAASB. (2024). International Standard on Sustainability Assurance 5000: General requirements for sustainability assurance engagements. International Auditing and Assurance Standards Board. Available online: https://www.iaasb.org/publications/international-standard-sustainability-assurance-5000-general-requirements-sustainability-assurance (accessed on 12 July 2026).
  25. IESBA. (2023). Final pronouncement: Technology-related revisions to the code. International Ethics Standards Board for Accountants. Available online: https://www.ethicsboard.org/publications/final-pronouncement-technology-related-revisions-code (accessed on 12 July 2026).
  26. IESBA. (2024). 2024 handbook of the international code of ethics for professional accountants, including international independence standards. International Ethics Standards Board for Accountants. Available online: https://www.ethicsboard.org/publications/2024-handbook-international-code-ethics-professional-accountants (accessed on 12 July 2026).
  27. ISSB. (2023a). IFRS S1 general requirements for disclosure of sustainability-related financial information. IFRS Foundation. Available online: https://www.ifrs.org/issued-standards/ifrs-sustainability-standards-navigator/ifrs-s1-general-requirements/ (accessed on 12 July 2026).
  28. ISSB. (2023b). IFRS S2 climate-related disclosures. IFRS Foundation. Available online: https://www.ifrs.org/issued-standards/ifrs-sustainability-standards-navigator/ifrs-s2-climate-related-disclosures/ (accessed on 12 July 2026).
  29. Jarrahi, M. H., Newlands, G., Lee, M. K., Wolf, C. T., Kinder, E., & Sutherland, W. (2021). Algorithmic management in a work context. Big Data & Society, 8(2), 1–14. [Google Scholar] [CrossRef] [Scilit]
  30. Kellogg, K. C., Valentine, M. A., & Christin, A. (2020). Algorithms at work: The new contested terrain of control. Academy of Management Annals, 14(1), 366–410. [Google Scholar] [CrossRef] [Scilit]
  31. Kokina, J., Blanchette, S., Davenport, T. H., & Pachamanova, D. (2025). Challenges and opportunities for artificial intelligence in auditing: Evidence from the field. International Journal of Accounting Information Systems, 56, 100734. [Google Scholar] [CrossRef] [Scilit]
  32. Kokina, J., & Davenport, T. H. (2017). The emergence of artificial intelligence: How automation is changing auditing. Journal of Emerging Technologies in Accounting, 14(1), 115–122. [Google Scholar] [CrossRef] [Scilit]
  33. KPMG. (2025). UK transparency report 2024. KPMG LLP. Available online: https://assets.kpmg.com/content/dam/kpmgsites/uk/pdf/2026/01/uk-transparency-report-2024.pdf (accessed on 12 July 2026).
  34. Lehner, O. M., Ittonen, K., Silvola, H., Ström, E., & Wührleitner, A. (2022). Artificial intelligence based decision-making in accounting and auditing: Ethical challenges and normative thinking. Accounting, Auditing & Accountability Journal, 35(9), 109–135. [Google Scholar] [CrossRef] [Scilit]
  35. Libby, R., & Luft, J. (1993). Determinants of judgment performance in accounting settings: Ability, knowledge, motivation, and environment. Accounting, Organizations and Society, 18(5), 425–450. [Google Scholar] [CrossRef] [Scilit]
  36. Logg, J. M., Minson, J. A., & Moore, D. A. (2019). Algorithm appreciation: People prefer algorithmic to human judgment. Organizational Behavior and Human Decision Processes, 151, 90–103. [Google Scholar] [CrossRef] [Scilit]
  37. Murikah, W., Nthenge, J. K., & Musyoka, F. M. (2024). Bias and ethics of AI systems applied in auditing: A systematic review. Scientific African, 25, e02281. [Google Scholar] [CrossRef] [Scilit]
  38. Nelson, M. W. (2003). Behavioral evidence on the effects of principles- and rules-based standards. Accounting Horizons, 17(1), 91–104. [Google Scholar] [CrossRef] [Scilit]
  39. Nelson, M. W. (2009). A model and literature review of professional skepticism in auditing. Auditing: A Journal of Practice & Theory, 28(2), 1–34. [Google Scholar] [CrossRef] [Scilit]
  40. NIST. (2023). Artificial intelligence risk management framework (AI RMF 1.0), NIST AI 100-1. National Institute of Standards and Technology, U.S. Department of Commerce. [CrossRef] [Scilit]
  41. PCAOB. (2020). 2019 inspection of Ernst & Young LLP. PCAOB release no. 104-2021-006A. Public Company Accounting Oversight Board. Available online: https://assets.pcaobus.org/pcaob-dev/docs/default-source/inspections/reports/documents/104-2021-006a-ey-2019.pdf (accessed on 22 August 2026).
  42. PCAOB. (2024). PCAOB updates its standards to clarify auditor responsibilities when using technology-assisted analysis. Public Company Accounting Oversight Board. Available online: https://pcaobus.org/news-events/news-releases/news-release-detail/pcaob-updates-its-standards-to-clarify-auditor-responsibilities-when-using-technology-assisted-analysis (accessed on 22 August 2026).
  43. PwC. (2024). UK transparency report 2024. PricewaterhouseCoopers LLP. Available online: https://www.pwc.co.uk/transparencyreport/assets/pdf/uk-transparency-report-2024.pdf (accessed on 12 July 2026).
  44. Stratopoulos, T. C., & Wang, V. X. (2025). Artificial intelligence and accounting research: A framework and agenda. International Journal of Accounting Information Systems, 56, 100760. [Google Scholar] [CrossRef] [Scilit]
  45. Sutton, S. G., Holt, M., & Arnold, V. (2016). “The reports of my death are greatly exaggerated”—Artificial intelligence research in accounting. International Journal of Accounting Information Systems, 22, 60–73. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Human–AI governance model for accountable professional judgment.
Figure 1. Human–AI governance model for accountable professional judgment.
Jrfm 19 00675 g001
Table 1. Normative layers relevant to AI-assisted professional judgment.
Table 1. Normative layers relevant to AI-assisted professional judgment.
LayerPrincipal SourcesJudgment ImplicationAI Governance Implication
Engagement and quality managementISA 200; ISA 220 (Revised); ISA 315 (Revised 2019); ISA 540 (Revised); ISQM 1The auditor retains responsibility for judgment, skepticism, evidence and quality.Tools must be embedded in engagement acceptance, risk assessment, review, documentation and monitoring.
Professional ethicsIESBA Code and technology-related revisionsIntegrity, objectivity, competence, confidentiality and professional behavior apply to technology use.Technology creates threats that require evaluation, safeguards and appropriate professional action.
General AI risk governanceEU AI Act; NIST AI RMF 1.0Human evaluation is necessary where system output affects a consequential decision.Purpose definition, risk management, data controls, testing, documentation, oversight, robustness and cybersecurity.
Sustainability reporting and assuranceCSRD; ESRS; IFRS S1; IFRS S2; ISSA 5000Materiality, estimates, source reliability and sufficiency of evidence remain professional judgments.AI use should preserve provenance, uncertainty, connected information and reviewability across financial and sustainability data.
Source: Author’s synthesis of the cited standards and legislation. The EU AI Act benchmark is applied analytically; the table does not classify every audit tool as a high-risk AI system.
Table 2. Public empirical corpus.
Table 2. Public empirical corpus.
FirmReporting PeriodPDF PagesPrincipal AI Evidence PagesDocument Status
Deloitte UKYear ended 31 May 20241558; 90–92; 108Official 2024 Transparency Report
EY UKYear ended 28 June 202416136–37; 60–61; 129–130Official 2024 Transparency Report
KPMG UKYear ended 30 September 202417736–37; 133Official 2024 Transparency Report
PwC UKYear ended 30 June 2024176110–112; 102–103Official 2024 Transparency Report
Source: Official firm reports. Total corpus: 669 pages. Reports were accessed on 12 July 2026.
Table 3. AI–Judgment Governance Disclosure Index codebook.
Table 3. AI–Judgment Governance Disclosure Index codebook.
DimensionScore 0Score 1Score 2: Explicit Mechanism
AI useNo audit AI locatedExperiment or general aspirationDeployed/scaled audit-specific tool or use case
Human oversightNo human role locatedGeneral professional judgment statementExplicit challenge, review or override of AI output
Validation/testingNo relevant disclosurePilot, approval or general evaluationAI-specific testing/validation and release criteria
TraceabilityNo relevant disclosureGeneral documentation/transparencyAI-use documentation, explanation or audit trail required
Data governance/securityNo relevant disclosureGeneral security/privacy controlAI-specific permitted-data, privacy, security or training-data control
AccountabilityNo owner or route locatedGeneral governanceNamed committee, owner, hub, approval or escalation route
AI learningNo relevant disclosureOptional/general technology learningStructured or mandatory AI-specific learning for auditors
Note: The index measures the completeness of public disclosure, not the operating effectiveness of controls.
Table 4. Firm-by-dimension AI-JGDI scores.
Table 4. Firm-by-dimension AI-JGDI scores.
DimensionDeloitteEYKPMGPwC
Audit-specific AI integration2222
Human judgment and oversight2222
Validation and testing1202
Explainability and documentation1112
Data governance and security2212
Accountability structures2222
AI-specific learning2122
Total (maximum 14)12121014
AI-JGDI85.785.771.4100.0
Note: Source: Author’s coding of the official 2024 UK transparency reports. Scores follow Table 3; selected named evidence and pages are presented immediately below, and the complete decision log is in Appendix A.
Table 5. Selected public-report evidence anchoring the AI-JGDI scores.
Table 5. Selected public-report evidence anchoring the AI-JGDI scores.
FirmAI-JGDIConcrete Evidence Anchoring Scores of 2Evidence Limitations Anchoring Scores of 0–1
Deloitte85.7PairD and AI/ML in Omnia (pp. 8, 90–92); Data Council and GenAI use-case governance (p. 108); AI-fluency workstream (p. 108).Validation = 1: piloting and risk thresholding, but no disclosed performance-testing or certification criteria (pp. 91, 108). Traceability = 1: first-review support, but no AI-use record requirement (p. 91).
EY85.7Globally scaled AI integrated with EY Canvas (p. 36); testing, piloting and certification before release (p. 37); global evaluation committee (p. 37).Traceability = 1: consistent documentation is described without an AI-specific engagement trace (p. 36). AI learning = 1: AI Badges are self-directed rather than mandatory (p. 44).
KPMG71.4Clara AI chat for all UK auditors and transaction scoring on nearly 900 audits (pp. 36–37); all auditors trained in prompt engineering (p. 37).Validation = 0: no AI-specific testing, certification or release criteria located. Traceability = 1: AI is expected to improve documentation without an AI-use record rule (p. 37). Data governance = 1: secure exchange and general data-risk oversight (pp. 36, 133).
PwC100.0Audit GenAI Hub and validation practices (p. 111); ChatPwC, approved use cases, documentation, secure data controls, human review and required training (p. 112).None under the stated disclosure rules. A score of 2 in every dimension indicates complete public evidence, not verified operating effectiveness.
Note: Page numbers refer to the pagination printed in the original transparency reports. A score of 0 means that no qualifying disclosure was located in the complete report; it does not establish that the internal control does not exist. A score of 2 indicates explicit public evidence, not verified operating effectiveness.
Table 6. Sensitivity of AI-JGDI disclosure profiles to alternative weighting schemes.
Table 6. Sensitivity of AI-JGDI disclosure profiles to alternative weighting schemes.
Weighting ScenarioDeloitteEYKPMGPwC
Equal weights85.785.771.4100.0
Control-risk emphasis80.085.060.0100.0
Human-accountability emphasis88.988.977.8100.0
Competence emphasis87.581.375.0100.0
Note: Control-risk emphasis doubles validation/testing, traceability and data governance/security. Human-accountability emphasis doubles human oversight and accountability. Competence emphasis doubles AI-specific learning. All other weights equal 1, and results are normalized to 0–100. These illustrative scenarios do not validate differential weights.
Table 7. AI disclosure governance and external FRC inspection context.
Table 7. AI disclosure governance and external FRC inspection context.
FirmAI-JGDIFRC 2024: Inspected Audits Good/Limited ImprovementsPermitted Interpretation
Deloitte85.794%Disclosure completeness and a risk-based inspection result are separate indicators.
EY85.776%No firm-level causal inference is made.
KPMG71.489%A lower disclosure score does not establish weaker internal practice.
PwC100.076%A complete disclosure score does not establish higher audit quality.
Note: Source: AI-JGDI from Table 4; inspection percentages from FRC (2024). The FRC sample is risk-based and should not be extrapolated to all audits.
Table 8. Public evidence of AI disclosure governance and sustainability-assurance integration.
Table 8. Public evidence of AI disclosure governance and sustainability-assurance integration.
FirmSustainability-Assurance EvidenceAI Disclosure Governance EvidenceExplicit LinkCoding Rationale
Deloitte“enhanced methodology for undertaking sustainability assurance engagements” (p. 94)PairD/Omnia and GenAI use-case governance (pp. 8, 90–92, 108)NoThe report presents both themes, but does not apply an AI tool or AI-specific safeguard to sustainability-assurance evidence.
EY“EY Sustainability Assurance Methodology… [a] consistent approach” (p. 41)AI integrated with Canvas and technology release controls (pp. 36–37)NoThe sustainability methodology and AI controls are separately described; no engagement-level methodological connection is stated.
KPMG“dedicated ESG Assurance team working closely with Audit teams” (p. 46)Clara AI chat and transaction scoring (pp. 36–37)NoAI deployment is not linked to ESG-assurance procedures, evidence or controls.
PwC“integrated financial and non-financial assurance team” (p. 116)Audit GenAI Hub, ChatPwC and approved use cases (pp. 111–112)NoThe report separately describes ESG/CSRD assurance (pp. 115–117); AI is not connected to a sustainability-assurance method.
Note: Page numbers refer to the pagination printed in the original reports. “No” means that no qualifying engagement-level methodological link was located after full-report searching and contextual review. It does not demonstrate that no internal practice exists.
Table 9. Accountable human-control protocol for financial and sustainability assurance.
Table 9. Accountable human-control protocol for financial and sustainability assurance.
Control StageMinimum Retained EvidenceProfessional Judgment QuestionPrincipal Normative Anchor
Approve purposeDefined use case, intended user, prohibited use and responsible ownerCan this tool appropriately support this task without determining the conclusion?ISQM 1; NIST Govern/Map
ValidateTest population, criteria, known limitations, approval and change/version historyIs performance sufficient for this purpose and population?NIST Measure; EU AI Act control concepts
Govern dataProvenance, permissions, confidentiality, retention and securityIs the input complete, lawful, reliable and appropriate?IESBA confidentiality; ISA 315; NIST Govern
Evaluate outputMaterial output, uncertainty, contradictory evidence and additional proceduresWhat does the output not establish, and what evidence could disconfirm it?ISA 200; ISA 540; ISSA 5000
Override and reviewAcceptance/override rationale, reviewer, escalation and final conclusionWould the same challenge be applied if the output moved the conclusion in the opposite direction?ISA 220; professional skepticism
Monitor and remediateIncidents, drift, user feedback, inspection findings and remediationShould the use case, control or training be changed or withdrawn?ISQM 1; NIST Manage
Note: Source: Author’s synthesis. Controls should be proportionate to the influence and risk of the use case.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Krasteva-Hristova, R. Professional Judgment and AI Disclosure Governance in Audit and Sustainability Assurance: Public Evidence from the UK Big Four. J. Risk Financ. Manag. 2026, 19, 675. https://doi.org/10.3390/jrfm19090675

AMA Style

Krasteva-Hristova R. Professional Judgment and AI Disclosure Governance in Audit and Sustainability Assurance: Public Evidence from the UK Big Four. Journal of Risk and Financial Management. 2026; 19(9):675. https://doi.org/10.3390/jrfm19090675

Chicago/Turabian Style

Krasteva-Hristova, Radosveta. 2026. "Professional Judgment and AI Disclosure Governance in Audit and Sustainability Assurance: Public Evidence from the UK Big Four" Journal of Risk and Financial Management 19, no. 9: 675. https://doi.org/10.3390/jrfm19090675

APA Style

Krasteva-Hristova, R. (2026). Professional Judgment and AI Disclosure Governance in Audit and Sustainability Assurance: Public Evidence from the UK Big Four. Journal of Risk and Financial Management, 19(9), 675. https://doi.org/10.3390/jrfm19090675

Article Metrics

Back to TopTop