1. Introduction
Financial reporting and assurance are becoming simultaneously more data-intensive and more judgment-intensive. Artificial intelligence can search contracts, identify anomalous transactions, score populations of journal entries, retrieve technical guidance, draft working papers and direct attention toward unusual patterns. Sustainability reporting adds heterogeneous operational, environmental and value-chain data, forward-looking assumptions, double-materiality assessments and information that may not have passed through financial-reporting control systems. These developments enlarge the evidence available to accountants and auditors, but they also increase the need to decide which evidence is relevant, reliable and sufficiently persuasive.
Professional judgment is therefore not a residual activity left after automation. It is the mechanism through which standards, evidence, expertise and ethical responsibility are converted into a defensible conclusion. Accounting judgment research has long shown that performance depends on task knowledge, experience, motivation and the decision environment (
Bonner, 1999;
Libby & Luft, 1993). Audit research adds professional skepticism: a questioning mind, alertness to contradictory evidence and a critical assessment of audit evidence (
Hurtt, 2010;
Nelson, 2009). AI changes this environment because a system can structure the evidence presented to the auditor, rank risks, recommend an action or create a first draft that becomes a powerful anchor. The system can improve consistency and coverage, but an apparently precise output can also conceal data limitations, model error, bias or a management-defined objective.
The behavioral evidence is deliberately cautionary. People may avoid algorithms after observing an error (
Dietvorst et al., 2015), yet they may also prefer algorithmic advice under other conditions (
Logg et al., 2019). In audit settings, reliance depends on task complexity, perceived control and the opportunity to provide input (
Commerford et al., 2022,
2024). AI can consequently produce either under-reliance or automation bias. The governance question is how to calibrate reliance: the professional should use the system where it is competent, challenge it where its limits matter, and remain accountable for the conclusion.
Public incidents make the risk concrete, but they also reveal an evidentiary boundary. Deloitte Australia’s Targeted Compliance Framework independent assurance review was republished to correct identified errors after the firm disclosed the use of generative AI in preparing the report (
Department of Employment and Workplace Relations [DEWR], 2026); this was an assurance review, not a statutory financial-statement audit. In the closest public audit analogue, the PCAOB found that five EY audits inspected in 2019 lacked sufficient testing of the accuracy and completeness of data and reports used in substantive procedures (
PCAOB, 2020). The PCAOB did not attribute those deficiencies to AI. As of August 2026, no public FRC or PCAOB enforcement decision identified by this study establishes AI misjudgment as the cause of a statutory audit failure. Instead, the record documents two failure mechanisms—unverified generated content and reliance on insufficiently tested electronic information—that AI can magnify. This is why the PCAOB amended AS 1105 and AS 2301 to reduce the risk that an auditor using technology-assisted analysis might issue an opinion without sufficient appropriate audit evidence (
PCAOB, 2024).
The normative environment is also moving quickly. The IESBA technology-related revisions preserve the fundamental principles of integrity, objectivity, professional competence and due care, confidentiality and professional behavior when technology is used (
IESBA, 2023,
2024). ISA 315 (Revised 2019) requires understanding of the entity and its information system and strengthens the exercise of professional skepticism in risk assessment (
IAASB, 2019). ISA 540 (Revised) is directly relevant to AI-supported estimates because it emphasizes estimation uncertainty, management bias and skeptical evaluation (
IAASB, 2018). The EU Artificial Intelligence Act establishes a risk-based regulatory architecture and, for relevant systems, requirements related to governance, technical documentation, logging, transparency, human oversight, accuracy, robustness and cybersecurity (
European Parliament and Council, 2024). The NIST AI Risk Management Framework organizes voluntary risk management around Govern, Map, Measure and Manage (
NIST, 2023). None of these instruments transfers an auditor’s responsibility to an algorithm.
At the same time, sustainability reporting and assurance have moved from voluntary practice toward formal requirements. The Corporate Sustainability Reporting Directive and European Sustainability Reporting Standards require extensive sustainability information and governance disclosures in the European Union (
European Commission, 2023;
European Parliament and Council, 2022). IFRS S1 and IFRS S2 establish an investor-focused global baseline organized around governance, strategy, risk management, and metrics and targets (
ISSB, 2023a,
2023b). ISSA 5000 supplies a global, framework-neutral standard for sustainability assurance engagements (
IAASB, 2024). AI may help process these data, but sustainability evidence often contains measurement uncertainty, estimates and value-chain information for which explainability and provenance are indispensable.
Existing research offers important insights but leaves three gaps. First, many studies discuss potential benefits and risks without examining what large audit firms publicly disclose about the governance of AI-assisted judgment. Second, corporate reports often combine technology, people, quality and risk narratives; a reproducible coding framework is needed to distinguish a deployed tool from a governance safeguard. Third, the relationship between AI governance and sustainability assurance is frequently asserted but rarely traced in comparable public documents. Recent work calls for research that examines specific AI configurations and the institutional arrangements surrounding them rather than treating AI adoption as a single binary variable (
Kokina et al., 2025;
Lehner et al., 2022;
Stratopoulos & Wang, 2025).
This study asks three research questions:
RQ1. How do the UK Big Four publicly describe the role of AI in audit work and professional judgment?
RQ2. Which safeguards for validation, documentation, data governance, accountability and learning are disclosed, and how complete are these disclosures across firms?
RQ3. To what extent do the transparency reports connect AI disclosure governance evidence with sustainability reporting and assurance?
The empirical setting is the United Kingdom. The Big Four operate under a common regulatory environment, publish transparency reports under comparable requirements and are subject to public inspection by the Financial Reporting Council (FRC). The corpus comprises the complete 2024 transparency-report cross-section for Deloitte, EY, KPMG and PwC UK. A structured content analysis codes seven governance dimensions and constructs an AI–Judgment Governance Disclosure Index (AI-JGDI). The FRC’s 2024 inspection results are used only as external context; the study does not infer that a disclosure score causes an inspection outcome.
The paper makes four contributions. It provides public empirical evidence about how leading audit firms frame human–AI responsibility. It introduces a transparent index whose items can be replicated or extended to other jurisdictions and years. It separates disclosure completeness from actual governance effectiveness and from regulatory audit quality. Finally, it develops a practical human–AI judgment protocol relevant to both financial-statement audit and sustainability assurance. To avoid overstating scope, the empirical evidence concerns AI governance in a financial-statement audit, whereas the treatment of sustainability assurance is conceptual: the protocol is extended analytically to ISSA 5000 engagements rather than tested on sustainability-assurance evidence. Stated precisely, the contribution is to measure the completeness of firms’ public accountability commitments, not to assess the effectiveness of their internal controls.
The theoretical contribution is a refinement rather than a new theory. The study extends the judgment-performance perspective in accounting (
Bonner, 1999;
Libby & Luft, 1993) and the professional-skepticism literature (
Hurtt, 2010;
Nelson, 2009) by adding an AI-mediated evidence layer to the judgment process. In this account the professional makes judgments not only about the reporting issue and the underlying evidence, but also about the technological system that structures that evidence—its inputs, assumptions and limits. The AI-mediated layer changes the object of skepticism: calibrated reliance and the reviewability of an AI-influenced conclusion become properties that professional judgment must actively manage. The paper operationalizes this refinement as a measurable disclosure construct (the AI-JGDI) and illustrates it empirically; it is therefore a conceptual refinement combined with a measurement contribution, not merely an empirical illustration.
A distinction is maintained throughout between AI governance and AI disclosure governance. AI governance refers to the internal structures, controls, responsibilities and practices through which a firm manages AI. AI disclosure governance refers to the subset of those arrangements described in public reports. Because the empirical evidence consists exclusively of public documents, the AI-JGDI measures AI disclosure governance rather than the existence or operating effectiveness of internal AI governance arrangements. Empirical claims therefore concern disclosed mechanisms; AI governance is reserved for the broader conceptual and normative framework.
4. Results
4.1. AI Is Disclosed as Deployed Augmentation, Not Autonomous Judgment
All four reports move beyond a generic expectation that AI may be used in the future. Deloitte describes PairD, AI and machine-learning functions in Omnia and audit pilots for research, document retrieval, first review and document creation (
Deloitte, 2024, pp. 8, 90–92). EY reports globally scaled AI integrated with EY Canvas and explicit testing and certification controls (
EY, 2024, pp. 36–37). KPMG reports Clara AI chat for all UK auditors and transaction scoring deployed to nearly 900 UK audits (
KPMG, 2025, pp. 36–37). PwC reports an Audit GenAI Hub, ChatPwC and approved engagement use cases (
PwC, 2024, pp. 111–112).
The common operational model is augmentation. Deloitte states that expertise, professional skepticism and judgment are used to challenge and assure the reliability of output. EY says that AI-enabled technology supports procedures but does not replace the professional’s experience and judgment. KPMG combines smart technology with curious and inquisitive minds and professional skepticism. PwC describes a human-led, technology-powered audit and requires skeptical review of GenAI outputs. These disclosures support a consistent answer to RQ1: public accountability remains attached to the auditor, even where AI is scaled across the practice.
This framing is important because tools perform different functions. An anomaly score reallocates attention; a technical chatbot retrieves or synthesizes guidance; a drafting assistant creates an initial artifact; and a transaction-scoring model analyses a population. None of these functions establishes by itself whether evidence is sufficient or whether an accounting estimate is reasonable. The reports generally recognize this boundary.
4.2. Validation and Traceability Are the Least Consistently Disclosed Safeguards
Validation disclosure varies more than adoption disclosure. EY provides the clearest lifecycle description: technology concepts pass through a global committee; testing with end users, piloting, feedback and certification are prerequisites for release. PwC describes prompt engineering and validation practices in the Audit GenAI Hub. Deloitte reports pilots, use-case approval, risk thresholding and a clearing house process, which demonstrate gatekeeping but provides less detail about performance validation. KPMG reports extensive deployment and responsible user challenge, but the 2024 report does not describe an AI-specific validation or certification mechanism at the same level of detail. Under the strict codebook, this produces scores of 2 for EY and PwC, 1 for Deloitte and 0 for KPMG on the validation dimension.
Explainability is also incompletely disclosed. The reports discuss transparent audit services, documentation and the ability of AI tools to support or improve working papers. However, only PwC explicitly states that clear documentation is required where approved GenAI use cases have been used. None of the reports supplies public model-level information such as performance thresholds, error rates, explainability methods, override frequency or post-deployment drift monitoring. Such details may exist internally and may be inappropriate to disclose fully for security or proprietary reasons. Yet an external reader cannot determine from most reports what minimum explanation must be retained in an engagement file when an AI output materially influences a judgment.
The result identifies a disclosure boundary rather than proving a control deficiency. Public transparency reports are designed for multiple regulatory and stakeholder purposes, not as model cards. Nevertheless, validation and traceability are the dimensions where the difference between a statement of responsible intent and an externally assessable control is greatest.
4.3. Data Governance and Accountability Are More Visible
Deloitte discloses a safe and secure environment for PairD and a firm-wide GenAI risk response involving use-case approval, data use and management, ethical use, cyber risk, a Global Data Council, risk thresholding and a Trustworthy AI framework (
Deloitte, 2024, p. 108). EY links responsible technology use with standardized development protocols, a global evaluation committee, information-security policies and privacy-impact assessments for new technology (
EY, 2024, pp. 36–37, 60–61, 129–130). KPMG reports Risk Committee deep dives on AI and data risk, secure interaction within Clara, and general information-security governance, but the AI section contains less explicit detail about AI-specific data lineage or permitted-data rules. PwC describes a secure environment, states that ChatPwC does not use prompts or responses to train the underlying model, limits allowable tools and use cases, and requires clear documentation (
PwC, 2024, pp. 111–112).
Accountability structures are explicit across the corpus. Deloitte identifies use-case governance and program ownership. EY identifies a global committee involving Professional Practice, the Assurance Quality Network and Technology. KPMG identifies central technology teams and board/Risk Committee oversight. PwC identifies an Audit GenAI Hub with audit subject-matter experts, data scientists and innovation managers. All four therefore receive the maximum accountability score. The reports differ less on whether someone owns AI governance than on what public evidence is supplied about validation outputs and engagement-level traceability.
4.4. Learning Is Treated as a Condition of Responsible Use
Three firms disclose broad AI-specific learning at a level meeting the maximum criterion. Deloitte’s principal-risk response includes an AI-fluency workstream. KPMG reports that all auditors were trained in prompt engineering at its 2024 Audit University so that they could engage with and challenge Clara AI chat responsibly. PwC makes training on GenAI fundamentals and audit business rules mandatory before granting access to ChatPwC. EY discloses AI badges and technology learning, but the 2024 Transparency Report is less explicit about a mandatory audit-wide AI curriculum; it therefore receives a score of 1 under the strict rule.
The emphasis on learning supports a dual-competence model. An auditor needs domain competence to recognize an implausible output and AI literacy to understand data, uncertainty, permitted use and limitations. Prompt skill alone is not professional competence. Conversely, a technically expert accountant who cannot evaluate model limitations may either reject useful evidence or accept output ceremonially. The reports generally treat technology training as complementary to professional skepticism rather than as its replacement.
4.5. Disclosure Scores and External Inspection Context
The cross-firm comparison reveals a common baseline of deployed audit-specific AI, retained human oversight and identifiable accountability structures. The principal differences concern the specificity of validation and testing mechanisms, engagement-level traceability, AI-related data controls and learning arrangements.
Table 4 presents the dimension-level coding and the resulting AI–Judgment Governance Disclosure Index (AI-JGDI).
To make the evidence directly visible alongside the scores,
Table 5 identifies the named tools and governance mechanisms supporting the most complete disclosures and the specific limitations responsible for scores of 0 or 1.
Because the 0–1–2 items are ordered categories, frequency patterns—not means or standard deviations—are used at the dimension level. All four firms score 2 for AI use, human oversight and accountability. Validation/testing contains one 0, one 1 and two 2s; traceability contains three 1s and one 2; data governance and AI learning each contain one 1 and three 2s.
Table 6 tests whether the firm-level disclosure profiles are sensitive to alternative weighting assumptions.
The broad endpoints are stable across the four specifications: PwC retains the most complete public disclosure profile and KPMG the least complete. The relative position of Deloitte and EY is not invariant: EY is higher under control-risk emphasis, whereas Deloitte is higher under competence emphasis. The analysis therefore supports a broad cross-firm disclosure pattern but not a precise ranking. Variation arises primarily from the specificity of validation, traceability, AI-related data controls and structured learning.
The AI-JGDI should therefore be interpreted as a profile of public disclosure completeness rather than as a ranking of actual governance effectiveness or audit quality. A maximum score indicates that explicit public evidence was identified for every coded dimension; it does not demonstrate that the disclosed controls operated effectively across all engagements. Correspondingly, a lower score indicates that the specified mechanism was not located in the transparency report and does not establish that the firm lacks an equivalent internal control. To avoid implying a governance ranking, the firm-level values are read as disclosure profiles rather than league-table positions, and the paper deliberately avoids describing any firm as having stronger or weaker governance on the basis of the index.
Several explanations for the observed differences are plausible, although none can be confirmed with public data and each is offered only to orient future research. First, the differences may reflect strategic disclosure choices: a firm may adopt a more expansive AI-disclosure posture to project market leadership, which would raise its completeness score without necessarily implying superior internal practice. Second, they may reflect report-architecture effects: where relevant material is dispersed across governance, technology and quality sections rather than consolidated, the strict coding rules are less likely to locate an explicit mechanism, which may partly explain a lower score such as KPMG’s on validation and data governance. Third, they may reflect differences in risk culture and communication style, which shape how much operational detail a firm is willing to place in a public document. Distinguishing these explanations would require interviews or engagement-level access and is left to future work.
The FRC’s risk-based inspection results provide an external context for interpreting the disclosure index. As shown in
Table 7, the ordering of the inspection outcomes does not correspond to the AI-JGDI ordering. This descriptive mismatch confirms that public AI disclosure governance and regulatory audit-quality inspection capture different aspects of accountability and should not be treated as interchangeable measures.
The comparison does not support a firm-level causal inference. The AI-JGDI measures the completeness of governance mechanisms disclosed in public transparency reports, whereas the FRC percentages reflect the outcomes of risk-based inspections of selected audit engagements. The two indicators differ in their objects of measurement, evidence bases and sampling conditions. Consequently, a more complete AI disclosure governance profile cannot be interpreted as evidence of higher audit quality, just as a lower disclosure score cannot be interpreted as evidence of weaker internal practice.
4.6. AI Disclosure Governance and Sustainability Assurance Remain Parallel Rather than Integrated Narratives
RQ3 was assessed using three pre-specified indicators. Sustainability assurance was coded present when a report discussed methodology, services, competence or engagement responsibilities. AI disclosure governance was coded present when a report disclosed an AI-specific governance mechanism. Integration required an explicit connection between an AI tool or safeguard and a sustainability-assurance procedure, evidence source or judgment.
The evidence shows concrete separation. Deloitte describes AI in Omnia and PairD (pp. 8, 90–92) and separately reports an “enhanced methodology for undertaking sustainability assurance engagements” (p. 94). EY describes AI integrated with Canvas (p. 36) and separately presents the EY Sustainability Assurance Methodology (p. 41). KPMG describes Clara AI chat and transaction scoring (pp. 36–37) and separately a “dedicated ESG Assurance team” (p. 46). PwC describes the Audit GenAI Hub and ChatPwC (pp. 111–112) and separately an “integrated financial and non-financial assurance team” (p. 116). None of these passages applies an AI tool, validation rule, provenance control or documentation requirement to sustainability-assurance evidence.
Table 8 summarizes this firm-by-firm evidence and the resulting integration assessment.
The systematic answer to RQ3 is therefore that all four firms disclose sustainability-assurance activity and separate AI disclosure governance evidence, but none provides an explicit public methodological link between them. This matters because sustainability evidence is especially exposed to inconsistent definitions, estimation uncertainty, missing value-chain data and narrative bias. Any future AI-assisted sustainability procedure would require explicit criteria for source provenance, validation, documentation and retained professional judgment.
5. Discussion
5.1. AI Redistributes Professional Judgment
The evidence supports a redistribution thesis. AI does not remove judgment; it moves judgment across the workflow. Before use, people decide the purpose, training or reference data, permitted population and performance threshold. During use, the auditor decides whether an output is relevant, whether contradictory evidence exists and whether further procedures are necessary. After use, reviewers decide whether documentation supports the conclusion and whether incidents require remediation. Professional judgment therefore operates both on the accounting or assurance issue and on the technology used to analyze it.
This result extends the accounting judgment literature. The traditional bounded judgment space is created by standards, transactions and uncertainty. AI adds an algorithmic evidence layer that determines what is salient and how alternatives are presented. A high anomaly score, generated summary or suggested conclusion can become an anchor. The professional must evaluate both the underlying economic question and the reliability of the mediation. This is why human oversight should be defined as effective intervention, not final approval after the system has framed the answer.
5.2. From Human-in-the-Loop to Accountable Human Control
The phrase human-in-the-loop is too weak if the human role is ceremonial. The public reports use stronger language—human-led, professional judgment, challenge and skepticism—but external accountability also requires operational evidence. An accountable human-control design should record: the approved purpose; the tool and version; the source and permissible use of data; relevant validation; the material output used; contradictory evidence; the professional’s evaluation; any override; the reviewer; and the final conclusion.
This design aligns engagement quality management with AI risk management (
Table 9). ISA 220 (Revised) assigns responsibility for managing and achieving quality at engagement level, while ISQM 1 requires a risk-based system of quality management (
IAASB, 2020a,
2020b). NIST’s Govern–Map–Measure–Manage sequence provides a compatible technology lens (
NIST, 2023). The EU AI Act’s concepts of documentation, logging, human oversight, robustness and cybersecurity provide additional design prompts even where a specific audit tool is not legally classified as high risk (
European Parliament and Council, 2024).
The table is not intended as a universal checklist. Controls should be proportionate to the influence of the system. A search assistant that retrieves paragraphs from an approved standards library may require different validation from a model that scores journal entries or generates a valuation range. The decisive factor is not whether the tool is labeled AI, but how its output can affect the nature, timing or extent of procedures and the resulting professional judgment. It should be emphasized that this protocol (
Table 9) is an analytical synthesis derived from the governance framework and the disclosure findings; it is a normative proposition rather than an artifact tested on engagement data in this study, and its effect on calibrated reliance and reviewability should be evaluated in future experimental and field research.
A brief illustration shows how the protocol is applied differentially. For a predictive journal-entry risk-scoring model, the approve-purpose stage records that the tool ranks items for attention but does not conclude on misstatement; validate records the test population, the known false-positive and false-negative behavior and the approved version; govern data records the ledger scope and permissions; evaluate output asks what a high score does not establish and what corroboration is required before a conclusion; and override records the auditor’s disposition of flagged and unflagged items. For a generative drafting assistant that prepares a first version of a working-paper narrative, the same stages are populated differently: validation centers on source grounding and hallucination checks; evaluate output asks which assertions are unsupported by cited evidence; and the override stage requires substantive rewriting and sign-off so that the generated text does not anchor the conclusion. The retained-evidence fields are identical in form; what changes is the risk each stage is guarding against.
5.3. Disclosure Completeness Is a Governance Outcome in Its Own Right
Transparency reports serve several functions: regulatory compliance, stakeholder communication and reputation. Measuring them cannot reveal every internal control. Yet disclosure completeness matters because audit committees, investors and regulators need a basis for informed dialog. A firm can reasonably protect proprietary model details while still explaining governance roles, testing categories, permitted-use boundaries, monitoring and the documentation expected when AI influences an engagement.
The AI-JGDI is therefore best understood as a conversation and research instrument. It shows where a public report supplies enough information to identify a mechanism and where it supplies only a general commitment. Equal weighting is the transparent primary specification;
Table 6 shows which broad conclusions survive alternative assumptions and which relative positions change. Future research can validate weights through regulator, audit-committee and practitioner elicitation and can test the index against engagement-level evidence.
5.4. Implications for Sustainability Assurance
Sustainability assurance creates a test of whether AI governance is genuinely integrated. CSRD/ESRS and IFRS S1/S2 require connected, decision-useful information, while ISSA 5000 requires appropriate evidence across diverse sustainability matters. AI can support document comparison, evidence classification, anomaly detection and consistency checks across narrative and quantitative information. It can also scale weak source data or produce fluent but unsupported explanations.
Firms should therefore define sustainability-specific AI use cases and evidence boundaries. A model used to classify value-chain evidence should retain links to source documents. A system used to compare disclosures with ESRS should not be treated as determining materiality. An emissions-estimation model should be evaluated for methodology, data completeness, uncertainty and sensitivity. A generative tool used to draft assurance documentation should not be allowed to convert absence of evidence into confident prose. These controls connect AI governance directly to the professional judgments required by ISSA 5000.
5.5. Implications for Regulators, Firms, Audit Committees and Education
Regulators can improve comparability by developing non-prescriptive disclosure expectations for material AI use in audit. Useful categories include use-case governance, validation, data controls, human oversight, documentation, incident monitoring and competence. This would avoid demanding proprietary source code while enabling stakeholders to distinguish aspiration from an operating governance process.
Audit firms can map approved AI use cases to their system of quality management. The map should identify the quality objective affected, the risk created, the control response, the owner, monitoring evidence and remediation route. Model or tool changes should trigger reassessment. Engagement teams should document material use in the same way they document specialists, data analytics or other sources of evidence.
Audit committees should ask focused questions: Which AI tools affected the audit? What data were used? How were the tools tested for the relevant purpose? What outputs were challenged or overridden? What remains a human judgment? Were any limitations communicated? For sustainability assurance, the committee should ask how source provenance and double-materiality judgments were protected from automated simplification.
Education should integrate accounting judgment and AI literacy rather than teach them separately. Case-based learning can require students to evaluate a model output against accounting standards, identify missing evidence, document an override and explain the conclusion to governance bodies. This approach preserves the professional identity of the accountant while preparing graduates for AI-mediated work.
6. Conclusions
This study examined how the UK Big Four publicly describe AI-assisted professional judgment using a complete 2024 cross-section of transparency reports and only public empirical data. All four firms disclosed audit-specific AI use and retained human professional responsibility. Governance ownership and learning were prominent. Validation, model-level explanation and engagement-file traceability were less consistently described. The AI-JGDI ranged from 71.4 to 100.0, but it measures public disclosure completeness and should not be interpreted as actual audit quality. FRC inspection results were reported separately and did not mirror the disclosure ordering.
The main theoretical conclusion is that AI redistributes rather than eliminates judgment. Accountants and auditors make judgments about the reporting issue, the evidence and the technological system mediating that evidence. The main practical conclusion is that human oversight must be evidenced through purpose approval, validation, data controls, documentation, review, override and assigned accountability. A professional signature alone does not demonstrate meaningful control.
The study also identifies a strategic gap for the Special Issue theme: AI and sustainability assurance are prominent but largely parallel disclosures. Public reports provide limited detail about how AI governance is adapted to sustainability evidence, materiality processes, emissions estimates or value-chain information. Extending the human–AI judgment protocol to ISSA 5000 engagements is therefore an immediate research and practice priority.
The limitations are material. The sample contains four firms in one jurisdiction and one reporting cycle. Transparency reports are self-reported, outward-facing documents that may be affected by social-desirability bias and report architecture. The ordered codebook uses equal primary weights and one coder; no inter-coder reliability statistic is reported. The complete decision log improves transparency and contestability but does not replace independent coding. Public disclosures cannot establish engagement-level operation or model performance, and FRC inspection results measure a different construct.
Future research should create a multi-year, multi-jurisdiction panel; use independent coders; validate the index with audit committees and regulators; and examine engagement-level evidence under confidentiality protections. Experiments can test whether the proposed documentation improves calibrated reliance. Field studies can compare financial audit with sustainability assurance. Research should also examine override direction: whether professionals challenge both unfavorable and favorable AI outputs with equal rigor. The public-document method and appendix supplied here provide a reproducible starting point. Extending the index to a multi-year, multi-jurisdiction panel would also allow disclosure trends to be tracked; a reasonable expectation is that AI-specific validation, explainability and data-governance disclosures will become more detailed as the EU AI Act takes effect and sustainability-assurance mandates mature, which the present single-year baseline is designed to measure against.