Next Article in Journal
The Mechanism and Spatiotemporal Variations in Digital Economy in Enhancing Resilience of the Cotton Industry Chain
Next Article in Special Issue
Occupational Gender Bias in Chinese Generative AI Models: Cross-Model Evidence of Stereotypical Amplification and Systematic Underrepresentation
Previous Article in Journal
Ontology Quality Improvement in the Semantic Web: Evidence from Educational Knowledge Graphs
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Stage-Aware Governance of Large Language Models: Managing Uncertainty and Human Oversight in AI-Assisted Literature Review Systems

1
School of Business, Konkuk University, 120 Neungdong-ro, Gwangjin-gu, Seoul 05029, Republic of Korea
2
Graduate School of Metaverse, Konkuk University, 120 Neungdong-ro, Gwangjin-gu, Seoul 05029, Republic of Korea
*
Author to whom correspondence should be addressed.
Systems 2026, 14(2), 153; https://doi.org/10.3390/systems14020153
Submission received: 29 December 2025 / Revised: 24 January 2026 / Accepted: 30 January 2026 / Published: 31 January 2026
(This article belongs to the Special Issue Ethics and Governance of Artificial Intelligence (AI) Systems)

Abstract

This study proposes a stage-aware governance framework for large language models (LLMs) that structures human oversight and accountability across different decision stages in AI-assisted literature review systems. Large language models (LLMs) are increasingly embedded in systematic review workflows, yet how human oversight and accountability should be structured across different decision stages remains unclear. This study evaluates three LLMs in a controlled two-stage literature review workflow—title-and-abstract screening and eligibility assessment—using identical evidence inputs and fixed inclusion criteria, with outputs benchmarked against expert consensus under fully reproducible conditions with standardized prompts and comprehensive logging. While LLMs closely matched expert decisions during screening (precision 0.83–0.91; F1 up to 0.89; Cohen’s κ 0.65–0.85), performance degraded substantially at the eligibility stage (F1 0.58–0.65; κ 0.52–0.62), indicating increased epistemic uncertainty when fine-grained criteria must be inferred from abstract-level information. Importantly, disagreements clustered in borderline cases rather than random error, supporting a stage-aware governance approach in which LLMs automate high-throughput screening while inter-model disagreement is operationalized as an actionable uncertainty signal that triggers human oversight in more consequential decision stages. These findings highlight the need for explicit oversight thresholds, responsibility allocation, and auditability in the responsible deployment of AI-assisted decision systems for evidence synthesis.
Keywords: large language models; AI governance; human oversight; literature screening; inter-model agreement large language models; AI governance; human oversight; literature screening; inter-model agreement

Share and Cite

MDPI and ACS Style

Kim, J.; Shin, H. Stage-Aware Governance of Large Language Models: Managing Uncertainty and Human Oversight in AI-Assisted Literature Review Systems. Systems 2026, 14, 153. https://doi.org/10.3390/systems14020153

AMA Style

Kim J, Shin H. Stage-Aware Governance of Large Language Models: Managing Uncertainty and Human Oversight in AI-Assisted Literature Review Systems. Systems. 2026; 14(2):153. https://doi.org/10.3390/systems14020153

Chicago/Turabian Style

Kim, Junic, and Haeyong Shin. 2026. "Stage-Aware Governance of Large Language Models: Managing Uncertainty and Human Oversight in AI-Assisted Literature Review Systems" Systems 14, no. 2: 153. https://doi.org/10.3390/systems14020153

APA Style

Kim, J., & Shin, H. (2026). Stage-Aware Governance of Large Language Models: Managing Uncertainty and Human Oversight in AI-Assisted Literature Review Systems. Systems, 14(2), 153. https://doi.org/10.3390/systems14020153

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop