1. Introduction
Many AI deployments now involve repeated interaction among people, organizations, and AI components rather than one-off model outputs. These systems shape what information is visible, what actions are feasible, and what consequences follow. It is therefore useful to view them as socio-technical multi-agent systems. In these systems, humans, organizations, automated assistants, recommendation and moderation components, and autonomous agents interact under partial observability and misaligned incentives.
Such systems exhibit familiar social-dilemma failure modes. Local incentives often favor short-term defection (free-riding, over-extraction of shared resources, or rule-breaking when undetected), while collective welfare depends on sustained cooperation [
1,
2]. Cooperative AI and multi-agent reinforcement learning (MARL) provide formal tools for studying these dynamics [
3,
4]. Yet a practical design question remains:
how should cooperation mechanisms be structured when they must also be acceptable, auditable, and contestable in human-facing settings?Deployed systems typically pursue cooperation through one of two broad strategies:
Behavioral influence: Systems steer behavior through interface design and personalization (nudges, friction, defaults, or hypernudges). Such interventions can be effective, but they may undermine autonomy and legitimacy when they become covert, manipulative, or difficult to contest [
5,
6,
7,
8,
9].
Accountability institutions: Systems establish enforceable norms through monitoring, record-keeping, adjudication, and contestable sanctions, ideally with due process and proportionality [
1,
10,
11,
12].
Both strategies carry ethical risks. Behavioral influence can improve outcomes while eroding agency. Institutional enforcement can support legitimacy while drifting toward surveillance or coercion if it is opaque, error-prone, or not meaningfully contestable [
13,
14,
15,
16]. This paper develops the second strategy—accountability institutions—as a framework for human–AI cooperation. The focus is on designs in which enforcement is transparent, auditable, and procedurally contestable (
Figure 1).
Three developments make institutional approaches timely for human–AI integration. First, agentic and tool-using systems are increasingly embedded in workflows such as planning, retrieval, execution, and verification. This creates repeated interaction among heterogeneous humans and AI components rather than isolated decisions. Second, these deployments operate in settings where formal rules already exist (policy compliance, safety constraints, access controls, and community standards), yet enforcement is often inconsistent or difficult to audit. Third, institutional mechanisms can themselves create harms when they are unaccountable: automated punishment with limited contestability can amplify error costs and distributional inequities, especially for vulnerable populations [
17,
18]. These considerations motivate an institutional layer whose monitoring and settlement actions are explicitly bounded by proportionality, transparency, and due process.
Section 2 frames cooperation mechanisms as governance institutions and motivates legitimacy-sensitive evaluation.
Section 3 positions IML relative to cooperative AI, behavioral influence, surveillance studies, and accountability infrastructure.
Section 4 develops the interaction-design requirements that guide the framework.
Section 5 and
Section 6 define IML formally and provide conservative incentive bounds under delayed settlement and imperfect monitoring.
Section 7 and
Section 8 translate institutional insights into design requirements and a broader evaluation agenda.
Section 9 presents the pilot SSD experiments.
Section 10 and
Section 11 discuss deployment vignettes and governance integration, and
Section 12 summarizes limitations and open problems.
This paper does not propose a new reward-shaping heuristic. Instead, it introduces an explicit mechanism class that externalizes enforcement into monitoring, evidence logging, review, and settlement. The paper’s main contribution is institutional: it formalizes and evaluates the accountability layer itself.
The paper combines four layers of contribution.
Section 2,
Section 3 and
Section 4 develop a conceptual and design-oriented framing.
Section 5 and
Section 6 formalize IML and derive conservative incentive checks.
Section 9 provides pilot simulation evidence in two canonical sequential social dilemmas. Human validation, broader substrate coverage, and deployment-oriented evaluation remain part of the future agenda in
Section 8 and
Section 12. Accordingly, the paper should be read as a framework paper with pilot evidence rather than as a fully validated empirical benchmark study.
1.1. Contributions
Framework and mechanism class: We formalize the Institutional Monitoring and Ledger (IML) augmentation for Markov games: monitor → evidence ledger → delayed settlement with review/appeal.
Interaction-design requirements: We outline design requirements for contestable AI institutions, grounded in transparency, procedural justice, and meaningful recourse.
Conservative incentive checks: We derive a one-shot deviation deterrence bound and extend it to imperfect monitoring and review, clarifying trade-offs among detection probability, adjudication delay, sanction magnitude, and wrongful penalties.
Institutional design translation: We translate governance insights from commons governance into concrete institutional parameters such as participatory norm-setting, graduated sanctions, contestability channels, and auditability requirements.
Pilot simulation study: We compare IML with vanilla PPO, Inequity Aversion [
19], Social Influence [
20], and IML ablations in two canonical sequential social dilemmas (
Harvest and
Cleanup) with
agents,
-step episodes, PPO training (
steps), and five training seeds per condition.
Agent-level burden analysis: Using IML ledger data, we quantify how flagged enforcement events, false-positive burden, overturn rates, and temporal burden shifts are distributed across agents.
Robustness and sensitivity additions: We add evaluation-seed robustness checks and a sensitivity analysis over institutional parameters, showing that measurement error can dominate welfare unless review and error control are calibrated carefully.
1.2. What This Paper Does Not Claim
This paper does not claim that monitoring and sanctions are sufficient for ethical governance, nor that institutional enforcement is always desirable. Norm specification is value-laden; institutions can be captured or misused; and monitoring can create dignitary harms even when it improves welfare. We also do not claim that benchmark simulations determine real-world acceptability. Our claim is narrower: if cooperation mechanisms are to be deployed in human–AI ecosystems, then transparent, auditable, and contestable accountability institutions should be treated as an explicit design option and evaluated on legitimacy dimensions—not only welfare or efficiency—with measurement error, proportionality, and due process treated as core design constraints.
2. From Incentive Design to Institutions: A Socio-Technical Framing
2.1. Why “Cooperation” Is an Institutional Question in Human–AI Systems
In MARL, cooperation is commonly pursued through reward shaping, learned social preferences, or centralized training. In socio-technical deployments, however, cooperation cannot be reduced to an optimization objective because it is inseparable from authority and legitimacy: who sets the norms, who monitors compliance, who adjudicates disputes, and what rights do affected parties have? These questions are institutional in the classical sense, and they become unavoidable once human stakeholders must live with, contest, or appeal the consequences of AI-mediated decisions.
Moreover, improved outcomes alone do not guarantee acceptance. Research on procedural justice emphasizes that perceived legitimacy and fairness predict compliance and cooperation even when outcomes are unfavorable [
10]. In algorithmic settings, people evaluate both distributive and procedural justice, and transparency interventions can have heterogeneous or even counterintuitive effects depending on what is revealed and to whom [
21,
22]. Related evidence on algorithm aversion suggests that technical superiority does not translate into adoption without trust, recourse, and legitimacy [
23]. These findings motivate a core thesis for cooperative human–AI systems: cooperation mechanisms are not merely technical instruments for increasing welfare; they are
forms of governance that must be constrained by societal values and evaluated as socio-technical institutions.
2.2. Internalizing Norms vs. Externalizing Enforcement
A recurring design choice in cooperative systems is whether to
internalize norms within the agent or to
externalize enforcement within an institution. Internalization modifies the agent’s objectives, preferences, or training signals so that prosocial behavior is intrinsically reinforced (for example via Inequity Aversion or other social preferences) [
19]. Externalization preserves the agent’s task objective while placing explicit accountability constraints around behavior through monitoring, adjudication, and settlement. Human societies routinely combine both modes: moral socialization and norm internalization coexist with external institutions such as courts, contracts, audits, and community governance [
1].
The distinction matters in human–AI integration because external institutional layers support governance properties that are difficult to guarantee through internalized objectives alone. First, externalization improves
auditability and contestability: an institution can publish evidence standards, document error rates, and report appeal outcomes, whereas internalized values can fail silently and may be hard to inspect or contest in practice [
12,
24]. Second, externalization better accommodates
pluralism and revision. Norms are contested and context-dependent; an institutional layer can encode procedures for changing norms and settlement rules through participation and oversight, whereas embedding “values” directly in objectives risks freezing contested assumptions and obscuring trade-offs [
25,
26]. Third, externalization creates a clearer
boundary against manipulation. Reward shaping, personalized influence, and choice architectures can drift into covert steering or “dark pattern” dynamics when incentives and interface design are not governable or contestable [
7,
8]. By making the locus of power visible—who monitors, what is logged, how disputes are settled, and on what grounds—institutions can be designed, audited, and held accountable.
This argument is not a rejection of internalization. Internal norms and prosocial preferences may reduce enforcement burden and enable smoother cooperation. The claim is narrower and governance-oriented: whenever cooperation is a public concern—involving power asymmetries, contested norms, or significant harms—external institutional layers offer governance advantages because the institutional layer can be made transparent, procedurally constrained, and contestable.
2.3. The Design-Space Boundary: Accountable Institutions vs. Covert Influence vs. Coercive Surveillance
Behavior-shaping mechanisms in human–AI systems occupy an ethically charged design space. One central axis concerns
transparency and contestability: whether affected parties can understand decisions, access reasons, and challenge outcomes or procedures [
11,
22]. A second axis concerns
intrusiveness and coercion: how much monitoring and sanctioning power is exercised, how broadly it is applied, and whether it remains proportionate to the harms being prevented [
1,
15]. A third, closely related dimension concerns
autonomy and manipulation risk: whether the mechanism preserves agency or subverts it through covert steering and exploitative choice architectures [
7,
27]. While this third dimension is not directly visualized in
Figure 1, it is conceptually important: many autonomy harms arise precisely when behavior change is opaque or difficult to contest, and thus the dimensions often co-vary in practice.
These axes help clarify why the two dominant deployment strategies can fail ethically in opposite directions. Nudges and hypernudges can be autonomy-compatible in some contexts yet become manipulative when hidden, personalized beyond reasonable expectations, or designed to exploit cognitive vulnerabilities [
5,
6,
7]. Conversely, surveillance-heavy enforcement can increase compliance while producing chilling effects and asymmetric power, particularly when monitoring is pervasive, error-prone, or insulated from meaningful oversight [
14,
16]. IML focuses on the part of this design space characterized by
high transparency and contestability with bounded intrusiveness, supported by explicit procedures for review/appeal and auditable institutional traces.
3. Related Work
IML sits at the intersection of (i) institutional theories of cooperation and sanctioning, (ii) open multi-agent governance and cooperative AI mechanism design, and (iii) ethical evaluation and accountability infrastructure for human-facing deployments. We therefore review the literature that jointly motivate institutional cooperation mechanisms and clarify how they should be evaluated when they are embedded in socio-technical systems.
3.1. Institutions for Cooperation: Commons Governance, Sanctioning, and Open Systems
A foundational literature in evolutionary and behavioral accounts of cooperation emphasizes kin selection, reciprocity, and reputation as stabilizing forces in repeated interaction [
28,
29,
30,
31]. In human societies, however, durable cooperation in commons dilemmas often depends on explicit institutions that define boundaries, monitor behavior, apply graduated sanctions, and provide low-cost dispute resolution [
1]. Experimental economics similarly shows that punishment can increase cooperation but also raises non-trivial welfare, legitimacy, and distributional questions [
32]. These results motivate a view of monitoring and settlement as
mechanism primitives that stabilize cooperation precisely when direct interpersonal enforcement is impractical.
Digital platforms and marketplaces have long relied on
reputation systems to support trust under anonymity and scale [
33]. Reputation, however, is only one instrument: it aggregates past behavior but may lack due process, can be gamed, and can amplify visibility and power asymmetries. In parallel, the “electronic institutions” tradition in multi-agent systems models the rules, roles, and permissible interaction protocols that structure open agent societies [
34]. The key move in this tradition is to treat interaction rules and enforcement procedures as part of the system specification rather than as an emergent equilibrium artifact.
IML is compatible with both lines while emphasizing a different governance bundle: it can incorporate reputation-like summaries in an auditable ledger, but it pairs them with structured evidence, explicit adjudication, and contestability. It can also be understood as a lightweight institutional wrapper that can govern modern learning agents without requiring shared internal architectures, thereby separating institutional authority from agent design.
3.2. Cooperative AI and Sequential Social Dilemmas: Internalization Versus Institutional Layers
In cooperative AI and MARL, sequential social dilemmas (SSDs) formalize how mixed incentives arise at the policy level in Markov games [
2]. Benchmark suites such as Melting Pot and Melting Pot 2.0 broaden this landscape and stress-test generalization to new partners and norms [
4,
35]. Within MARL, cooperation is frequently pursued by
internalizing norms into agent objectives or representations, including social-preference shaping such as Inequity Aversion [
19], intrinsic motivation for social influence [
20], and explicit contract modifications [
36]. The Cooperative AI agenda argues that designing systems to cooperate with humans and other agents is a central objective for human-compatible AI [
3]. Normative multi-agent systems research similarly treats obligations, permissions, and sanctions as explicit components of agent societies [
37].
IML targets a complementary axis. Rather than primarily changing the agent’s intrinsic objective, it preserves base rewards and introduces a separate institutional layer that can be audited and governed. This externalization matters for human-facing deployments because it makes enforcement procedures and their error properties explicit (e.g., false positives, review outcomes), enabling contestability and oversight while leaving open the choice of learning architecture for the agents that operate within the institution.
Recent work has further expanded the landscape of cooperation mechanisms. Dell’Anna et al. [
38] propose runtime revision of sanctions in normative multi-agent systems, allowing the system to adapt its enforcement over time. Centeno et al. [
39] develop adaptive sanctioning mechanisms for open MAS regulated by norms. Christoffersen et al. [
40] show that formal contracts—voluntary zero-sum reward modifications—can mitigate social dilemmas in MARL. These approaches share IML’s emphasis on explicit institutional mechanisms but differ in their treatment of human oversight and contestability. The novelty of IML is to make that human-facing accountability layer—not only the sanction logic—the central object of formalization, analysis, and evaluation.
3.3. Contestable AI and Human–AI Interaction
As AI systems are deployed in high-stakes domains, the need for accountability and contestability has become paramount. Contestability—the ability of users to understand and challenge algorithmic decisions—is increasingly recognized as a fundamental requirement for legitimate AI governance [
41]. Recent empirical work by Yurrita et al. [
42] identifies the specific needs of algorithmic decision subjects for meaningful contestability, emphasizing that the right to contest must be accompanied by practical capacity to do so. Guidelines for human–AI interaction emphasize the importance of clear feedback mechanisms and multimodal interfaces [
43]. The IML framework operationalizes these principles in the context of multi-agent systems, providing a structured interaction design for monitoring, logging, and appealing AI behaviors.
3.4. Ethical Risk Boundaries: Opacity, Manipulation, Surveillance, and Legitimacy
Institutional mechanisms can fail ethically even when they improve aggregate outcomes, especially when technical abstractions omit social context or procedural constraints. Opacity arises for reasons beyond model complexity—including intentional secrecy, mismatches between technical representations and social meaning, and differences in what stakeholders can understand or contest [
24]. Related critiques of technical “fairness” highlight how abstraction and modularity can sever interventions from the normative goals they claim to serve [
25]. For institutional cooperation mechanisms, this implies that “monitoring” and “violations” cannot be treated as purely technical facts; they are socio-technical constructs that require contextual justification, procedural safeguards, and governance.
A parallel set of concerns arises from behavioral influence. Choice architecture and nudges can shape behavior without formal coercion [
5], and digital personalization enables dynamic “hypernudges” at scale [
6]. Ethical critiques emphasize that influence becomes problematic when it is hidden, exploitative, or substitutes the designer’s judgment for the individual’s deliberation [
7,
27]. HCI research documents manipulative “dark patterns” in real systems [
8,
9]. These works in the literature are relevant for IML because cooperation mechanisms can be implemented either as transparent institutions or as opaque influence systems; IML explicitly targets the former and treats transparency and contestability as design requirements rather than optional features.
Surveillance studies further emphasize that monitoring infrastructures are not merely data-collection mechanisms but power structures with implications for autonomy, inequality, and citizenship [
14,
15]. Classic accounts of disciplinary power highlight how being potentially observed can induce self-regulation and reshape subjectivity [
16]. In contemporary AI systems, these concerns manifest as chilling effects, function creep, and mission creep, where monitoring introduced for safety or cooperation expands into broader control. Accordingly, IML treats privacy, proportionality, and governance constraints as primary design parameters rather than afterthoughts.
Finally, legitimacy and procedural justice predict compliance and acceptance beyond deterrence [
10]. In algorithmic decision-making, stakeholders evaluate both distributive and procedural justice, and transparency interventions interact with perceived outcome control [
21,
22]. Algorithm aversion further suggests that errors by algorithms can reduce trust even when algorithms outperform humans [
23], and resistance to AI is domain-dependent and especially pronounced in high-stakes contexts [
44]. These findings motivate an evaluation agenda that treats legitimacy, contestability, and error harms as important outcomes for institutional cooperation mechanisms.
3.5. Accountability Infrastructure, Audits, and Standards for Governance
The growing literature proposes technical and organizational infrastructure for accountability, including internal auditing practices [
12], accountable algorithmic systems [
11], and persistent recording mechanisms, such as ethical “black boxes” [
45]. Algorithmic impact assessments are proposed as pre-deployment accountability mechanisms [
46,
47], and external audit methodologies support third-party scrutiny [
48]. Across governance practice, high-level ethics principles are widespread but insufficient without institutionalization and enforceable processes [
26,
49]. This motivates a shift from purely aspirational principles toward operational artifacts: logs, review procedures, audit trails, and contestable decisions.
Standards and regulatory frameworks increasingly require risk management, transparency, and accountability (e.g., NIST AI RMF, ISO/IEC 23894, OECD principles, EU AI Act) [
50,
51,
52,
53]. Complementary perspectives in human–AI interaction emphasize human-centered design and meaningful human control [
54,
55,
56]. IML is intended to be compatible with these governance directions by making monitoring and settlement explicit, auditable, and procedurally constrained and by enabling parameterized trade-offs among cooperation, error costs, and institutional overhead.
4. Interaction Design for Contestable AI Institutions
An accountability mechanism is not only a formal object. In human-facing systems, it also depends on
interaction design: how the institution explains decisions, surfaces evidence, and supports recourse. If users cannot understand the rules being enforced, inspect the relevant evidence, or contest outcomes, the institution does not provide meaningful accountability. This section summarizes the interaction-design requirements that follow from that view, drawing on recent advances in contestable AI [
41,
42] and human–AI interaction design [
43].
4.1. Transparency Through Auditable Interaction Traces
Transparency in human–AI interaction requires more than open-source code or explainable models; it requires an accessible, structured record of institutional actions and their justifications. The IML framework operationalizes transparency through the auditable ledger, which serves as the primary interaction surface between the institution and its stakeholders. Each ledger entry records the detected event, the evidence supporting the detection, the resulting institutional action (sanction, review, or overturn), and the timestamps of each step. From an interaction design perspective, this ledger must present information in a format that is comprehensible to diverse stakeholders—combining structured textual logs with visual evidence such as state snapshots, trajectory replays, or annotated timelines. The multimodal presentation of institutional traces is essential: purely textual logs are insufficient for conveying the spatial and temporal context of multi-agent interactions, while purely visual representations may lack the precision needed for formal appeals.
4.2. Meaningful Contestability as an Interaction Pattern
Contestability—a user’s practical ability to challenge an algorithmic decision—is increasingly recognized as a core requirement for legitimate AI systems [
41]. In the IML framework, contestability is operationalized through the
review and appeal channel, which transforms the institution from a unilateral enforcement mechanism into a bidirectional interaction. Designing for meaningful contestability requires attention to three interaction properties:
Clear notification: Stakeholders must be promptly and clearly notified when a violation is detected and a sanction is pending, with sufficient context to understand the allegation.
Accessible evidence: Stakeholders must have easy access to the evidence recorded in the ledger, presented in a format that supports informed evaluation of the detection’s accuracy.
Low-friction appeal process: The interface for submitting an appeal must be intuitive and not overly burdensome, ensuring that the right to contest is practically accessible rather than merely theoretical.
Recent empirical work on contestability [
42] emphasizes that affected individuals need not only the
right to contest but also the
capacity—including cognitive resources, time, and interface support—to do so effectively. This motivates the design of appeal interfaces that actively support the user rather than simply receiving appeals.
4.3. Procedural Justice in Institutional Interaction
Procedural justice—the perceived fairness of the processes used to make decisions—is a well-established predictor of compliance, trust, and legitimacy in both human institutions and algorithmic systems [
10,
22]. In human–AI interaction, procedural justice is closely tied to the temporal and communicative structure of the accountability mechanism. The IML framework supports procedural justice through several interaction design choices:
Delayed settlement: Sanctions are not applied immediately or automatically. The built-in settlement delay creates a deliberative window during which the stakeholder can review evidence and initiate an appeal, transforming the interaction from a punitive algorithmic action into a structured institutional process.
Graduated response: Sanctions are proportional to the severity of the violation and the confidence of the detection, avoiding the “one-size-fits-all” penalties that undermine perceived fairness.
Outcome communication: The result of each institutional action—whether a sanction is upheld, reduced, or overturned—is communicated back to the stakeholder with an explanation, closing the interaction loop.
These interaction design principles are not merely aspirational; they have direct implications for the engineering of IML systems. As the empirical study later illustrates (
Section 9), the accumulation of measurement error over long horizons can create an “institutional tax” that materially affects welfare outcomes. The interaction design of the review and appeal channel determines how much of this tax can be mitigated through human oversight, making contestability not just an ethical requirement but a functional requirement for deployable systems.
5. Formal Framework: Institutional Monitoring and Ledger
This section specifies Institutional Monitoring and Ledger (IML) at two levels: as a formal augmentation of an
n-agent Markov game and as a socio-technical institution with roles, interfaces, and explicit governance degrees of freedom.
Figure 2 summarizes these components and the main information flows among the environment, the institution, and the participants.
5.1. Base Markov Game and Norm Specification
We consider an n-agent Markov game with state space , joint action space , transition kernel , and per-agent base rewards . Each agent i executes a policy given observation (full observability is a special case). Returns are discounted with .
To make “violations” explicit (and therefore auditable), we assume a
norm specification that induces an operational predicate over trajectory prefixes. Let
denote a trajectory prefix up to time
t (e.g., including the relevant state/action/observation history available for monitoring). A minimal form is a predicate
which indicates whether a norm-relevant event occurred at time
t (e.g., unauthorized access, harmful action, exceeding a quota, deception in a protocol). This separates
norm definition (a governance choice) from
norm enforcement (the mechanism studied here). Following critiques of abstraction [
25], we treat
v as socially situated: its definition must be justified, contestable, and revisable.
5.2. Definition: IML Augmentation
Definition 1 (IML augmentation)
. An Institutional Monitoring and Ledger (IML) is an augmentation that introduces: (i) a monitor
M that maps a trajectory prefix (and the norm specification) to an evidence event , where ⌀
denotes “no evidence” under imperfect detection; (ii) a ledger state that records or summarizes evidence and institutional actions over time (an audit log); (iii) an adjudication and settlement rule Σ
that, after a delay D, maps ledger-supported evidence to a vector of transfers/penalties (including the possibility of no action), and (iv) a contestability interface that enables inspection, explanation, and dispute of evidence and outcomes. The environment emits modified rewards, for each agent i, IML is a wrapper: it need not change the base dynamics P or base rewards . It can implement sanctions (negative transfers), restitution (compensating harmed parties), or bounded “taxes” that fund shared resources. The ledger makes enforcement auditable: both evidence and settlement logic can be inspected, challenged, and governed.
5.3. IML as an Institutional Wrapper: Roles and Interfaces
In human–AI deployments, IML must be specified not only by its technical components but also by the roles that make contestability meaningful. We assume
participants (users and agents) whose actions are subject to norms; an
institution operator that runs
M, maintains the ledger, and executes
;
auditors or oversight bodies (internal or external) that can inspect the ledger and evaluate compliance with governance constraints [
12,
48]; and
adjudicators (human reviewers or hybrid processes) that handle contested cases. These roles determine what “contestability” means in practice: who can access evidence, what explanations are provided, what time bounds apply, and what remedies are available.
5.4. Operational Pseudocode: Wrapping an Environment with IML
While IML is a conceptual institution, it can be implemented as a wrapper around any interactive system (simulator, platform workflow, robotic controller) that exposes events and permits institutional feedback. Algorithm 1 sketches a minimal wrapper that logs evidence and applies delayed settlement, making explicit the separation between base rewards and institutional adjustments.
| Algorithm 1 Institutional Monitoring and Ledger (IML) wrapper for a Markov game |
![Mca 31 00069 i001 Mca 31 00069 i001]() |
5.5. Design Degrees of Freedom (Governance Knobs)
IML exposes parameters that are meaningful both technically and institutionally. Monitoring can vary in scope and privacy posture: which signals are observed, whether monitoring is local or centralized, and whether privacy-preserving methods are used, ideally aligned with contextual integrity [
13]. Monitoring also has detection-quality parameters (true/false positive rates) and must be evaluated for robustness to strategic gaming of evidence channels. Settlement introduces explicit timing and procedure: the delay
D, the form of
(automatic fines, hybrid processes, or human-in-the-loop adjudication), and the error properties induced by these choices. Sanction design raises boundedness and proportionality concerns, including the use of graduated sanctions and calibrated magnitude constraints [
1]. Finally, contestability is not a binary feature but a design space spanning explanation style, evidence access, appeal pathways, and remedies [
11,
22]. In socio-technical deployments these are not merely hyperparameters: they are policy decisions that should be documented, justified, monitored, and revisable under governance processes [
50,
51].
Figure 3 gives the corresponding operational wrapper view, showing how actions, observations, rewards, and governance constraints interact in one deployment loop.
5.6. Threat Model and Failure Modes
An accountable institution should be designed against foreseeable failures, including failures of
normative choice and failures of
implementation. Wrong norms and value capture occur when
v encodes contested values without participation or when institutional authority is captured, thereby entrenching power. Opacity and abstraction failures arise when evidence and settlement are difficult to inspect or contest, collapsing accountability claims even if the mechanism improves aggregate outcomes [
24,
25]. Surveillance drift and chilling effects occur when monitoring scope expands beyond the original purpose or becomes normalised as routine control [
14,
15]. Error and wrongful punishment arise when false positives or biased evidence pipelines generate unjust sanctions and erode legitimacy, especially under long-horizon accumulation. Finally, strategic gaming occurs when agents exploit blind spots, manipulate evidence channels, or displace harms outside the monitored interface. These failure modes motivate the design principles and evaluation stress tests developed in
Section 7 and
Section 8.
6. Incentive Analysis: Deterring Deviations Under Delayed Settlement
This section develops conservative, auditable incentive checks for IML. The results below are one-shot sufficient conditions for local deterrence under fixed opposing policies; they are not a full equilibrium or repeated-deviation analysis for partially observable Markov games. Their purpose is to make the central trade-offs legible to designers, auditors, and affected parties: deterrence requires detection and timely settlement, while legitimacy requires error control and meaningful review.
6.1. One-Shot Deviation Deterrence Bound
Let
denote a candidate cooperative joint policy in the base game
. Consider the standard one-shot deviation test: at some state
s visited under
, one agent deviates for a single step and then all agents revert to
thereafter. Define the maximal one-shot deviation advantage in the
base game as
where
is the discounted action-value function in the base game when, after the one-step deviation, play returns to
. Importantly,
upper-bounds
total discounted gain from the deviation, not only the immediate reward difference.
IML introduces a monitor-and-settlement pipeline. Let be a lower bound on the probability that a norm-violating deviation at s results in an institutional penalty being ultimately upheld and applied (i.e., after any review), and let be a lower bound on the magnitude of that penalty. Let be the (deterministic) settlement delay in steps and the discount factor.
Proposition 1 (Sufficient condition for deterring one-shot deviations)
. Fix and consider with discount factor . If, for all relevant states s on trajectories induced by ,then no agent can improve its discounted return by a one-shot deviation from at any such state, holding other agents fixed and then reverting to . Proof. By definition, the deviation yields discounted gain at most
. With probability at least
p, the institutional pipeline applies a penalty of magnitude at least
K after
D steps, which has discounted value at least
. Hence the deviation’s expected discounted cost is at least
. Condition (
2) implies expected cost weakly exceeds maximal gain, so the deviation is not profitable. □
Condition (
2) makes explicit a three-way trade-off: weaker upheld-detection (
) requires either larger penalties (
) or shorter delays (
). In human-facing systems, increasing
K is ethically and legally constrained; thus, governance-relevant design effort should prioritize improving evidence quality (raising the probability that true violations are upheld) and reducing unnecessary delay.
If settlement delay is stochastic (e.g., due to queueing for review), the same argument yields the sufficient condition , which separates procedural latency (a governance design variable) from deterrence strength.
6.2. Imperfect Monitoring, Review Accuracy, and Wrongful Punishment
Real institutions face both false negatives and false positives. To make the review/appeal channel explicit, distinguish raw monitoring from post-review settlement. Let denote a lower bound on the monitor’s true-positive rate (flagging when a violation occurs), and let denote an upper bound on the monitor’s false-positive rate (flagging when no violation occurs). Let the review/appeal procedure have two accuracy parameters: is a lower bound on the probability that a false accusation is overturned (true-negative correction), and is an upper bound on the probability that a valid accusation is wrongly overturned (false-negative due to review). Penalties are applied only after review.
Under these definitions, the effective probability that a true violation yields an upheld penalty is at least
while the effective probability that a
non-violation yields a wrongful upheld penalty is at most
Substituting (
3) into Proposition 1 shows how improved review (smaller
) strengthens deterrence for fixed monitoring, while substituting (
4) quantifies how review (larger
) reduces wrongful punishment.
For a compliant agent, a conservative upper bound on the expected discounted
wrongful penalty from a single time step is
In long-horizon deployments, the cumulative wrongful-penalty burden can dominate welfare even when each step’s false-positive probability is small. Over a horizon
T with approximately stationary rates, a conservative bound on the cumulative discounted wrongful penalty is
and in the near-undiscounted regime (
) this scales approximately linearly with
T (a “sanction tax” from measurement error). This makes false positives a primary design constraint rather than a secondary nuisance.
To bound wrongful punishment risk, a policy-maker can impose an “error budget”
and require
or, for long-horizon settings, an episode-level constraint using (
6). Together with the deterrence condition (
2) (with
), these inequalities define a feasible region in
: increasing
K may improve deterrence but worsens the impact of residual false positives, whereas improving monitoring quality (higher
, lower
) and strengthening due process (higher
, lower
) can improve both deterrence and legitimacy.
6.3. Institutional Delay as a Governance Choice
Delay D is often treated as an implementation artifact, but in human-facing institutions it is a governance variable that mediates the tension between deterrence, due process, and perceived arbitrariness. Shorter delays strengthen deterrence through reduced discounting but can reduce opportunities for explanation, contestation, and careful adjudication. Longer delays can support due process, evidentiary development, and appeals but weaken deterrence and may be experienced as opaque or punitive if outcomes arrive after participants can no longer connect them to actions.
From an institutional-design perspective,
D should therefore be justified relative to the speed of harm accumulation, the operational feasibility of review (including capacity constraints), and the procedural requirements for contestability and remedy. This aligns with broader critiques that ethics principles require translation into operational institutional practice rather than remaining aspirational [
26].
7. Design Principles for IML as an Accountable Institution
This section translates institutional theory and human–AI ethics into concrete requirements for IML. The aim is not to offer an exhaustive moral theory. Rather, it identifies design commitments that are (i) operationally checkable, (ii) compatible with institutional governance practice, and (iii) responsive to the legitimacy risks of monitoring and sanctioning in human-facing systems.
7.1. Principle Set
We propose nine design principles for IML in human–AI integration, stated as institution-level requirements rather than agent-level desiderata.
Legitimacy-sensitive evaluation: IML should be evaluated and designed for acceptance, perceived fairness, and procedural justice, not only welfare or efficiency. Procedural legitimacy can predict compliance even when outcomes are unfavorable, and algorithmic systems are judged on both distributive and procedural dimensions [
10,
22].
Transparency and explainability of the institution: IML should disclose what is monitored, how settlement decisions are made, and what the institution’s error properties are in practice; it should also provide intelligible explanations for individual decisions. This requirement applies to the
institutional process (rules, thresholds, escalation, review) rather than only to underlying models [
11,
21].
Contestability and recourse: IML should provide meaningful avenues for contesting evidence and outcomes, including appeal pathways, timely resolution, and (where appropriate) outcome control. Contestability is not a cosmetic feature; it is a structural condition for procedural justice and for limiting wrongful punishment [
22,
57].
Proportionality and bounded sanctions: Institutional responses should be proportionate to the stakes and severity of violations, should prefer graduated sanctions and restitution where feasible, and should treat sanction magnitude as bounded (cap
K) to avoid coercive escalation. This aligns with institutional insights on graduated sanctions and legitimacy in commons governance [
1].
Contextual integrity and data minimization: Monitoring should be limited to norm-relevant signals, with data flows justified relative to context, role, and expectations; retention and access should be minimized and governed. This requirement treats privacy not as an optional add-on but as a condition for legitimate monitoring [
13].
Robustness to gaming and feedback effects: IML should be stress-tested against strategic adaptation, evidence-channel manipulation, and Goodhart-style dynamics in which proxies become targets. Robustness requires both technical adversarial testing and institutional monitoring for drift and unintended incentives.
Non-manipulation: IML should avoid covert influence mechanisms that substitute transparent accountability with hidden interface steering, especially when persuasion and enforcement blur. Where behavior change is sought, the locus of power should remain visible and governable rather than personalized and opaque [
7,
8].
Participatory norm-setting: Because norm specification is value-laden, IML should incorporate participatory processes for defining and revising norms and governance rules, consistent with value-sensitive design and with the need to handle pluralism and dissent [
58].
Auditability and accountability artifacts: IML should maintain logs and documentation that support internal and external audits, impact assessments, and reproducible incident reconstruction, including versioning of monitors, rules, and settlement logic [
12,
46,
47].
7.2. Operationalizing the Principles: Metrics and Documentation Hooks
High-level principles often fail in practice when they are not translated into observable commitments and auditable artifacts [
26]. For IML, we recommend treating each principle as a set of operational questions with measurable proxies and documentation hooks so that different forms of review—technical, organizational, and ethical—can evaluate the same institutional design on commensurable evidence.
Table 1 translates the institutional principles into measurable commitments and documentation artifacts.
This operational stance supports interdisciplinary evaluation. Technical reviewers can inspect detection error rates, robustness tests, and versioning; organizational reviewers can inspect incident workflows and audit readiness; and ethics reviewers can inspect legitimacy measures, participation, and the existence (and accessibility) of redress.
7.3. Mapping Commons Governance Principles to IML
Commons governance provides a mature vocabulary for legitimate cooperation institutions.
Table 2 maps core design principles [
1] to IML requirements in human–AI systems, emphasizing that enforcement mechanisms must be embedded in procedures that constrain mission creep, support contestability, and enable revision under pluralism.
7.4. When Not to Deploy IML
IML is inappropriate when norms cannot be specified without unacceptable value conflict, when monitoring would violate contextual integrity or create unacceptable abuse potential, when sanctions would be coercive relative to the stakes, or when governance capacity (audits, appeals, oversight) does not exist in practice. In such cases, the more appropriate response is often to redesign the system to reduce the need for enforcement—for example by changing incentives, limiting harmful affordances, or restructuring the interaction—rather than strengthening monitoring and sanctioning.
8. Broader Evaluation Agenda: Welfare and Legitimacy
A full evaluation of cooperation mechanisms for human–AI integration requires more than reporting “higher cooperation rates.” Because IML is an institutional intervention, the broader evaluation agenda should separate two axes: (A) welfare and safety outcomes (how the system behaves and what harms it prevents) and (B) legitimacy and acceptability outcomes (whether those interventions are perceived as fair, contestable, and compatible with autonomy). Keeping these axes separate is methodologically important: mechanisms that improve aggregate welfare can still be rejected if they lack procedural justice, and procedures that increase legitimacy can alter behavior even when outcomes are unchanged. The pilot simulations in
Section 9 speak primarily to part A and to mechanism-level proxies related to part B; legitimacy in human-facing settings requires the methods outlined in
Section 8.3.
8.1. Benchmark Suites, Baselines, and Ablations
For simulation-based evidence, we recommend beginning with mixed-motive Markov games that expose commons dilemmas and conflict dynamics. Sequential social dilemmas (SSDs) provide a canonical substrate for this purpose [
2], and broader suites such as Melting Pot and Melting Pot 2.0 extend evaluation to diverse substrates and cross-play protocols that stress-test generalization to new partners and norms [
4,
35]. Where the research question concerns explicit agreement structures, contracting domains offer complementary benchmarks for testing contract-style augmentations [
36]. Across these settings, IML should be compared not only to “no institution” (vanilla MARL) but also to prominent alternatives that internalize cooperation within agents, including social-preference shaping such as inequity aversion [
19], intrinsic motivation for social influence [
20], and explicit contract augmentation [
36].
Beyond headline comparisons, institutional mechanisms require ablations that expose where performance comes from and where it fails. For IML, the core ablation family varies monitoring and settlement parameters (or their implementation-level analogues), as well as privacy constraints (monitoring scope and retention) and contestability features (appeals, explanation interfaces, evidence access). Because institutional interventions are strategic objects, evaluation should also include adversarial or “gaming” agents that attempt to exploit monitoring blind spots, manipulate evidence channels, or Goodhart the monitored proxy.
8.2. Outcome Families
Table 3 summarizes a minimal outcome taxonomy. Welfare/safety outcomes capture efficiency and harm prevention; legitimacy/acceptability outcomes capture procedural justice, autonomy, and perceived rightfulness. A methodological warning from socio-technical “fairness” research applies: these outcome families should not be collapsed into a single scalar objective without explicit normative justification [
25]. In practice, reporting should therefore include both mean performance and distributional summaries (e.g., inequality, tail risk, and agent-level burden allocation), as well as explicit error metrics for monitoring and review (false positives/negatives and their costs) since error burdens are central to legitimacy. In the simulation results below, we use
reliability narrowly to refer to bounded cross-seed variability, avoidance of unstable training collapse, and sensitivity to enforcement error; claims about legitimacy or acceptability remain outside the scope of simulation and require the human-subject methods outlined in
Section 8.3.
8.3. Human-Subject Studies and Mixed Methods
For human-facing deployments (e.g., moderation assistance, collaborative tools, shared-resource management), legitimacy cannot be inferred from simulation alone. Controlled studies should therefore vary institution design factors that plausibly shape procedural justice and acceptance. At minimum, the study design should manipulate transparency (disclosed versus opaque monitoring scope and explanation styles), recourse and outcome control (no appeal versus contestable decisions with meaningful redress) [
22], sanction proportionality (low versus high penalties, including graduated sanctions) [
1], and monitoring intrusiveness (minimal versus expansive evidence collection, aligned with contextual integrity and surveillance concerns) [
13,
15]. These factors are precisely those that can shift an institutional mechanism from “accountable enforcement” toward either covert influence or coercive surveillance.
Three complementary empirical designs are especially useful for separating behavioral effects from legitimacy perceptions. Interactive laboratory tasks can instantiate repeated cooperation problems in which participants interact with other humans and/or AI agents under randomized institutional procedures; this design can reveal whether identical welfare trajectories are judged differently under different procedural regimes. Vignette-based surveys can scale quickly by presenting realistic scenarios (e.g., content takedown, workplace compliance flag, robotics incident) and manipulating evidence disclosure, explanation quality, and redress; vignettes are particularly valuable for probing cultural and contextual heterogeneity, but should be triangulated with interactive evidence. Where feasible, field and organizational studies can analyze real appeal logs, audit outcomes, and stakeholder interviews in settings where institutions already produce records (e.g., security logs, moderation workflows), improving external validity and surfacing institutional dynamics such as power asymmetries, informal workarounds, and governance capacity constraints.
A rigorous analysis plan should treat both welfare and legitimacy outcomes as primary endpoints. We recommend preregistering hypotheses when possible, reporting effect sizes and uncertainty, and publishing ablations that identify which institutional features drive which effects. Because institutions distribute burdens, reporting should include distributional analyses and ethically appropriate subgroup analyses, and should document how monitoring scope and retention choices were communicated to affected parties. Mixed-method work—including interviews, analysis of disputes, and institutional ethnography—is important for capturing the gap between formal mechanism design and lived governance practice, especially where coercion, dignity, and autonomy concerns arise.
8.4. Auditability and Organizational Stress Tests
Because IML is explicitly a governance mechanism, evaluation should include organizational stress tests in addition to behavioral metrics. At a minimum, the institution should be tested for strategic gaming: whether agents can exploit monitoring blind spots, manipulate evidence channels, or shift harms outside the monitored interface. Auditability should be evaluated as a functional capability: whether auditors can reconstruct why a settlement occurred, including the evidence relied upon, the rule invoked, and the review outcome, consistent with accountability-oriented auditing practices [
11,
12]. Finally, the evaluation should produce documentation artifacts analogous to algorithmic impact assessments, including versions of norms, monitors, settlement rules, and observed error rates, so that accountability claims are tied to concrete records rather than aspirational principles [
46,
47].
9. Empirical Evaluation on Sequential Social Dilemmas
We complement the conceptual framing and conservative incentive analysis with a focused pilot study in a canonical class of multi-agent reinforcement learning (MARL) benchmarks:
sequential social dilemmas (SSDs) [
2]. The aim is mechanism illustration rather than broad empirical validation. Specifically, we use these environments to show how an auditable institutional layer—with explicit monitoring, sanctioning, and review—changes learning dynamics, welfare, inequality, agent-level burden, and policy exploration relative to standard MARL baselines and representative alternative cooperation mechanisms.
9.1. Experimental Design and Baselines
We evaluate two SSD environments: Harvest, a renewable resource dilemma, and Cleanup, a public-goods dilemma. Both environments include an action that can directly harm others (the “beam”). Our institution targets this action via the deliberately narrow rule no_punishment_beam. The rule is behaviorally unambiguous by construction and is used here to isolate enforcement error and contestability; it does not resolve the broader problem of contested norm specification.
All experiments use
agents and fixed-horizon episodes of
steps. Agents are trained with a parameter-shared PPO implementation [
59] for 2 × 10
6 environment steps per run. We run five independent training seeds per condition, resulting in 70 total training runs. The seed count is still modest, so the experiments should be read as a focused pilot evaluation rather than a definitive statistical benchmark. Accordingly, we emphasize uncertainty intervals, ablations, and full seed-level reporting, and we report exploratory inferential comparisons later in this section.
Table 4 summarizes the setup used in the SSD study.
Appendix A provides the full training and evaluation configuration, and
Appendix C reports the evaluation-only sensitivity sweep over false-positive rate and review probability.
We compare seven conditions across two broad categories:
These baselines are intended to be representative rather than exhaustive. PPO provides the no-institution reference, while IA and SI represent two established ways of internalizing cooperative pressure within the agent. Other alternatives, including contract-based and additional norm-aware methods, remain outside the scope of the present pilot study and are noted again in
Section 12.
9.2. Learning Dynamics and Baseline Comparisons
Figure 4 summarizes the main experimental results across all conditions. A notable pattern is the difficulty of making internal reward-shaping baselines behave robustly in these long-horizon settings. Inequity Aversion (IA) exhibits severe value-loss instability (value loss > 50,000 in both environments;
Table 5), suggesting that the dense, non-stationary intrinsic penalty for inequity destabilizes the PPO value function. Social Influence (SI) collapses to low-entropy policies in both environments. In
Cleanup, it still attains high mean return with low inequality. In
Harvest, SI also attains a much higher mean episode return than Baseline or IML, but with extremely large cross-seed dispersion, as shown later in the held-out evaluation summary table. We therefore treat SI as a useful behavioral comparison rather than as a stable or accountability-oriented substitute for the institutional layer.
In contrast, the external institutional wrapper (IML) maintains stable learning while materially altering policy exploration. Here, “stable” is used narrowly to refer to the absence of the severe optimization pathologies seen in IA and to less collapse into near-deterministic policies in these runs. As shown in
Figure 5, Baseline PPO agents rapidly collapse to deterministic policies (entropy
in
Cleanup,
in
Harvest). IML agents, however, sustain substantially higher entropy throughout training (
in
Cleanup,
in
Harvest;
Table 5). This pattern suggests that the institutional wrapper helps prevent collapse into rigid strategies and maintains a more exploratory joint policy.
9.3. Ablation Study: Isolating Institutional Mechanisms
To understand which components of IML drive these effects, we conducted an ablation study (
Table 6 and
Figure 6).
The results clearly isolate the causal mechanisms of the institution:
Observation is insufficient: The Monitor Only condition closely tracks the Baseline (entropy vs. ). Simply logging violations without enforcement does not materially alter agent behavior.
Sanctions drive exploration: The Sanction (No Review) condition achieves the vast majority of IML’s effect (entropy in Cleanup, in Harvest). The threat of external penalty forces agents to maintain more stochastic policies to avoid predictable violations.
Review modulates the effect: Adding the review mechanism (Full IML) slightly increases entropy in Cleanup () but decreases it in Harvest (). Increasing the review probability to (High Review) further modulates this effect. This suggests that the contestability interface is not merely a post hoc fairness mechanism, but an active institutional parameter that shapes system behavior.
9.4. Evaluation Outcomes and the Institutional Tax
While IML improves policy robustness and exploration, it introduces a structural cost under imperfect monitoring.
Figure 7 shows the episode returns and Gini coefficients for Baseline, SI, and IML.
In
Cleanup, Baseline converges to an approximately zero-return regime. Under IML, the held-out evaluation mean shifts to
per agent, and the last-200-episode training summary in
Table 7 shows a similar
. This welfare loss is explained primarily by false-positive sanctioning. Because monitoring noise is applied per agent-step, the expected number of false detections per episode scales linearly with the horizon
T. With
,
, and
, the expected false positives are ≈100 per episode. Even with a
review rate overturning a share of erroneous flags, the remaining upheld actions impose a persistent “institutional tax” on the population.
9.5. Interpretation and Interaction Design Implications
Taken together, these pilot results support a narrower claim. In these two SSD settings, the external IML wrapper was easier to stabilize than the representative internalization baselines tested here, but it also introduced a measurable error-mediated welfare cost. The experiments therefore illustrate how institutional design parameters become part of the learning problem. They do not establish broad substrate generality, legitimacy in human-facing settings, or deployment readiness.
From an interaction-design perspective (
Section 4), the ablation study underscores the importance of the contestability interface. In the present experiments, review is implemented as an automated probabilistic proxy for due process. It is useful for mechanism analysis, but it is not a validated human contestation interface. In real deployments, review would involve stakeholders interacting with the ledger through a multimodal interface, examining evidence, submitting appeals, and receiving outcomes.
The quality of that interaction affects the effective review rate (
) and, consequently, the magnitude of the institutional tax. Poorly designed review interfaces that are cognitively burdensome would lower the effective review rate (closer to the Sanction-No-Review ablation) and increase false-positive harm. By contrast, clearer interfaces and more efficient appeal processes (closer to the High Review ablation) would mitigate these costs while preserving the behavioral benefits of the institution.
Appendix C makes this dependence more explicit by varying
and
while holding the learned policy fixed.
A capacity-constrained review pipeline would generally worsen this picture. If appeal volume exceeds reviewer throughput, the effective review process is not only lower
but also larger settlement delay
D, and
Section 6.3 implies that both changes push the system toward higher residual sanction burden. The present simulations vary
while keeping
D fixed, so they should be interpreted as a lower-dimensional proxy for review bandwidth rather than as a full model of queueing or institutional gridlock.
Given
training seeds per condition, the inferential results in
Table 8 should be read as exploratory. We therefore rely primarily on directional consistency across seeds, uncertainty intervals, effect sizes, and ablation structure rather than on binary significance claims.
9.6. Per-Agent Fairness and False-Positive Burden
A central concern for accountability institutions is how enforcement burdens are distributed across agents. Aggregate welfare can mask whether some agents bear disproportionate false-positive costs, even when mean outcomes appear acceptable. To address this, we analyze IML ledger records at the agent level, tracking flagged enforcement events, true violations, false positives, and overturned cases across runs.
Table 9 reports the per-agent enforcement profile in both environments. In these symmetric benchmark settings, the measured burden is nearly uniform across agents. The burden Gini coefficient is
in
Cleanup and
in
Harvest, and the per-agent flagged-event counts cluster tightly around
events per episode.
Figure 8 visualizes the per-agent enforcement burden. The narrow spread across agents is consistent with the structural symmetry of the monitored rule and monitoring process in these environments. This result is useful, but it should be interpreted narrowly: it does not imply that institutional fairness is guaranteed under heterogeneous roles, asymmetric observability, or differential monitoring in real deployments.
The ledger also reveals an important temporal shift.
Figure 9 shows that true-positive flagged events are concentrated early in training, when agents still violate the monitored rule. As policies adapt, true violations fall sharply, but flagged events do not disappear because residual enforcement is increasingly driven by monitoring error.
Figure 10 makes this transition explicit. In both environments, the phase-level false-positive rate rises from about 0.88 in the early phase to above 0.997 in later phases. As agents learn to comply, the fraction of flagged events that are false positives therefore rises toward 100%. This is the same long-horizon “institutional tax” predicted by the incentive analysis: once real violations become rare, even a small per-step false-positive rate can dominate the remaining enforcement burden.
Together, these results refine the paper’s fairness claim. In the present symmetric benchmarks, IML does not concentrate measured enforcement burden on particular agents. The more durable lesson, however, is temporal rather than distributive: mature compliance regimes make review and appeal more important, not less so, because the residual burden increasingly reflects measurement error rather than genuine violations.
10. Illustrative Vignettes for Human–AI Integration
To make the institutional questions concrete, we outline four illustrative vignettes. They are not empirical case studies. Instead, they show where design choices concentrate ethical risk, what kinds of evidence and recourse are feasible, and how the same technical primitives can support either accountable governance or coercive control depending on how they are scoped and governed.
10.1. Platform Community Governance: Moderation and Harassment
Consider a large platform that uses AI assistants to support harassment mitigation, misinformation moderation, and community rule enforcement. The socio-technical system is inherently multi-agent: users, creators, recommendation/ranking components, automated classifiers, reporting tools, and human moderators interact repeatedly, often under asymmetric information. Norms are expressed through community standards (e.g., harassment, coordinated abuse, doxxing), but these standards are contested, context-dependent, and frequently require interpretation rather than mechanical application.
In this setting, an IML-style institution is attractive because it can make enforcement
inspectable rather than merely effective. The design challenge is to define monitoring scope without collapsing into pervasive surveillance. Monitoring must specify what signals are collected (content, metadata, interaction graphs) and justify these flows under contextual integrity constraints [
13]. Evidence pipelines must avoid opacity failures in which users cannot understand what evidence was used or why a classification was treated as actionable [
24]. Contestability must be treated as a core mechanism property rather than a customer-support add-on: users need a practicable path to appeal takedowns or sanctions, and explanations must be intelligible and matched to the kinds of mistakes the system makes [
22]. Finally, the institution must explicitly consider power and chilling effects. Even when enforcement improves aggregate safety, monitoring that is experienced as pervasive or unpredictable can deter participation and reshape expressive behavior [
15,
16].
A concrete instantiation can be deliberately minimalist. Instead of “collect everything” logging, the ledger can record (i) the specific norm invoked (policy clause and version), (ii) a pointer to the content/action at issue, (iii) the evidence features actually relied upon (e.g., classifier scores, report provenance, limited contextual signals), and (iv) the settlement outcome (warning, takedown, time-limited restriction) with graduated escalation [
1]. Where content retention is high-risk, the ledger can store commitments/hashes and access-controlled references rather than full payloads. Crucially, the contestation interface should expose the invoked norm and provide a path to human review, because legitimacy depends on being heard and having recourse, not only on outcome accuracy [
10,
22]. The same infrastructure can become illegitimate if monitoring scope expands without explicit boundary-setting and oversight; in this vignette, the main value of IML is precisely that it makes boundaries and procedures governable.
10.2. Enterprise Agentic Workflows: Policy Compliance and Data Access
Organizations increasingly deploy agentic assistants that plan tasks, call tools, query internal databases, and coordinate work across teams. Here, “cooperation” often means compliance with access controls, adherence to data-minimization policies, truthful reporting of uncertainty, and non-circumvention of safety constraints. The relevant multi-agent ecology includes employees, teams, contractors, internal systems, AI components, and security/compliance functions.
IML naturally fits this environment as an accountability layer for norm-relevant events such as access requests, data exports, and tool calls, with settlement delayed to support review. The ethical risk, however, is acute: workplace logging can drift into surveillance, and auditability can be repurposed for performance discipline rather than policy compliance [
14,
15]. Accordingly, governance must specify who can access logs, what is retained, and how individuals can contest interpretations of logged events, especially where false positives can trigger reputational or employment harm.
A defensible instantiation treats IML as a policy compliance institution rather than a productivity-monitoring apparatus. This implies role-based access to logs, strict retention limits, and aggregation where feasible (e.g., counts of policy-relevant events rather than raw content), along with a clear separation between security monitoring and HR/performance evaluation. Contestability should be operational: employees should be able to challenge whether a logged event was truly a violation (including disputing false positives) and whether the applied settlement was proportionate, with documented resolution timelines. In such settings, the ledger’s primary governance value is to enable audit and incident reconstruction without normalizing pervasive workplace surveillance.
10.3. Autonomous Fleets: Shared Safety Norms
Robot fleets in warehouses, delivery, or medical logistics share resources and coordinate under safety and efficiency constraints. Cooperation in this vignette takes the form of safe spacing, deconfliction, right-of-way adherence, and compliance with priority rules, with repeated interactions and potentially high-stakes externalities.
Here, an IML-style institution can act as a shared “incident ledger” that supports both deterrence and post hoc investigation. The ledger can record norm-relevant events (near-misses, rule violations, overrides), while settlement can implement graded responses such as capability throttling, mandatory safety re-calibration, or temporary access restrictions after repeated unsafe behavior. This is aligned with proposals for accountable recording in autonomous systems (ethical “black boxes”) [
45]. Because these deployments are safety-critical and organizationally embedded, meaningful human control requires that humans can override institutional actions and audit why they occurred, and that responsibility attribution is supported by usable records [
54,
55].
A concrete instantiation should prioritize investigability without maximizing surveillance. The ledger can store time-bounded sensor snapshots (subject to privacy limits), policy and firmware versions, and a trace of actions leading to near-misses, enabling reproducible incident reconstruction. Settlement is often better framed as capability and safety management (e.g., throttling after repeated violations) rather than purely punitive transfers, reducing incentives to hide incidents and aligning with safety governance norms. As in other vignettes, monitor error rates and settlement rules should be audited periodically: wrongful “incidents” can trigger costly downtime and erode trust in both the robots and the institution.
10.4. Public-Sector Services: Eligibility, Compliance, and the Risk of “Automated Punishment”
Public-sector agencies increasingly rely on data-driven systems to allocate scarce resources, flag fraud, and prioritize service provision. The socio-technical system is multi-agent: citizens, caseworkers, contractors, and decision-support models interact over time. These settings are institutionally hard because power asymmetries are large, stakes are high, and procedural legitimacy is not optional. Empirical accounts caution that automated systems can produce punitive error cascades when monitoring and enforcement are opaque, unaccountable, or disconnected from meaningful recourse [
17,
18].
If IML were deployed at all in this domain, it should be used to
raise the procedural floor rather than accelerate sanctions. Evidence standards must be explicit (what counts as actionable evidence, how uncertainty is represented, and what human review is required), and due process must be designed for real access, including offline appeal pathways for people with limited digital connectivity or literacy. Monitoring scope must be justified against contextual expectations and the risk of surveillance harms [
15]. A ledger can help conditionally by forcing documentation and contestability: each flag is tied to a norm, an evidence trail, a settlement rule, and a review outcome, and the institution can be audited for error rates and disparate burdens. Yet the same infrastructure can be used to intensify monitoring or routinize punitive automation; this vignette therefore underscores a boundary condition for IML: in high power-asymmetry settings, the legitimacy constraints in
Section 6.2 and the governance lifecycle in
Section 11 are baseline requirements rather than optional best practices.
11. Governance Integration: From Principles to Practice
High-level ethics principles are widespread yet repeatedly fail to constrain practice when they are not institutionalized into roles, procedures, and inspectable artifacts [
26,
49]. IML is designed to support such institutionalization by making enforcement
procedural and
auditable: norms are explicit, monitoring scope is governable, settlements are logged, and review outcomes are recorded. This section describes how an IML deployment can be integrated into governance processes in a way that is operational rather than aspirational.
11.1. Lifecycle Governance
Institutional mechanisms require lifecycle governance because norms, evidence pipelines, and organizational practices evolve. We therefore propose a governance lifecycle with five linked phases (
Figure 11). The lifecycle begins with
participatory norm-setting, in which the norm predicate
v and settlement goals are defined with relevant stakeholders rather than silently encoded. It continues with a
pre-deployment impact assessment that documents anticipated harms, monitoring scope, contestability pathways, and mitigation plans in the style of algorithmic impact assessment practice [
46,
47]. Deployment should then occur
with audit hooks by design: evidence and settlements are logged, error metrics are measured and tracked, and access/retention rules are enforced as governance constraints rather than engineering defaults. Once deployed, IML must support
ongoing redress through functioning appeals and dispute resolution, including aggregate reporting on outcomes and error burdens. Finally, the institution must be
audited and updated through internal and/or external review, including analysis of appeals and revision of norms and settlement policies when evidence or stakeholder feedback indicates drift, abuse potential, or unjust burdens [
12,
48]. The key governance claim is that institutions should be revisable: contestability is not merely about individual cases but about the ability to change institutional rules when they fail.
11.2. Alignment with Standards and Regulation
IML can support compliance-oriented risk management by producing inspectable artifacts that many frameworks increasingly expect: explicit monitoring policies, settlement rules, auditable logs, and documented error rates [
50,
51,
52,
53]. In practice, these artifacts can help organizations meet requirements for transparency, accountability, and post-incident reconstruction. However, alignment should not be treated as a check-box exercise: an institution can satisfy formal documentation requirements while still being illegitimate for affected communities. Accordingly, standards alignment should be treated as necessary but not sufficient; legitimacy depends on whether affected parties experience contestability, proportionality, and contextual integrity in practice.
11.3. From Framework Requirements to Concrete Artifacts
To avoid “principles without practice” [
26], IML deployments should produce artifacts that can be inspected by reviewers, auditors, and affected communities.
Table 10 illustrates one natural mapping from IML outputs to common risk-management functions, using the NIST AI RMF categories as a convenient organizing frame [
50]. The intent is not superficial compliance, but operational hooks that allow institutions to be scrutinized, challenged, and revised when they fail.
A practical implication is that IML should ship with a minimum viable governance package—a baseline set of commitments that make the institution auditable and contestable even under constrained organizational capacity. At minimum, this package should document the enforced norms and their justification (including stakeholder participation and dissent handling); the monitoring scope, data sources, retention limits, and access controls; the settlement logic, sanction bounds, and the points where humans are in the loop; measured error rates and a plan for continuous monitoring of those errors; a functioning appeal and redress mechanism accompanied by aggregate statistics on reversals and resolution times; and an update process that prevents silent changes to institutional rules and enables rollback when harm is detected. These artifacts are not merely administrative: they are the concrete mechanisms by which institutions become governable rather than opaque.
By requiring explicit artifacts and revision pathways, IML also directly targets failure patterns documented in critiques of opaque algorithmic governance, where automated decisions are difficult to contest, error burdens cascade, and institutional accountability is diffuse [
17,
18].
12. Limitations and Open Problems
Several limitations of the current work should be acknowledged.
While we articulate interaction design principles for contestable AI institutions (
Section 4), we do not evaluate these principles with human participants. A user study examining how stakeholders interact with the auditable ledger, how effectively they can contest false positives, and how the interface design affects perceived procedural justice would materially strengthen the empirical grounding of our interaction design claims. This is a priority for future work.
Our empirical evaluation now includes direct quantitative comparisons against inequity aversion and social influence, which helps position IML relative to representative internalization approaches. The baseline set remains selective, however, and does not yet include contract-based approaches, additional norm-aware methods, or environment-specific non-learning baselines.
The evaluation uses two environments with a single narrow rule. This choice was intentional because it isolates monitoring error and due-process effects, but it leaves unexamined how IML behaves under more contested norms, multiple simultaneous rules, and richer interaction dynamics.
The review and appeal channel is simulated probabilistically rather than implemented as an actual human-in-the-loop interaction. The effective review rate in real deployments would depend on the quality of the contestability interface, the cognitive load on reviewers, and the volume of appeals. Capacity bottlenecks could also create queues and longer settlement delays, which our simulations do not model directly.
The simulations should be read as a focused pilot study. Five training seeds per condition are sufficient to surface large directional effects and failure modes, but they are still limited for strong claims about stability, tail behavior, or corrected statistical significance. Future work should expand the seed count, broaden substrates, and report pre-specified reliability metrics such as collapse frequency, variance reduction, and outlier sensitivity.
We now report per-agent sanction and false-positive burden, and in these symmetric environments, the burden is nearly uniform across agents. That result should not be overgeneralized: future work should test whether sanctions, reversals, or review delays concentrate on particular agents, roles, states, or demographic groups when monitoring is heterogeneous or strategically manipulated.
This paper is deliberately scoped as a framework contribution with pilot experiments: it argues for institutional cooperation mechanisms as a governance-relevant alternative to covert influence, formalizes a minimal IML wrapper, and demonstrates empirically how monitoring error can dominate welfare in long-horizon settings. Several limitations and open problems follow directly from this scope.
First, our incentive analysis is intentionally conservative. The one-shot deviation deterrence bound is designed as an auditable check rather than a full characterization of strategic behavior. Equilibrium analysis in partially observable Markov games with heterogeneous learning dynamics, delayed settlement, and endogenous monitoring remains technically challenging, and our bounds do not capture the richness of multi-step deviations, coordinated deviations, or equilibrium selection effects.
Second, norm specification is not a purely technical problem. Norms are value-laden and contested; even when an institution is transparent, disagreement can persist about what should count as a violation, what evidence is sufficient, and what sanctions are legitimate. Participatory and value-sensitive approaches are therefore necessary, but operationalizing such participation at scale—especially across heterogeneous communities and cultures—remains difficult and politically fraught [
58].
Third, institutional design does not eliminate power. Transparent institutions can still be captured by powerful actors, and procedural features can be implemented in ways that are formally present but substantively ineffective. Meaningful checks and balances, credible oversight, and enforceable constraints on operators are therefore essential for any real deployment, yet such governance capacity varies widely across organizations and jurisdictions.
Fourth, monitoring changes behavior beyond deterrence. Monitoring can produce chilling effects that reduce exploration, creativity, and willingness to engage, which is especially salient for human–AI collaboration and expressive domains [
15]. Even when monitoring improves aggregate welfare, these dignitary and autonomy costs can undermine legitimacy, and they are not easily captured by standard MARL benchmarks.
Finally, legitimacy measurement itself is a moving target. Survey-based legitimacy metrics can be gamed, influenced by framing, or divorced from longer-term institutional experience. For this reason, legitimacy evaluation likely requires mixed methods, longitudinal observation, and attention to how institutions interact with organizational incentives and social power rather than relying on short-run proxies.
12.1. Research Agenda: Interaction Design for Accountable AI Institutions
The interaction design challenges identified in this paper suggest several concrete research directions:
Multimodal evidence presentation. How should complex multi-agent trajectories be visualized to support human review of institutional decisions? What combination of textual logs, spatial visualizations, and temporal replays is most effective for different stakeholder roles?
Contestability interface evaluation. What interface designs minimize the cognitive burden of contesting false positives while maintaining the integrity of the appeal process? How does the friction of the appeal interface affect the effective review rate and, consequently, system-level welfare?
Adaptive monitoring interfaces. Can the monitoring interface be designed to adapt its sensitivity and reporting based on the observed false-positive rate and the capacity of the review system, creating a feedback loop between interaction design and institutional performance?
Cross-cultural interaction norms. How do cultural differences in attitudes toward authority, contestation, and procedural justice affect the design requirements for accountability interfaces in global AI deployments?
12.2. Research Agenda: Institutions for Agentic Ecosystems
Despite these limitations, IML suggests a concrete interdisciplinary research program for agentic ecosystems in which humans and AI components interact repeatedly under partial observability and power asymmetries. The first priority concerns legitimacy under partial automation: which combinations of explanation, appeal, and outcome control produce stable perceptions of legitimacy when enforcement is partly automated, and how do these procedural features interact with outcome quality over time [
10,
22]? A second priority concerns institutional pluralism and conflict. Digital systems increasingly operate under overlapping communities, jurisdictions, and rule systems; understanding how institutions behave when norms conflict—and how “nested enterprises” should be implemented in digital form—is central for practical deployment [
1].
A third priority is privacy-preserving monitoring. Institutional enforcement requires evidence, yet the evidence required for accountability must be balanced against contextual integrity and the risk of surveillance harms [
13,
15]. This motivates research on what minimal evidence suffices for legitimate enforcement, how to protect evidence while enabling audit, and how to design monitoring scopes that are technically enforceable rather than aspirational. A fourth priority concerns learning dynamics and Goodhart effects: learning agents will adapt to institutional incentives over long horizons, so institutions must detect and correct gaming, proxy exploitation, and feedback-loop pathologies, ideally with auditable diagnostics rather than opaque patching.
Finally, power and capture must be treated as design constraints, not externalities. Future work should examine what checks and balances reduce institutional capture by powerful actors, what transparency is necessary for meaningful external oversight, and how auditing and redress can be made credible when incentives to conceal error are strong [
12]. Addressing these questions is essential for connecting optimization-centric accounts of cooperation with institution-centered governance in socio-technical systems.
A primary lesson of the pilot study is that measurement error is a primary issue for institutional cooperation mechanisms. When monitoring is applied at the granularity of agent-steps and decisions are delayed, even small false-positive rates can accumulate into predictable welfare losses unless they are explicitly budgeted, procedurally mitigated, and audited. For this reason, IML should be read not as a finished governance solution but as a research framework with pilot empirical evidence: a way to make cooperation interventions legible as institutions with tunable error budgets, contestability pathways, and accountability artifacts, enabling systematic study of the welfare–legitimacy trade-offs that matter in practice.
13. Conclusions
As AI systems become more deeply integrated into socio-technical settings, the design of accountability mechanisms becomes as important as the underlying learning algorithms. This paper makes a narrower claim than a full empirical validation study. We conceptualize cooperation control as an institution-design problem, formalize IML as a transparent wrapper for Markov games, derive conservative incentive checks, and illustrate the mechanism in two canonical sequential social dilemmas.
The pilot study does not establish legitimacy in human-facing settings or deployment readiness. It does show that, in these two SSD environments, an external accountability layer can change learning dynamics, expose monitoring error through auditable records, and make review design part of system performance. In these runs, the representative internalization baselines were harder to stabilize, while IML incurred a persistent welfare cost when false positives accumulated. The added agent-level analysis shows that measured burden is nearly uniform across agents in the present symmetric benchmarks, but increasingly dominated by residual false positives as compliance improves.
The paper therefore argues for treating cooperation mechanisms as institution-design problems as well as reward-design problems. The next step is broader empirical and human-facing validation: more environments, richer norm classes, human-subject evaluation, explicit capacity-constrained review models, and fairness tests under heterogeneous monitoring.