1. Introduction
Mitigating free-riding in the provision of public goods fundamentally relies on addressing coordination failures and incentivizing voluntary contributions. A growing body of the literature highlights “leading by example” as an important mechanism to promote cooperation (
Glowacki & von Rueden, 2015;
Pietraszewski, 2020). In the present study, we use the term leader in a specific experimental sense: a leader is the participant who makes the first contribution decision in a sequential public goods game, and this decision is observed by the other group members before they make their own contribution decisions. By contrast, in leaderless groups, all group members make their contribution decisions simultaneously, without observing any other member’s contribution beforehand. Leaders can influence group outcomes by signaling commitment, setting a clear cooperative benchmark, and coordinating behaviors (
Antonakis et al., 2022;
Gächter et al., 2012;
Gächter & Renner, 2018;
Moxnes & Van der Heijden, 2003;
Zhou et al., 2015). However, existing evidence suggests that leadership effectiveness is highly dependent on the institutional design of the selection process. Accordingly, our analysis focuses on evidence directly relevant to sequential public goods games and leader-selection institutions.
A key insight from the literature is that leading by example may operate through informational and reciprocal channels in sequential interactions (
Arbak & Villeval, 2013;
Hermalin, 1998;
Potters et al., 2007;
Stanca et al., 2009). The effectiveness of these channels depends on the leader-selection process and the institutional consequences of accepting or rejecting leadership. These factors influence both the informativeness of leadership behaviors and followers’ willingness to reciprocate.
Two challenges have emerged from the literature regarding this process. First, leadership mechanisms with low entry barriers do not always succeed in generating a leader (
Arbak & Villeval, 2013;
Haigner & Wakolbinger, 2010;
Rivas & Sutter, 2011). When a selection mechanism fails to yield a leader, the resulting absence of leadership sends a strong negative signal, leading to pessimism among followers and lower cooperation (
Xu et al., 2025). Second, when leadership is distributed among multiple individuals, coordination difficulties and diffused accountability may diminish its effectiveness.
We conceptualize leadership as the influence exerted by the first mover in a sequential public goods game and investigate how alternative leader-selection mechanisms address these challenges and affect voluntary contributions compared to a leaderless baseline, focusing on two main objectives.
First, we provide a systematic comparison of three institutional processes for leader selection: exogenous leadership, voluntary leadership, and delegated leadership following refusal. We implement four treatments. Treatment 1 (T1, control) employs a leaderless voluntary contribution mechanism (VCM). Treatment 2 (T2, Random Leadership; RL) randomly appoints one group member as the leader. To isolate the effect of individual consent from exogenous assignment, Treatment 3 (T3, Voluntary Leadership from a Random Candidate; VL-RC) allows a randomly selected candidate to accept or reject the leadership role; rejection results in no leader. Additionally, we introduce a novel mechanism, Treatment 4 (T4, Delegated Leadership by a Random Candidate; DL-RC), which also begins with a randomly selected candidate. If the candidate accepts, the game proceeds with a single leader. If the candidate refuses, the remaining three group members move first jointly, while the original candidate becomes the sole follower. This allows us to examine whether refusal yields different behavioral outcomes depending on whether it leads to no leadership or to a collective leadership structure.
Second, we compare fixed matching with random rematching to assess whether the effectiveness of each mechanism differs when stable-group membership and repeated-partner reputational incentives are present versus when such incentives are weakened by stranger matching.
Our experimental design employs a unified
framework that bridges and extends two strands of the existing literature. Prior studies have investigated standard leadership mechanisms under either random matching (
Xu et al., 2025), which weakens stable repeated-partner reputational incentives, or fixed matching (
Haigner & Wakolbinger, 2010;
He & Zheng, 2024), which allows repeated interactions; however, they have not simultaneously varied leader-selection rules and matching protocols within a single experimental design.
Existing studies have investigated a wide range of sequential-move structures, including one-leader games and alternative move orderings such as fully sequential and multi-leader games (
Eichenseer, 2023;
Haigner & Wakolbinger, 2010;
Helland et al., 2018). However, the behavioral implications of delegating first-mover responsibility following a candidate’s refusal—which explicitly transfers the leadership burden to the remaining group members—have not been studied. Our DL-RC mechanism enables us to evaluate the overall performance of this refusal-triggered delegation rule under both fixed and random matching. Because the refusal event and the three-first-mover structure are jointly generated by the mechanism, the design does not separately identify the independent effect of leader multiplicity. Therefore, we examine whether the complete DL-RC institutional arrangement affects cooperation and use realized-state comparisons to determine where its aggregate performance is concentrated. We hypothesize that reputational incentives under fixed matching encourage leadership and contributions. In contrast, random matching is likely to exacerbate free-riding behaviors when leadership is either declined or shared among multiple members.
Using this framework, we identify five main findings while distinguishing descriptive rankings from confirmatory inference. First, leadership does not uniformly increase contributions. DL-RC exhibits the lowest observed mean contribution overall, particularly under random rematching; however, the primary omnibus and Holm-adjusted model-based pairwise tests do not establish a general mechanism effect. Second, within VL-RC and DL-RC, observed contributions differ substantially across realized acceptance and refusal branches. These comparisons are conditional on endogenous candidate decisions and do not identify causal effects of leadership acceptance, refusal, leader absence, or leader multiplicity. Third, first movers generally contribute more than later movers in leader-present groups, and later-mover contributions are positively associated with observed first-mover contributions. Fourth, the endogenous mechanisms produce different realized group structures; the lower observed contribution in DL-RC is descriptively concentrated in the refusal-contingent three-first-mover branch. Finally, fixed matching has a higher observed mean contribution than random rematching, but the pooled difference is not statistically significant; the clearest within-mechanism matching contrast occurs in DL-RC.
The rest of this paper is organized as follows.
Section 2 reviews the literature.
Section 3 describes the experimental design,
Section 4 presents the results, and
Section 5 concludes.
2. Literature Review
The economic analysis of leadership as an institutional mechanism for addressing collective-action problems originates with
Hermalin (
1998), who conceptualizes “leading by example” as a sequential-move process in which a first mover coordinates beliefs and behaviors. In studies of experimental public goods games, leadership typically operates through two interrelated channels. The first is an informational channel: the leader’s contribution provides a benchmark that can shift followers’ expectations and serve as a credible signal of intent or commitment (
Antonakis et al., 2022;
Gächter et al., 2012;
Pietraszewski, 2020;
Potters et al., 2007). The second is a reciprocity channel: followers respond conditionally to the perceived cooperativeness of the leader, rewarding positive examples and punishing negative ones through their own contribution decisions (
Arbak & Villeval, 2013;
Stanca et al., 2009). The interaction between these two channels determines whether leadership functions as a catalyst for cooperation or as a trigger for collective free-riding.
The existing literature has increasingly explored the effects of selection mechanisms on leadership outcomes. For exogenous leadership—where the first mover is assigned randomly—the empirical evidence is mixed. Some studies report significant improvements in efficiency (
Dannenberg, 2015;
Levati et al., 2007;
Moxnes & Van der Heijden, 2003), while others report minimal or statistically insignificant effects on cooperation (
Gächter & Renner, 2018;
Haigner & Wakolbinger, 2010). In contrast, endogenous leadership operates as a costly signal, reflecting a leader’s willingness to bear the burden of leadership and to forgo the advantage associated with moving second (
Eichenseer, 2023). Consequently, some studies find that voluntary leadership promotes higher levels of cooperation than externally imposed leadership (
Dannenberg, 2015;
Haigner & Wakolbinger, 2010). This variability suggests that the informational and reciprocal significance of leadership signals depends on the underlying selection mechanisms. The importance of institutional credibility is further emphasized in applied contexts such as climate cooperation.
Helland et al. (
2018) demonstrate that the effectiveness of conditional commitments is fundamentally linked to their credibility and to the perceived fairness of the institutional arrangement.
A second institutional dimension concerns the allocation of responsibility and influence in contexts of shared leadership. Empirical behavioral evidence indicates that distributing decision-making responsibility may dilute accountability and foster increased self-interested behavior compared to situations involving a single, clearly accountable decision maker (
Charness & Jackson, 2009). In public goods games, collective leadership may similarly inhibit prosocial behavior, as group leaders tend to exhibit lower levels of cooperation than individual role models (
Zhou et al., 2015). Supporting this perspective,
Xu et al. (
2025) demonstrate that unrestricted voluntary leadership (VL)—which often results in multiple leaders—can yield suboptimal outcomes due to coordination difficulties and diminished accountability.
We extend this body of the literature by examining a unique case in which shared leadership is not voluntarily adopted but is institutionally mandated following an act of refusal. Unlike settings where multiple leaders emerge voluntarily (
Xu et al., 2025), the DL-RC framework investigates a specific institutional response to refusal, wherein the remaining group members are assigned collective first-mover responsibility. This design allows us to evaluate whether the refusal-triggered delegated arrangement is associated with different contribution outcomes. However, because refusal and the multiple first-mover structure occur simultaneously, the design does not isolate responsibility diffusion from alternative explanations such as negative inferences from refusal, coordination difficulties, or concerns about the legitimacy of the delegated roles. Therefore, we interpret the realized DL-RC branches as conditional outcomes generated by the complete institutional package rather than as direct evidence of any single psychological mechanism.
Furthermore, this study contributes to the literature by examining how leader-selection rules interact with matching protocols. Previous studies have typically evaluated particular leadership mechanisms within a single matching environment. For example, fixed matching allows stable-group membership and thereby creates scope for repeated-interaction, reciprocal, and reputational incentives (
He & Zheng, 2024), whereas random rematching weakens stable-group-specific relationships while preserving the observability of leadership behavior within each period (
Xu et al., 2025). By implementing the same set of leader-selection mechanisms under both protocols, our design provides a unified framework for assessing whether their behavioral performance varies with the availability of stable repeated-partner incentives. This comparison motivates our hypotheses concerning the relative performance of RL, VL-RC, and DL-RC under fixed matching and random rematching.
3. Experimental Design
3.1. Basic Setting
The experimental design is based on the voluntary contribution mechanism (hereafter, VCM). Throughout the experiment, we define a leader operationally as a first mover in the contribution sequence. In leader-present groups, the leader, or leaders in the case of delegated leadership, choose their contributions before the later mover(s), and these first-mover contributions are revealed to the later mover(s) before they make their own contribution decisions. In leaderless groups, all four participants make contribution decisions simultaneously, and no participant observes another group member’s contribution before deciding. In each session, participants are assigned to groups of four and engage in interactions over 10 periods. Depending on the treatment, group composition either remains constant throughout the experiment (fixed matching) or is randomly re-matched in each period (random matching).
In each period, subjects are endowed with 20 tokens. Although the sequence of moves varies across treatments, the payoff structure is identical in all treatments. Participants decide how to distribute their endowment between a private account and a public account. The Marginal Per Capita Return (MPCR) from the public account is 0.4, meaning every token contributed to the public pool yields a total of 1.6 tokens shared equally by the group. Let
denote the individual
i’s contribution to the public account in period
t, restricted to an integer that satisfies
. Let
denote individual
i’s earnings from their private account and the public account in period
t, given the contributions of all group members
:
3.2. Matching Protocols and Leader-Selection Mechanisms
The experiment employs a factorial design, interacting two matching protocols with four distinct leader-selection mechanisms, yielding a total of eight treatments.
Matching Protocols. To investigate the interaction between institutional mechanisms and repeated-interaction incentives, we implement two matching rules.
1 - 1.
Fixed Matching: Group membership is fixed for all periods. This protocol creates stable-group-level repeated interaction and may allow for reciprocal or reputational incentives within the group.
- 2.
Random Rematching: Groups are randomly re-formed in each period. We use this protocol to weaken stable-group-specific reputational incentives and repeated interaction within the same group. Because the same participants remain in the experiment for 10 periods and may learn from previous outcomes, this protocol should not be interpreted as a sequence of fully independent one-shot games. Rather, it approximates stranger matching by preventing stable-group membership across periods.
Leader-Selection Mechanisms. Within each matching protocol, we implement the following four distinct mechanisms. These mechanisms differ in their selection processes (exogenous vs. endogenous) and in the default outcome when the leadership role is not accepted.
- 1.
Control (Baseline): Subjects participate in a standard VCM. All group members make their contribution decisions independently and simultaneously.
- 2.
RL (Random Leadership): One randomly selected group member is exogenously appointed as the leader by the computer. The game proceeds sequentially: the leader first decides on their contribution, which is then revealed to the three followers. After observing the leader’s decision, followers decide simultaneously on their respective contributions.
- 3.
VL-RC (Voluntary Leadership by the Randomly Selected Candidate): One group member is randomly selected as a candidate and offered the leadership role.
If the candidate accepts, they become the sole voluntary leader and contribute first (identical to the RL treatment).
If the candidate declines, the leadership position is dissolved. The group reverts to the standard VCM (control), where all members contribute simultaneously.
- 4.
DL-RC (Delegated Leadership by the Randomly Selected Candidate): One group member is randomly selected as a candidate and offered the leadership role.
If the candidate accepts, the procedure follows the same one-leader sequence as in the accepted branch of VL-RC: the candidate becomes the sole voluntary leader and contributes first.
If the candidate declines, leadership responsibility is delegated to the other three group members. In this sub-game, the three delegated leaders contribute first, while the original candidate becomes the sole follower and decides only after observing the delegated leaders’ contributions.
In all treatments, subjects receive anonymous feedback on group members’ roles, contributions, and payoffs at the end of each period.
The existing literature has comprehensively examined the standard leadership mechanisms (control, RL, VL-RC) under distinct matching protocols. While
Xu et al. (
2025) investigated these mechanisms under a random-rematching protocol that weakens stable repeated-partner incentives,
He and Zheng (
2024) provided a counterpart analysis under a fixed-matching protocol that allows reputation building within stable groups.
Despite these contributions, the strategic implications of delegation remain unexplored—particularly the scenario where rejecting leadership explicitly shifts the coordination burden onto other group members. To fill this gap, we introduce a novel DL-RC (delegated leadership) mechanism and implement it across both matching protocols to establish a complete 2 × 4 comparative framework. The design enables us to identify the overall effect of assignment to the DL-RC institutional condition relative to the control, RL, and VL-RC and to examine how its performance varies across different matching environments.
3.3. Hypotheses
Following the game-theoretic framework of
Hermalin (
1998), we posit that leaders facilitate public goods provision through signaling and reciprocity. We hypothesize that the varying effectiveness of our mechanisms originates from differences in the selection process. Specifically, the selection rule dictates whether the act of leading serves as an informative signal of the leader’s cooperative intent.
We view the selection mechanism as shaping how informative leadership behavior is to followers. Upon observing the leader’s contribution, followers update their beliefs regarding the leader’s type (altruistic vs. self-interested) and the group’s cooperative potential. However, the clarity of this signal depends critically on the counterfactual consequences of declining the leadership role.
The experimental design examines the effects of alternative leader-selection and matching institutions on observable role choices and contribution behavior. Although we do not directly measure psychological mediators such as beliefs, intentions, or perceived accountability, mechanisms such as signaling, responsibility allocation, and reputation provide theoretically grounded channels that inform our testable predictions. Consequently, our empirical tests focus on differences in leadership acceptance, conditional contribution behavior, and group outcomes across institutional conditions.
Table 1 summarizes the structural differences and strategic properties of the three active leadership mechanisms:
We start by discussing the key design features of the treatments with respect to the number of leaders. DL-RC is the only mechanism that generates multiple first movers (Fact F1). While RL is an exogenous leader-selection mechanism that guarantees a single leader in every period, VL-RC and DL-RC are endogenous, as the randomly selected candidate may either accept or decline the leadership role. In VL-RC, a rejection results in a leaderless group (Fact F2). Conversely, DL-RC precludes a leaderless outcome by shifting the first-mover responsibility to the remaining group members upon the candidate’s rejection, thereby potentially creating multiple leaders (Fact F1).
Table 1 further highlights whether a mechanism respects individual willingness not to lead (Attribute A1). VL-RC and DL-RC incorporate this feature by allowing the selected candidate to decline the role, whereas RL imposes leadership exogenously, regardless of the subject’s underlying motivation.
Candidates who voluntarily accept roles in VL-RC and DL-RC effectively select themselves into the first-mover position. This observable selection is expected to make their subsequent contributions more informative to followers than those of randomly appointed leaders, thereby potentially enhancing the value of their actions as signals of cooperative intent. In contrast, leaders in RL are randomly assigned and may lack intrinsic motivation to promote group cooperation.
A distinctive feature of DL-RC is that rejection by the candidate induces multiple delegated leaders. In this multiple-leader sub-game, each delegated leader may expect the others to shoulder the responsibility for public good provision, giving rise to a free-riding incentive among delegated leaders (Effect E1).
Finally, VL-RC is the only mechanism in which the absence of a leader conveys information. However, this signal is inherently noisy (Effect E3): a leaderless outcome arises solely because the single randomly selected candidate refused, which does not preclude the possibility that other group members would have been willing to lead. Consequently, the no-leader state in VL-RC provides an imprecise signal of low cooperative intent.
The hypothesized behavioral implications summarized in
Table 1 may also depend on the matching protocol. Under fixed matching, stable-group membership creates scope for repeated-interaction, reciprocal, and reputational incentives. These incentives may strengthen the behavioral significance of voluntary acceptance (E2) and may limit the adverse consequences associated with shared first-mover responsibility in the refusal-triggered branch of DL-RC (E1). Under random rematching, stable-group-specific incentives are weakened, although leadership choices and contributions remain observable within each period. In this environment, voluntary leadership signals may carry less influence across periods, while the refusal-triggered multiple-first-mover arrangement may be more vulnerable to coordination difficulties or reduced individual accountability.
Based on these considerations, we hypothesize the following efficiency ranking. VL-RC combines positive selection (E2) with the absence of multiple-leader free-riding (E1), making it a strong candidate for the most effective mechanism. RL suffers from the lack of leader motivation (weak E2). DL-RC represents an intermediate case: while it screens for willingness in the acceptance branch (E2), it entails the structural risk of multiple forced leaders (E1) following rejection. We expect this disadvantage to be mitigated under fixed matching, where reputational incentives discourage rejection and thereby reduce the inefficient outcome.
3.4. Procedure
A total of 256 college students from Tsinghua University participated in the experiment between April and July 2020. The experiment was conducted as a controlled online laboratory experiment during the COVID-19 pandemic. Participants were recruited from the university subject pool, signed an informed consent form before participation, and were allowed to participate only once. This study followed standard ethical protocols and was approved by the Institutional Review Board (IRB) prior to data collection.
We implemented a between-subjects design with eight experimental conditions, corresponding to four leader-selection mechanisms crossed with two matching protocols. Each experimental condition included 32 participants and was conducted in two sessions of 16 participants each. Thus, each session consisted of four groups of four participants, and each participant made decisions over 10 periods.
Each session implemented one experimental condition. Within each session, group assignment and role assignment were implemented by the z-Tree program. In the fixed-matching condition, participants were randomly assigned to groups of four at the beginning of the first period, and group membership then remained unchanged for all 10 periods. In the random-rematching condition, groups of four were randomly re-formed in each period. In the RL treatment, one group member was randomly assigned as the first mover in each period. In the VL-RC and DL-RC treatments, one group member was randomly selected as the leader candidate in each period and then decided whether to accept the leadership role. Participants were informed of the relevant matching and role-assignment procedures before making decisions.
Participants signed in simultaneously at the scheduled session time and completed the controlled online laboratory experiment remotely using z-Tree version 4.1.6 (
Fischbacher, 2007) under online supervision by the experimenter. The experimenter guided participants through the instructions in an online meeting. Participants were instructed not to communicate with one another during the experiment, and any questions were answered privately by the experimenter. At the end of each period, participants received anonymous feedback on current-period roles, contributions, total group contribution, and payoffs. Role histories were not displayed to participants. No participants were excluded from the analysis, and there was no attrition or technical failure during the experiment.
Each session lasted approximately 20 min. Earnings were accumulated over 10 periods at an exchange rate of 1 token = 0.05 CNY. The average payment received was about 17.92 CNY per subject, including a 5 CNY show-up fee, which is comparable to standard online payment levels in mainland China.
4. Results
This section presents the experimental findings, structured as follows. We begin by analyzing the average contribution levels across all treatments, followed by a comparison between groups with and without a designated leader. Subsequently, we explore the behavioral dynamics inherent in leader–follower interactions. Additionally, we investigate the strategic incentives for free-riding within endogenous leadership mechanisms. Finally, we assess the impact of matching protocols by comparing outcomes from fixed and random matching rules.
4.1. Comparison of Average Contributions
Result 1. Leadership does not produce a uniform increase in contributions across different mechanisms. DL-RC exhibits the lowest observed mean contribution overall, especially under random rematching; however, the primary confirmatory comparisons do not reveal a statistically significant general effect of the mechanism.
Table 2 reports average individual contributions across the four leader-selection mechanisms. The overall averages are 7.052 in control, 7.813 in RL, 7.698 in VL-RC, and 5.578 in DL-RC. Thus, DL-RC has the lowest descriptive average contribution overall.
Table 3 presents the results separately by matching protocol. In fixed matching, the average contributions are similar in all four mechanisms: 8.038 in control, 7.794 in RL, 8.053 in VL-RC, and 7.044 in DL-RC. Under random rematching, the corresponding averages are 6.066, 7.831, 7.344, and 4.113. The largest descriptive differences, therefore, arise under random rematching.
The mixed-effects estimates are reported in
Appendix B Table A1. The overall mechanism test is not statistically significant (
,
). The protocol-specific omnibus tests are also not statistically significant under fixed matching (
,
) or random rematching (
,
). In addition, the mechanism-by-matching interaction is not statistically significant (
,
).
Appendix B Table A2 reports all six pairwise mechanism contrasts for the overall sample and separately within each matching protocol, with Holm correction applied within each comparison family. None of the primary model-based contrasts is statistically significant at the 5% level. The largest protocol-specific estimate occurs under random rematching, where DL-RC is 3.719 tokens lower than RL (unadjusted
; Holm-adjusted
).
Appendix B Table A3 reports a session-level HC2 sensitivity analysis. The mechanism-by-matching interaction remains statistically insignificant (
), and most pairwise contrasts are also insignificant. Under random rematching, however, DL-RC is estimated to be 3.231 tokens lower than VL-RC (unadjusted
; Holm-adjusted
). Taken together, the lower observed mean in DL-RC is a consistent descriptive pattern supported by this one session-level sensitivity comparison, but it is not established as a broadly robust treatment effect by the primary omnibus or Holm-adjusted model-based pairwise analyses.
Figure 1 presents average individual contributions across periods, pooling the two matching protocols. Contributions display an overall downward tendency in all four mechanisms, although the trajectories fluctuate over time. DL-RC begins at the highest descriptive level in Period 1, declines sharply after Period 3, and has the lowest average contribution from Period 4 onward. The mixed-effects trajectory model rejects the joint null that the mechanism-by-period and mechanism-by-matching-by-period interaction terms are zero (
,
, see
Appendix B Table A4). This result indicates that contribution paths vary jointly across mechanisms and matching protocols, but it does not identify a specific pairwise difference in trajectories.
4.2. Comparison of Average Contributions Between Groups with and Without a Leader
Result 2. Within the endogenous mechanisms, observed contributions differ significantly between the realized acceptance and refusal branches. In VL-RC, the accepted one-leader branch exhibits higher observed contributions than the refusal-contingent leaderless branch. In DL-RC, the refusal-contingent three-first-mover branch shows lower observed contributions than the accepted one-leader branch. Since these branches are endogenously selected, the comparisons are conditional and not causal.
The comparisons in this subsection are conditional on realized leadership states. In VL-RC and DL-RC, these states are generated by the selected candidate’s decision to accept or decline the leadership role and are therefore endogenously selected. The results describe how contributions differ across the branches generated by each mechanism, but they do not identify the causal effects of accepting leadership, declining leadership, remaining leaderless, or having one rather than three first movers.
Table 4 presents average individual contributions separately by the realized leadership state and matching protocol. Within VL-RC, average contributions are higher when the selected candidate accepts the leadership role than when the candidate declines and the group remains leaderless. The corresponding averages are 12.043 versus 2.371 under fixed matching and 11.038 versus 3.829 under random rematching.
Within DL-RC, groups in the accepted one-leader branch contribute more than groups in the refusal-contingent three-leader branch. The corresponding averages are 10.163 versus 5.785 under fixed matching and 7.750 versus 3.056 under random rematching.
All four comparisons remain statistically significant after Holm correction (). However, these realized states are generated by endogenous acceptance or refusal decisions. The comparisons therefore describe conditional differences between observed branches and should not be interpreted as causal effects of leadership acceptance, leadership absence, or leader multiplicity. In particular, the three-leader DL-RC branch combines the candidate’s refusal, a change in the contribution sequence, and the assignment of multiple first movers.
4.3. Comparison of Leader–Follower Interactions
Result 3. In a leader-present environment, first movers contribute more on average than later movers. After accounting for the dependence between decisions made within the same group-period and correcting for multiple comparisons, the first-mover versus later-mover gap remains statistically significant under fixed matching only in the accepted one-leader branch of DL-RC. In contrast, under random rematching, this gap is statistically significant in all four realized leadership states. Across leader-present states, later-mover contributions are also positively correlated with the observed mean contribution of the first mover or movers in the same group-period.
Table 5 reports descriptive individual contributions by role and realized leadership state. Under fixed matching, first movers contribute an average of 10.300 tokens in RL and 14.191 tokens in the accepted VL-RC branch, compared with 6.958 and 11.326 tokens among later movers, respectively. Within DL-RC, the corresponding first- and later-mover averages are 15.130 and 8.507 in the accepted one-leader branch, but only 6.450 and 3.789 in the refusal-contingent three-leader branch.
A similar descriptive ordering appears under random rematching. First movers contribute 11.225 tokens in RL and 16.231 tokens in the accepted VL-RC branch, compared with 6.700 and 9.308 tokens among later movers. Within DL-RC, first- and later-mover contributions average 13.667 and 5.778 in the accepted one-leader branch, compared with 3.591 and 1.452 in the refusal-contingent three-leader branch.
Since first- and later-mover decisions are generated within the same group-period, formal inference is based on the within-group-period contribution gap, defined as the mean contribution of the first mover or movers minus the mean contribution of the later mover or movers. This approach directly compares the decisions generated within the same group-period rather than treating the two roles as independent samples.
Under fixed matching, the estimated gap is positive in all four realized leadership states. The adjusted gaps are 3.347 tokens in RL (95% CI ; Holm-adjusted ), 3.960 tokens in the accepted VL-RC branch (95% CI ; Holm-adjusted ), 6.875 tokens in the accepted one-leader DL-RC branch (95% CI ; Holm-adjusted ), and 2.568 tokens in the refusal-contingent three-leader DL-RC branch (95% CI ; Holm-adjusted ). After Holm correction, only the accepted one-leader DL-RC gap remains statistically significant under fixed matching.
Under random rematching, the estimated gaps are 4.528 tokens in RL (95% CI
), 6.911 tokens in the accepted VL-RC branch (95% CI
), 8.057 tokens in the accepted one-leader DL-RC branch (95% CI
), and 2.095 tokens in the refusal-contingent three-leader DL-RC branch (95% CI
). All four gaps remain statistically significant after Holm correction (
;
Appendix B Table A8).
The lower later-mover contribution in the accepted one-leader DL-RC branch is not explained by weaker first-mover contributions. Under fixed matching, accepted first movers contribute 14.191 tokens in VL-RC and 15.130 tokens in the accepted one-leader DL-RC branch. The adjusted DL-RC-minus-VL-RC difference is 1.781 tokens (95% CI ; Holm-adjusted ). Under random rematching, the corresponding means are 16.231 and 13.667, and the adjusted difference is tokens (95% CI ; Holm-adjusted ). Thus, accepted first movers do not contribute significantly less in DL-RC than in VL-RC under either matching protocol.
More importantly, the focused later-mover analysis shows that the same observed first-mover contribution is followed by less cooperation in DL-RC. After controlling for the observed first-mover contribution, matching protocol, and period fixed effects, later movers in the accepted one-leader DL-RC branch contribute 3.015 tokens less than those in the accepted VL-RC branch (, 95% CI , ). A one-token increase in the observed first-mover contribution is associated, on average, with a 0.521-token increase in the later-mover contribution (, 95% CI , ). Taken together, these estimates indicate that the DL-RC environment is less effective at translating a given first-mover contribution into subsequent cooperation.
The interaction analysis further identifies where this weaker transmission arises. Under fixed matching, a one-token increase in the observed first-mover contribution is associated with a 0.756-token increase in later-mover contributions in VL-RC, but only a 0.356-token increase in the accepted one-leader DL-RC branch. The DL-RC-minus-VL-RC difference in response slopes is (, 95% CI ; unadjusted ; Holm-adjusted ). Later movers therefore respond significantly less strongly to the first-mover contribution in DL-RC under fixed matching.
Under random rematching, the corresponding response slopes are 0.353 in VL-RC and 0.472 in DL-RC. Their difference is 0.119 (, 95% CI ; Holm-adjusted ). The difference between the fixed- and random-matching slope contrasts is itself statistically significant (difference , , 95% CI , ). The weaker follower responsiveness in DL-RC is therefore concentrated in the fixed-matching environment.
The broader later-mover analysis yields the same central pattern that first-mover behavior is transmitted to subsequent contributors. Using all leader-present states, a one-token increase in the observed mean contribution of the first mover or movers is associated with a 0.526-token increase in the later-mover contribution (
, 95% CI
,
;
Appendix B Table A11). This result confirms a strong positive relationship between the contribution signal sent by first movers and the contribution subsequently selected by later movers.
The distribution of contributions further clarifies where the lower observed contributions in DL-RC are concentrated. Defining a low first-mover contribution as no more than 5 of the 20 available tokens, the share of low first-mover observations under fixed matching is 54.12% in DL-RC but only 10.64% in VL-RC. Under random rematching, the corresponding shares are 73.04% and 7.69%. Thus, low first-mover contributions occur substantially more often in DL-RC, particularly in the refusal-contingent three-first-mover branch. Along with the lower follower responsiveness estimated under fixed matching, these results reveal two associated behavioral patterns: more frequent low first-mover contributions in the refusal-contingent branch and weaker transmission of a given first-mover contribution in the accepted branch. However, they do not establish the unmeasured psychological mechanisms underlying these patterns.
Figure 2 and
Figure 3 present the contribution trajectories by mover role under fixed matching and random rematching, respectively. Across most periods, first movers contribute more than later movers. The figures also show that the low DL-RC contributions are concentrated in the refusal-contingent three-leader branch, where both first- and later-mover contributions remain substantially below those in the accepted one-leader branch.
4.4. Discussion of Endogenous Leadership
Result 4. Endogenous leadership mechanisms generate different realized group structures. DL-RC produces the refusal-contingent three-first-mover structure more frequently than VL-RC produces a leaderless structure, and the lower observed contribution in DL-RC is descriptively concentrated in that refusal-contingent branch. These realized-state comparisons remain conditional on endogenous candidate decisions.
Table 6 presents candidate acceptance rates and the resulting leadership structures. Under fixed matching, candidates accept in 47 of 80 VL-RC candidate-periods, corresponding to an acceptance rate of 58.75%, compared with 23 of 80 candidate-periods, or 28.75%, in DL-RC. Under random rematching, the corresponding acceptance rates are 48.75% in VL-RC and 22.50% in DL-RC.
We formally compare candidate acceptance using mixed-effects logistic regression models that account for repeated candidate decisions. In the full candidate-period model, the adjusted probability of acceptance is 57.06% in VL-RC and 27.11% in DL-RC. The difference between DL-RC and VL-RC is therefore percentage points (, 95% CI , ). Candidates are considerably less willing to enter the voluntary one-leader branch when refusal transfers first-mover responsibility to the other three group members.
The mechanism difference does not vary significantly across matching protocols, , . Nevertheless, the protocol-specific estimates differ in precision. Under fixed matching, the adjusted acceptance probabilities are 62.97% in VL-RC and 30.53% in DL-RC, resulting in a difference of percentage points (, 95% CI ; Holm-adjusted ). Under random rematching, the adjusted probabilities are 48.28% and 25.21%, respectively, yielding a difference of percentage points (, 95% CI ; Holm-adjusted ).
A session-level sensitivity analysis using HC2 standard errors maintains the negative direction of all three comparisons but yields substantially wider confidence intervals due to the analysis including only two sessions per mechanism-by-matching cell. In this sensitivity analysis, the random-rematching difference remains statistically significant after Holm correction, whereas the overall and fixed-matching estimates are imprecisely estimated. The complete mixed-effects and session-level results are reported in
Appendix B Table A13 and
Table A14.
The differences in acceptance decisions directly result in distinct realized group structures. Under fixed matching, VL-RC produces one voluntary first mover in 58.75% of group-periods and remains leaderless in the remaining 41.25%. By contrast, DL-RC produces one voluntary first mover in only 28.75% of group-periods and enters the three-delegated-first-mover branch in 71.25% of group-periods. Consequently, the average number of first movers per group-period is 0.588 in VL-RC but 2.425 in DL-RC.
The same structural contrast is observed under random rematching. VL-RC produces one voluntary first mover in 48.75% of group-periods and no first mover in 51.25%. DL-RC produces one voluntary first mover in 22.50% of group-periods and three delegated first movers in 77.50%, resulting in an average of 2.55 first movers per group-period.
Figure 4 and
Figure 5 illustrate the period-by-period development of these structures. In VL-RC, the average number of total first movers remains below one because a candidate’s refusal results in the group having no first mover. In DL-RC, the average number of total first movers generally exceeds two, as candidate refusal triggers the three-delegated-first-mover branch. The figures also indicate that the majority of first movers observed in DL-RC are delegated rather than voluntary.
These structural findings clarify the origin of the realized branches examined in Results 2 and 3. The low-contribution three-first-mover branch is not an occasional outcome within DL-RC; rather, it represents the mechanism’s modal realized state under both matching protocols. However, the current design does not independently manipulate the refusal event and the number of first movers. Therefore, the evidence from Result 4 establishes that DL-RC generates a substantially different acceptance pattern and realized role structure.
4.5. Comparison Between Fixed Matching and Random Matching
Result 5. Fixed matching exhibits a higher observed overall mean contribution than random rematching; however, the pooled difference is not statistically significant. The clearest statistically significant within-mechanism matching contrast is observed in DL-RC.
Table 7 first reports the overall comparison and then separates group-periods according to whether a leader is present. Across the complete sample, the average group-period contribution is 7.732 under fixed matching and 6.338 under random rematching. The mixed-effects model estimates a fixed-minus-random difference of 1.394 tokens, but this difference is not statistically significant (
, 95% CI
,
). Adding leader-selection-mechanism fixed effects leaves the estimate unchanged (
, 95% CI
,
).
Conditional on the realized leadership structure, the descriptive fixed-matching advantage remains visible but is not statistically significant. Among leader-present group-periods, the adjusted fixed-minus-random difference is 1.231 tokens (, 95% CI ; raw ; Holm-adjusted ). Among leaderless group-periods, the corresponding estimate is 1.264 tokens (, 95% CI ; raw ; Holm-adjusted ). Thus, the data do not support a general conclusion that fixed matching increases contributions either whenever a leader is present or whenever no leader emerges.
Table 8 examines the matching comparison within a fixed leader-selection mechanism. Among leader-present group-periods, contributions are virtually identical across matching protocols in RL. The estimated fixed-minus-random difference is
tokens (
, 95% CI
; Holm-adjusted
). In the realized leader-present branch of VL-RC, the corresponding estimate is 0.428 tokens (
, 95% CI
; Holm-adjusted
).
A different pattern emerges within DL-RC. Contributions average 7.044 under fixed matching and 4.113 under random rematching. The adjusted fixed-minus-random difference is 2.931 tokens (, 95% CI ; raw ; Holm-adjusted ). By contrast, neither of the leaderless comparisons is statistically significant. The estimated difference is 1.972 tokens in control (, 95% CI ; Holm-adjusted ) and 0.583 tokens in the realized leaderless branch of VL-RC (, 95% CI ; Holm-adjusted ).
These estimates indicate that the observed matching-protocol difference is not a general feature of leader-present or leaderless groups. Instead, the clearest difference is concentrated within the complete DL-RC institutional environment. Because DL-RC combines candidate acceptance or refusal, the resulting first-mover assignment, and the associated role structure, this comparison does not isolate a single behavioral mechanism through which fixed matching affects contributions.
We next examine whether the DL-RC difference is concentrated among first movers or later movers.
Table 9 reports role-specific comparisons. Pooling all leader-present mechanisms, leaders contribute an average of 9.165 under fixed matching and 7.570 under random rematching, whereas followers contribute 8.028 and 6.552, respectively. After controlling for the mechanism and period and preserving within-group-period dependence, neither pooled role-specific contrast is statistically significant. The adjusted fixed-minus-random difference is 0.690 tokens for leaders (
, 95% CI
; Holm-adjusted
) and 1.578 tokens for followers (
, 95% CI
; Holm-adjusted
). The matching-by-role interaction is also statistically insignificant,
,
.
Within DL-RC, the fixed-minus-random difference is positive for both roles. The leader difference is 3.246 tokens (, 95% CI ; raw ; Holm-adjusted ). The follower difference is 2.587 tokens (, 95% CI ; raw ), but it does not remain statistically significant after correction across the six mechanism-by-role comparisons (Holm-adjusted ). No role-specific matching contrast is statistically significant in RL or VL-RC. The DL-RC pattern is therefore shared directionally by first and later movers, but the statistical evidence is clearest among first movers.
The supplementary mechanism-by-matching interaction in the full sample is statistically insignificant, , . Accordingly, the significant conditional DL-RC comparison should not be interpreted as formal evidence that the matching effect is statistically larger in DL-RC than in every other mechanism.
The session-level HC2 analyses produce the same overall direction but wider uncertainty. The overall fixed-minus-random estimate remains 1.394 tokens (, 95% CI , ). Within leader-present DL-RC, the session-level estimate is 2.931 tokens (, 95% CI ; raw ), but the Holm-adjusted p-value is 0.071. The DL-RC leader estimate is also positive in the session-level analysis but does not remain significant after adjustment (Holm-adjusted ). Because each mechanism-by-matching condition contains only two sessions, these estimates are interpreted as sensitivity evidence.
Taken together, fixed matching does not produce a statistically significant contribution increase across the full sample or across all leader-present or leaderless group-periods. The main conditional pattern is concentrated in DL-RC, where fixed matching is associated with higher contributions and the role-specific difference is most precisely identified among first movers. Stable-group membership may therefore be particularly relevant within the refusal-contingent delegation environment represented by DL-RC, but the experiment does not directly identify reputation, reciprocity, coordination, or accountability as the mechanism producing this pattern.
5. Conclusions
We study how leader-selection institutions affect cooperation in a sequential public goods game. Using a unified experimental design, we compare a leaderless baseline (control) with three leadership-selection rules (RL, VL-RC, and a novel DL-RC) under both fixed and random matching.
Our analysis yields five main conclusions. First, leadership does not uniformly increase contributions. DL-RC exhibits the lowest observed mean contribution overall, particularly under random rematching; however, the primary omnibus and Holm-adjusted model-based pairwise tests do not establish a general mechanism effect. One session-level sensitivity comparison supports the descriptive ranking. Second, observed contributions differ substantially across the realized acceptance and refusal branches of VL-RC and DL-RC. Because candidates endogenously choose whether to accept leadership, these branch comparisons are conditional and do not identify causal effects of acceptance, refusal, leader absence, or the number of first movers. Third, first movers generally contribute more than later movers, and later-mover contributions are positively associated with observed first-mover contributions. Fourth, the lower observed contribution in DL-RC is descriptively concentrated in the refusal-contingent three-first-mover branch. Finally, fixed matching has a higher observed overall mean than random rematching, but the pooled difference is not statistically significant; the clearest within-mechanism matching contrast occurs in DL-RC.
This study extends prior research (
He & Zheng, 2024;
Xu et al., 2025) by comparing identical leader-selection mechanisms across fixed-matching and random-rematching environments. The findings reveal an association between the complete DL-RC institutional arrangement and lower observed contributions. In DL-RC, candidate refusal transfers first-mover responsibility to the remaining three group members, and the lower observed contributions are descriptively concentrated in this refusal-contingent branch. This pattern may reflect coordination difficulties, negative inferences from refusal, concerns about the legitimacy of delegated roles, or diminished individual accountability; however, none of these mechanisms were directly measured.
A key limitation of the design is that it does not separately identify the effect of leader multiplicity from the refusal event that triggers delegated leadership. Three delegated first movers emerge only when the initially selected candidate declines the role. Therefore, we interpret the branch comparison as evidence of conditional outcomes within the DL-RC institutional package, rather than as a causal estimate of the effect of multiple leadership. A natural extension would be to assign the first three movers exogenously from the outset, without a prior refusal, thereby isolating leader multiplicity from refusal-contingent delegation.
The experiment is a controlled online laboratory study based on standard VCM parameters and a student sample. Future research could examine whether the observed associations generalize across alternative institutional settings, including heterogeneous endowments, asymmetric information, communication opportunities, and larger groups, as well as to participant populations for whom leadership and collective-action decisions are substantively relevant. Incorporating direct measures of beliefs, expectations, perceived accountability, legitimacy, fairness, and reputational concerns would also help identify the behavioral mechanisms underlying responses to voluntary and delegated leadership.