Next Article in Journal
The Relationship Between Stablecoin and Cryptocurrency Returns During Periods of Market Stress
Previous Article in Journal
Predictive Analytics Approaches to Modeling Bitcoin Prices During Periods of Financial Uncertainty
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Evidence on Settlement-Window Price Divergence in Bitcoin Prediction Markets

School of Computing Sciences and Computer Engineering, The University of Southern Mississippi, Hattiesburg, MS 39406, USA
*
Author to whom correspondence should be addressed.
FinTech 2026, 5(3), 67; https://doi.org/10.3390/fintech5030067
Submission received: 8 May 2026 / Revised: 19 July 2026 / Accepted: 22 July 2026 / Published: 1 August 2026

Abstract

This paper investigates whether prediction market settlements create incentives for temporary price pressure in Bitcoin spot markets. Using high-frequency data from February 2025 to January 2026 and actual contract-level data from Polymarket and Kalshi to identify economically relevant contract strikes, we document basis divergence between settlement oracle exchanges (Coinbase) and non-constituent exchanges (Binance) during expiry windows. Employing a difference-in-differences framework with month fixed effects, we find that a one standard deviation increase in strike proximity is associated with a 6.7 basis point constituent exchange price deviation during settlement windows. The estimate is precise under the baseline minute-level HAC specification, while exact paired-month permutation inference based on 12 settlement events yields p = 0.0256 ; equal-weight event aggregation produces a larger negative estimate, indicating event heterogeneity. Monthly directional patterns are suggestive, though stricter event-level and above-versus-below-strike tests provide mixed evidence on directional asymmetry. Taken together, these findings provide reduced-form evidence consistent with settlement-related incentives and may raise broader settlement-design considerations for decentralized financial systems. However, the analysis does not directly observe trader intent or the underlying mechanism.

1. Introduction

The rapid growth of prediction markets in 2024–2025 represents one of the most significant innovations in decentralized finance. Platforms such as Polymarket and Kalshi have facilitated billions of dollars in trading volume on outcomes ranging from political elections to cryptocurrency prices. These markets have emerged as important mechanisms for information aggregation [1,2], with recent evidence suggesting improved accuracy in long-run forecasting [3,4]. Unlike traditional derivatives, these platforms offer binary options with discontinuous payoff structures: contracts pay $1.00 if a condition is met at settlement and $0.00 otherwise. This paper examines an important design question in financially-settled prediction markets: the potential for settlement rules to create incentives that affect trading in the spot markets used as settlement oracles. When the expected profit from securing a favorable settlement outcome exceeds the cost of temporarily distorting spot prices on constituent exchanges, rational actors may face settlement-related incentives that are reflected in observed price pressure patterns.
In this paper, a settlement oracle is the external price source or benchmark used to determine whether a prediction-market contract pays $1.00 or $0.00. An oracle-constituent exchange is a spot venue whose trades can enter the settlement benchmark or reference price in the institutional setting studied here; a non-constituent exchange is a comparison venue whose trades do not directly enter that benchmark, although it may still co-move with constituent venues through arbitrage and information transmission. In the baseline empirical design, Coinbase is treated as the benchmark-relevant venue and Binance as the non-constituent comparison venue; Kraken and Bitstamp are discussed only within the institutional setting of the Bitcoin benchmark family.
We focus on Bitcoin prediction markets for three reasons. First, Bitcoin contracts represent a substantial share of financial prediction market volume. Second, Bitcoin’s fragmented global liquidity across multiple exchanges creates natural variation in settlement index composition. Third, high-frequency data availability enables rigorous microstructure analysis of settlement-window pricing dynamics. Our empirical strategy exploits the fact that the CME CF Bitcoin Reference Rate—used by regulated prediction markets including Kalshi—is constructed from a subset of exchanges (Coinbase, Kraken, Bitstamp) while excluding the world’s largest spot venue (Binance). This creates an empirical comparison: if settlement rules are associated with settlement-related incentives, prices on constituent exchanges may decouple from non-constituent exchanges during settlement windows, particularly when Bitcoin trades near strike prices where payoff sensitivity is especially high. A key methodological innovation in our study is the use of actual contract-level data from Polymarket and Kalshi to identify economically relevant contract strikes. We select the strike nearest to the settlement window price with meaningful trading volume (exceeding $100,000). This approach helps ensure that we measure incentives at economically relevant strikes where any settlement-related effects are most likely to be visible in the data.
Using a difference-in-differences framework with 12 months of minute-level data, we document statistically significant and economically relevant price divergence. After controlling for Tether premium fluctuations—a critical confound given Binance’s USDT denomination—we find that a one standard deviation increase in strike proximity is associated with a 6.7 basis point cross-exchange price deviation between constituent and control exchanges during settlement windows ( p < 0.001 ). At high incentive levels, the estimated magnitudes reach roughly 14–20 basis points. Monthly directional patterns are suggestive of the proposed incentive framework. When Bitcoin settles near strikes, the sign of the expiry-minus-placebo basis often aligns with the payoff-relevant direction. However, formal above-versus-below-strike regression tests are weaker, so the directional evidence should be interpreted as supportive but not definitive.
Our robustness analysis suggests that strike identification is important for detecting settlement-window price divergence associated with incentives.
Our findings contribute to three literatures. First, we extend research on limits to arbitrage [5] by identifying settlement windows as periods when arbitrage may break down due to asymmetric information about the source of observed price pressure. Second, we build on the benchmark manipulation literature [6] by documenting analogous pricing patterns in decentralized markets. Third, we add to the growing body of work on oracle risks in blockchain-based financial systems.
The remainder of this paper proceeds as follows. Section 2 reviews related literature. Section 3 develops theoretical predictions. Section 4 describes data and methodology. Section 5 presents results. Section 6 discusses implications and possible design considerations. Section 7 concludes.

2. Literature Review

Recent work on modern prediction markets documents cross-platform price disparities, arbitrage opportunities, and price-discovery differences across Polymarket, Kalshi, PredictIt, and Robinhood [7]. This evidence motivates treating platform design, liquidity segmentation, and settlement rules as central features of the current prediction-market environment.

2.1. Limits of Arbitrage and Benchmark Manipulation

Our theoretical framework builds on [5], who demonstrate that arbitrage opportunities can persist when arbitrageurs face capital constraints, implementation costs, and fundamental risk. In our setting, the settlement window represents a unique period where traditional arbitrage logic breaks down: while global market prices may deviate from constituent exchange prices, arbitrageurs cannot be certain whether the deviation reflects temporary settlement-related price pressure (which may revert after settlement) or genuine information (which should be incorporated). Documented distortions in financial benchmarks have been extensively discussed in traditional markets. Regulatory enforcement actions documented attempted manipulation and false reporting in LIBOR submissions [8], illustrating the vulnerability of benchmark fixings to strategic input. This experience also motivated work on robust benchmark design [9]. Evans [6] documents manipulation of the WM/Reuters 4pm London FX fix, showing that traders “bang the close” by concentrating orders during the brief calculation window. Our contribution is to extend this literature to decentralized prediction markets, where patterns consistent with the proposed incentive framework arise from binary option payoffs rather than directional trading profits.

2.2. Option Expiration Effects

Ni et al. [10] document “pinning” in equity options, where stock prices cluster near strikes with high open interest on expiration days. They attribute this to delta-hedging by market makers who unwind positions as expiration approaches. However, prediction markets create fundamentally different incentives. Binary options cannot be effectively delta-hedged near expiry due to discontinuous payoffs—gamma exhibits mathematically discontinuous behavior (approaching a Dirac-delta function) at the strike. Within this stylized setting, the market maker problem shifts from “hedge exposure” to price movements around settlement-relevant levels.
Pearson et al. [11] further show that pinning effects are strongest for options with large open interest relative to underlying spot volume. This observation motivates our focus on Bitcoin, where prediction market notional has grown to meaningful fractions of spot market liquidity during specific settlement windows.

2.3. Cryptocurrency Market Microstructure

The Tether adjustment in this paper addresses a narrow mechanical denomination issue: Binance quotes BTC/USDT, whereas Coinbase quotes BTC/USD. Any departure of USDT from one U.S. dollar therefore enters an unadjusted cross-exchange basis mechanically, so we convert Binance prices to USD-equivalent terms using the contemporaneous USDT/USD rate. Prior work showing that Tether can interact with Bitcoin pricing more broadly provides useful background [12], but our adjustment is not an application of that empirical design. Recent stablecoin research and institutional evidence suggest that stablecoin pegs remain imperfect but that large deviations are episodic rather than a normal state of trading [13,14,15]. Consistent with that interpretation, USDT/USD deviations in our February 2025–January 2026 event-window sample are modest: across the nonmissing one-minute observations used for the paper’s event windows, the median absolute deviation from parity is approximately 2.2 basis points and the maximum is approximately 17.6 basis points. The adjustment should therefore be interpreted as a necessary denomination correction, not as the main economic driver of the settlement-window result. Because the maximum observed USDT/USD deviation is of the same order as the estimated settlement-window effects, measurement error in the minute-level conversion rate could affect estimated magnitudes. The USD-denominated Coinbase–Kraken comparison provides partial reassurance that the baseline result is not purely a USDT-conversion artifact, although it is not a direct validation using an independent USDT/USD rate. More broadly, cryptocurrency markets exhibit unique microstructure features—24/7 trading, fragmented liquidity, minimal regulation—that may amplify settlement-related price-pressure opportunities relative to traditional markets [16]. The absence of circuit breakers, consolidated tape, or market surveillance creates an environment where localized price pressure can persist without triggering intervention. Furthermore, the rise of automated market makers in decentralized prediction markets [17] has created new channels of oracle-related settlement risk that differ fundamentally from traditional centralized exchange dynamics. Recent DeFi evidence also emphasizes that protocol design, incentive structure, arbitrage frictions, transparency, and settlement mechanisms shape market outcomes in blockchain-based financial systems [18]. In prediction markets specifically, Rahman et al. [19] organize decentralized market design around infrastructure, trading, resolution, settlement, and archiving, highlighting manipulation resistance as a core design concern.

2.4. Oracle Risk in Decentralized Finance

Recent theoretical work has identified “Oracle Extractable Value” (OEV) as an emerging risk in blockchain-based financial systems [20]. When smart contracts rely on external data feeds (oracles) to determine payoffs, attackers may find it profitable to manipulate the underlying oracle inputs rather than the contracts themselves [21]. Our evidence on settlement-window divergence is suggestive that OEV may be relevant in these settings beyond a purely theoretical concern. This framing is closely related to the broader oracle economics literature, which studies the trade-offs between scalability, decentralization, and truthfulness in systems that depend on external data inputs [22]. It is also related to the broader literature on maximal extractable value, where public blockchain infrastructure can create strategic extraction opportunities and allocative inefficiencies through market-design frictions [23].

3. Theoretical Framework and Hypotheses

3.1. Settlement Incentives

Consider a trader who holds N binary option contracts with strike K that settle based on the price observed on constituent exchange i at time T. The contract pays:
Payoff = N × $ 1.00 if P i , T K 0 if P i , T < K
If the global market price P global , T is trading just below K as settlement approaches, the “No” position holder (who profits if P i , T < K ) may face incentives to trade in ways that keep constituent exchange i below K even if the global market breaks above it. The potential net payoff from such settlement-related trading pressure is:
Π = N × $ 1.00 C impact ( Q , σ ) C inventory
where:
  • C impact is the cost of selling quantity Q on exchange i to place downward pressure on price
  • σ is market volatility (higher volatility increases slippage)
  • C inventory is the risk of holding the resulting short spot position
Settlement-related price pressure may become more attractive under this stylized framework when Π > 0 . As N (open interest) increases, the required Q (trading volume needed to affect settlement-relevant pricing) decreases relative to the payoff, increasing the scope for such behavior consistent with the incentive logic developed above.

3.2. Testable Predictions

This stylized framework suggests three testable predictions:
Hypothesis 1 (Decoupling). 
During settlement windows, prices on constituent exchanges may diverge from non-constituent exchanges if settlement-related incentives are operative, particularly when spot price is near a strike with high open interest.
Hypothesis 2 (Directionality). 
Any such divergence may be directional under the stylized incentive setting considered here: negative (suppression) when price is below strike, positive (support) when price is above strike.
Hypothesis 3 (Reversion). 
Price convergence may resume after settlement is finalized.
H1 is evaluated with the main interaction regression and the expiry-placebo basis comparisons. H2 is evaluated with the monthly directional table and exact sign-alignment test. H3 is evaluated descriptively with the event-time figures and the August case study, which examine whether divergence is temporary rather than persistent.

3.3. Cross-Asset Qualitative Comparison: Centralized Closing Auctions

As a qualitative comparison, applying a similar construction to U.S. equity markets (SPY and QQQ) does not reveal a comparable settlement-window pattern, consistent with the more centralized and coordinated structure of equity index pricing. Unlike Bitcoin, equity settlement is organized around centralized closing auctions with deeper liquidity, tighter surveillance, and different venue coordination. The equity setting also differs in data sources, venue definitions, and event construction, so this comparison should be interpreted as institutional context rather than as a formal test.

4. Data and Methodology

4.1. Data Sources and Settlement Rules

We construct a high-frequency dataset covering February 2025 through January 2026. The dataset includes:
  • Spot prices: 1-min OHLCV data for BTC/USD (Coinbase) and BTC/USDT (Binance)
  • Tether rates: 1-min USDT/USD exchange rate from Coinbase Pro
  • Settlement dates: Monthly expiry schedule for Bitcoin prediction markets
  • Strike prices: Identified from actual Polymarket and Kalshi contract data (see Section 4.3)
All timestamps are synchronized to UTC. The settlement window for CME CF Bitcoin Reference Rate is 15:00–16:00 London time, corresponding to the hourly calculation period specified in the CME methodology.
Table 1 summarizes the reduced-form research design and its principal identifying assumptions.

4.2. Event Windows and Placebo Construction

Our sample focuses on windows surrounding each monthly settlement event from 60 min before settlement begins through 90 min after settlement starts. The interval spans 150 elapsed minutes but, because both endpoints are included, contains 151 one-minute timestamps per event window. We also collect identical windows on the 15th of each month as placebo controls where available in the paper-facing regression panel. This yields a final regression sample of 3624 min-level observations across 12 expiry event windows and 12 placebo event windows ( 24 × 151 = 3624 ). For event-time visualizations below, we center the analysis at settlement initiation. Figure 1 reports a broader descriptive window, while Figure 2 uses a symmetric window of t [ 60 , + 60 ] for visualization clarity. The regression specification uses the full 150-min span, or 151 included one-minute timestamps, for estimation.

4.3. Strike Identification

We do not identify the relevant strike using the highest-volume Bitcoin contract on a given expiry date. That rule is economically incorrect because total monthly volume can be concentrated in contracts that are far from the settlement price or tied to longer-horizon speculation rather than the discrete payoff threshold that becomes most settlement-relevant during the window. An incentive-based measure should instead target the strike that is locally payoff-relevant at the moment settlement is determined.
Our revised procedure anchors strike selection to the settlement window itself. For each monthly expiry, we first compute the average Coinbase Bitcoin price during the 15:00–16:00 London settlement window. We then collect all Polymarket Bitcoin binary contracts expiring on that exact date, restrict attention to contracts with meaningful trading activity (volume greater than $100,000), and select the strike with the smallest absolute distance to the settlement-window average price. We retain only strikes that lie within 5 percent of the settlement price, which ensures that the contract remained economically live during settlement. Table 2 illustrates the construction for one representative minute, and Table 3 reports the identified strikes along with corresponding settlement prices and contract volumes.
This procedure improves economic relevance but also creates an ex-post concern: the settlement-window price is observed after the event has occurred. The concern is not that the strike itself is invented ex post—the contracts and strikes existed before settlement—but that selecting the nearest economically live contract using the realized settlement-window average may mechanically favor strikes that look relevant after the fact. We therefore treat the strike-selection robustness checks below as an important part of the design rather than as cosmetic sensitivity analysis.
This refinement improves identification. The selected Polymarket strike is, on average, $904.84 from the settlement-window average price (median $585.56), with 75.0% of months within $1000 and 100.0% of months within 5.0% of the settlement price. Kalshi provides a secondary comparison benchmark. The corresponding Kalshi strike is even closer on average: $351.92 mean absolute distance (median $239.44), with 83.3% of months within $500 and 91.7% within 1.0% of the settlement price. Kalshi is closer than Polymarket in 75.0% of expiry months. These statistics indicate that the revised selection rule isolates strikes that were plausibly relevant for settlement incentives rather than contracts that merely accumulated the most total volume over the month.
We use the Polymarket-based strike as the primary strike measure because the empirical question concerns incentives generated by prediction-market settlement. Kalshi is used as an external comparison device rather than a substitute measure. In the panel analysis, we construct the incentive variable as the inverse distance between the spot price and the selected strike, standardized over the full sample. The robustness checks below show that the main result is stable across alternative strike-selection and specification choices.

4.4. Basis and Tether Adjustment

A critical methodological challenge arises from comparing USD-denominated prices (Coinbase) to USDT-denominated prices (Binance). When Tether trades away from $1.00 parity, this creates mechanical basis that is unrelated to settlement-related incentives.
We address this by adjusting Binance prices for real-time Tether premium:
P Binance , USD - equivalent , m τ = P Binance , USDT , m τ × R USDT / USD , m τ
where R USDT / USD , m τ is the contemporaneous Tether exchange rate for month m and event-time minute τ . This ensures our basis calculation compares like-for-like USD values.
Basis (Dependent Variable):
Y m τ = ln ( P Coinbase , m τ ) ln ( P Binance , adjusted , m τ )
Under market efficiency, E [ Y m τ ] 0 (after accounting for stable fee premia). Significant deviations indicate localized inefficiency. The dependent variable is already a Coinbase-minus-Binance basis, so the baseline specification is not an exchange-by-time panel with separate exchange fixed effects.

4.5. Incentive Metric

To capture non-linear pressure near strikes, we construct:
I n c e n t i v e m τ = 1 | P m τ K m | + ϵ
where K m is the selected contract strike for month m and ϵ = $ 50 prevents division by zero. This metric spikes when price is close to the strike. We standardize I n c e n t i v e m τ to have mean 0 and standard deviation 1 for interpretability.
Treatment Indicator:
D Expiry , m τ = 1 if minute τ belongs to the expiry - date window for month m 0 otherwise

4.6. Empirical Specifications

We employ a difference-in-differences estimator with month fixed effects as a reduced-form test of whether settlement-window divergence is stronger when settlement-related incentives are greater:
Y m τ = α + μ m + β 1 D Expiry , m τ + β 2 I n c e n t i v e m τ + β 3 ( D Expiry , m τ × I n c e n t i v e m τ ) + X m τ γ + ε m τ
This specification captures pricing patterns consistent with the proposed incentive framework, but it does not directly observe trader intent, position concentration, or the specific mechanism generating order flow.
Where:
  • μ m are month fixed effects
  • β 1 captures the average difference associated with being in an expiry window
  • β 2 captures the relationship between incentive and basis in non-expiry periods
  • β 3 is our coefficient of interest: the incremental association between incentive pressure and basis during expiry windows
  • X m τ denotes auxiliary controls examined in robustness analyses; in those specifications, volume refers to aligned minute-level traded volume on Coinbase ( V o l Target ) and Binance ( V o l Control ), while volatility is studied separately through subsample splits rather than included in the baseline table

4.7. Inference

Standard errors are estimated using the Newey-West HAC estimator with 15 lags to account for serial correlation and heteroskedasticity in high-frequency financial data [24].
The key identifying assumption is parallel trends: in the absence of expiry treatment, the basis would have evolved similarly on expiry days and placebo days. In this setting, that assumption supports a reduced-form comparison of whether observed pricing patterns are stronger when predicted settlement-related incentives are greater. We formally test this in Section 5.2.

4.8. Robustness and Falsification Design

The monthly-actual specification is the paper’s reference design. The baseline choices of a 5% strike band, a $100,000 contract-volume threshold, a 150-min event span with inclusive one-minute endpoints, and HAC(15) standard errors are used as middle-ground implementation values for tractability and interpretability; Section 5.8 shows that the main interaction result remains stable across nearby alternatives, so these choices should be read as baseline conventions rather than as uniquely optimal parameters.
  • Placebo date (15th of the month). The 15th provides a mid-month benchmark that preserves the same calendar month and intraday clock structure as the expiry analysis while remaining far from month-end settlement. This makes it a natural falsification date without the same settlement incentives.
  • Event window (150 min). The baseline window covers 60 min before settlement begins and 90 min after settlement starts, which is intended to capture pre-settlement, settlement, and immediate post-settlement dynamics without extending so far that unrelated intraday noise dominates. This is also a standard event-study compromise between coverage and precision.
  • Strike band (5%). Restricting the sample to strikes within 5% of the settlement-window average price keeps the analysis focused on economically live contracts while excluding far out-of-the-money contracts with limited settlement relevance. The 5% cutoff also preserves a usable number of monthly observations.
  • Volume threshold ($100,000). Requiring at least $100,000 in contract volume filters out thinly traded markets that are less likely to generate economically meaningful settlement incentives. This threshold also reduces noise from idiosyncratic illiquid contracts.
  • Smoothing parameter ( ϵ = 50 ). The incentive metric uses ϵ = 50 to prevent the inverse-distance measure from becoming unstable when the spot price is extremely close to the strike. The sensitivity analysis over alternative ϵ values indicates that the sign and significance of the main interaction are not driven by this smoothing choice.
  • HAC lag choice (15). Newey-West HAC(15) standard errors are used to accommodate serial correlation and heteroskedasticity in minute-level basis data. Alternative lag choices are reported in the specification grid and do not change the sign of the main coefficient.
  • Controls and heterogeneity definitions. When volume controls are added, they correspond to aligned minute-level traded volume on Coinbase ( V o l Target ) and Binance ( V o l Control ) from the same timestamped panel. Volatility is not part of the baseline regression table; instead, the heterogeneity exercise uses a trailing 30-day realized-volatility measure, computed as the non-annualized standard deviation of daily log returns over the prior 30 days. High-volatility observations are those above the sample median of that measure, and low-volatility observations are those at or below it.
The confirmatory emphasis of the paper is on the reference monthly-actual specification and its close variants. By contrast, the volatility partition and the broader specification grid are exploratory robustness exercises used to assess how sensitive the reduced-form pattern is to reasonable implementation choices.
Software and reproducibility. The analysis is implemented in Python 3.11.11 using pandas, numpy, statsmodels, pyarrow, requests, and matplotlib. The project code constructs schedules, collects public exchange data, aligns event windows, builds incentive variables, estimates the regressions, and exports paper figures and tables. The processed analysis panels, validation files, generated figures, and generated LATEX tables are retained in the project directory so that the paper-facing regressions can be reproduced from processed data. Full raw-data replication is more limited because exchange APIs, historical endpoint availability, and external derivatives/order-book data access can change over time.

5. Empirical Results

5.1. Main Regression Results

Table 4 presents our primary findings. Model 1 shows a baseline specification without the interaction term. Model 2 introduces the key interaction between expiry status and incentive pressure, along with month fixed effects. The coefficient of interest, β 3 = 0.000674 ( SE = 0.000134 , p = 4.73 × 10 7 ), is highly statistically significant. This is consistent with a 6.74 basis point decline in the constituent-exchange basis relative to the control exchange during expiry windows when incentive pressure (price moving closer to strike) rises by one standard deviation. This pattern is suggestive of settlement-related effects, but it should be interpreted as reduced-form evidence rather than as direct evidence of trader intent or a uniquely identified mechanism. The sign of the interaction coefficient reflects the dominant marginal response at high incentive levels, rather than the unconditional average across all observations. Because the regression conditions on incentive intensity and includes month fixed effects, β 3 should be interpreted as a within-month marginal effect rather than as a visual summary of the raw average paths shown in Figure 1 and Figure 2.
Economic Magnitude Interpretation: To contextualize the coefficient, our sample mean incentive during high-pressure periods (price within $100 of strike) is 2.1 standard deviations above the mean. This corresponds to a predicted basis of:
Predicted basis = 2.1 × ( 0.000674 ) × 10,000 = 14.2 bps
This is the same order of magnitude as the observed aggregate decoupling across high-incentive expiry windows and is consistent with the linear specification capturing the broad settlement-window divergence pattern. At extreme incentive levels (3+ standard deviations), the predicted effect reaches about 20 basis points, consistent with the maximum observed deviations of 24–38 basis points across the monthly events reported later. For a Bitcoin price near $100,000, a 6.7 basis point basis movement corresponds to roughly $67 per BTC, while 14–20 basis points corresponds to roughly $140–$200 per BTC during a short settlement-relevant window. These magnitudes are informative about the size of the observed association, but they do not by themselves identify the underlying mechanism.
Hypothesis status: The main interaction result supports H1 in the limited reduced-form sense that settlement-window basis divergence is stronger when the price is closer to economically relevant strikes.
Because the panel is minute-level but the design is organized around monthly settlement events, the effective independent variation is much closer to 12 settlement events than to 3624 independent observations. The baseline HAC specification addresses serial correlation within the minute panel, but it should not be read as creating thousands of independent settlement experiments. Table 5 therefore reports more conservative inference checks that cluster or aggregate closer to the event level. These checks directly address the event-level inference concern: event-date clustering preserves the minute-level estimand, event aggregation gives each event equal weight, and exact paired-month permutation relabels treatment within matched treatment-placebo months. The coefficient remains negative across these checks, while the associated p-values weaken relative to the baseline HAC presentation.
The equal-weight event-level estimate is larger in magnitude than the minute-level estimate because the two procedures weight the underlying variation differently. Event aggregation gives each event window the same weight, whereas the minute-level regression is influenced by within-event minute variation and by the number of usable observations. The magnitude difference therefore indicates meaningful cross-event heterogeneity and possible influence from a subset of settlement events. It should not be interpreted as an independent replication or as mechanically stronger evidence; the common negative sign across inference procedures is the more stable feature.
Note on R-squared: The model without month fixed effects explains a small share of minute-level variation, while month fixed effects raise the reported R 2 by absorbing persistent month-level differences. This is typical for high-frequency return differentials (basis changes), where noise and event-specific conditions dominate short-term fluctuations. Our focus is on the economic magnitude and statistical significance of the settlement-window divergence effect, not the overall explanatory power of the model.

5.2. Event-Window Comparability Diagnostics

The validity of our reduced-form comparison depends on local event-window comparability: absent settlement-specific pressure, expiry and placebo windows in the same monthly setting should display similar basis behavior around the same clock-time structure. We assess this using descriptive event-time plots and pre-window tests.

5.2.1. Visual Diagnostic

Figure 1 plots average basis for expiry days versus placebo days across a broader event-time window centered around settlement. In the rebuilt 24-event panel, the raw average expiry-placebo difference over the pre-window [ 60 , 10 ] is 7.93 basis points, and the exact paired-month sign-flip test for that mean pre-window difference yields p = 0.296 . During the post-settlement display window, the average treatment-placebo difference is 2.58 basis points, with the largest raw gap equal to 9.38 basis points at t = 14 . These patterns are consistent with local settlement-window divergence, but they should be interpreted as event-window diagnostics rather than as a conventional long-panel parallel-trends validation.
The positive average divergence in Figure 1 is not contradictory to the negative interaction coefficient in Table 4. The figure reports an unconditional average across months, while the regression coefficient β 3 is a conditional within-month marginal effect that accounts for incentive intensity and month fixed effects. The two objects therefore answer different questions.

5.2.2. Formal Statistical Test

We formally assess pre-window comparability using exact paired-month sign-flip tests and clustered event-time pre-bin tests. For the average pre-window difference over [ 60 , 10 ] , the exact paired-month test yields p = 0.296 across 12 matched month pairs. The corresponding immediate pre-window test over [ 10 , 2 ] is weaker ( p = 0.110 ), and the clustered event-time regression rejects equality over the immediate pre-window ( p = 0.00205 ). These results indicate that broad pre-window comparability is not rejected, but local differences immediately before settlement are visible in the formal event- time specification.
For this reason, the figures and tests should be read as local diagnostics rather than as proof of a clean untreated counterfactual path. They support the reduced-form comparison only in a limited sense and do not eliminate all settlement-specific confounding.

5.3. Formal Event-Time Evidence

Figure 2 plots formal event-time interaction coefficients for expiry relative to placebo windows. Event time is centered at t = 0 , defined as the settlement midpoint, and the regression uses the symmetric window t [ 60 , + 60 ] , month fixed effects, t = 1 as the reference bin, and standard errors clustered by event date. The line reports the expiry-minus-placebo coefficient at each event minute and the shaded band reports the corresponding 95% confidence interval.
The coefficients vary around settlement and the largest absolute post-settlement estimate is 14.3 basis points at t = 6 . However, the confidence intervals are wide and the immediate pre-period joint test rejects equality, so the figure is not a clean parallel-trends validation and does not identify the mechanism or trader behavior behind the observed pattern.
Table 6 summarizes the same formal regression shown in Figure 2. It uses 24 event dates (12 expiry and 12 placebo windows). The immediate pre-period joint Wald test over 10 t 2 rejects equality in the local pre-window ( p = 0.00198 ), reinforcing the need to interpret the dynamic estimates as a reduced-form event-window diagnostic rather than as clean causal evidence.
Hypothesis status: The event-time evidence is suggestive for H3 because settlement-window divergence is visible around the event, but the immediate pre-window differences mean it should be interpreted cautiously rather than as clean evidence of post-settlement reversion.

5.4. Monthly Heterogeneity and Directional Alignment

Table 7 reports settlement outcomes for all 12 expiry events in our sample, providing complete transparency regarding effect heterogeneity. The monthly classification is rule-based. Predicted directional pressure is upward (“Support”) when the settlement price ends above the relevant strike and downward (“Suppress”) when the settlement price ends below the strike; months with very small basis changes are labeled “Neutral” using the threshold described in Table 7.
The results show substantial directional consistency: 10 of 12 months exhibit basis movements in the theoretically predicted direction (83.3% success rate). The two near-neutral months (September 1.53 bps, November +0.39 bps) fall inside the | Delta | < 2 basis-point neutral band and are therefore retained but counted as non-matches. February now uses the observed February 15 placebo window, so all monthly placebo entries are computed from same-month control data. Months with the largest positive deltas (March +29.61 bps, December +38.12 bps, April +10.53 bps, October +5.93 bps) coincide with Bitcoin settling above strike prices, while the largest negative deltas (May 24.65 bps, July 14.09 bps, June 12.38 bps, August 6.30 bps) coincide with Bitcoin settling below strikes. This directional pattern is compatible with incentive-driven explanations, though it remains reduced-form evidence and alternative settlement-specific explanations cannot be fully ruled out.
Hypothesis status: The monthly directional evidence is suggestive for H2 in the sense that most settlement months move in the predicted sign, while the small number of months and the neutral-band classification require caution.

5.5. Formal Directional Test

The descriptive sign-alignment exercise in Table 7 yields 10 matches in 12 months (one-sided exact binomial p 0.019 ; two-sided p 0.039 ), but this result is sensitive to the small event count and the neutral-band classification. Table 8 applies stricter event-level and slope-based criteria: Table 7 uses expiry-minus-placebo monthly mean basis with a ±2 basis-point neutral band, whereas Table 8 uses directionalized event-level scores and above-versus-below-strike slope tests. Under these stricter tests, the evidence is mixed; the treatment-incentive slopes are negative on both sides of the strike, but the slope difference is not statistically significant. We therefore interpret the monthly sign pattern as suggestive rather than as evidence of a clean directional-asymmetry mechanism.

5.6. Descriptive Microstructure Evidence

To address mechanism-related concerns, we examine whether the main settlement-window basis result is accompanied by visible changes in coarse one-minute microstructure proxies. Using one-minute OHLCV proxy measures, we do not find a statistically sharp settlement-window increase in abnormal volatility relative to placebo windows ( β = 0.322 , p = 0.093 ), abnormal Coinbase log volume ( β = 0.045 , p = 0.788 ), or immediate signed return differentials ( β = 0.027 , p = 0.736 ). These coarse proxy measures do not identify order-book mechanism or trader behavior, but they suggest that the main basis result is not simply mirrored by an obvious aggregate spike in minute-level volume or volatility.
These measures are not substitutes for order-book imbalance, aggressive trade direction, signed trade volume, or depth-depletion tests. Those mechanism tests would require synchronized historical order-book and trade-direction data for Coinbase and Binance. Appendix A provides supporting descriptive statistics and the August case study. The appendix evidence should be interpreted as illustrative only, not as causal evidence of a specific microstructure mechanism.

5.7. Robustness Checks

The results in this section correspond to the reference monthly-actual specification described in Section 4.8 and represent the paper’s primary confirmatory evidence.
We conduct several robustness tests to assess the stability of our findings and examine alternative explanations.

5.7.1. Test 1: Placebo Date Falsification

We re-run Model 2 using data from the 15th of each month (identical 150-min spans with inclusive one-minute endpoints, but no expiry event). If our results reflected time-of-day effects, day-of-month patterns, or other spurious factors unrelated to settlement, we should observe similar coefficients on placebo dates. Result: In the main specification, the interaction coefficient is approximately β 3 = 0.000674 . When we re-estimate the same regression on non-expiry placebo dates, the interaction coefficient becomes β 3 = + 0.001547 ( p 0.20 ). The placebo estimate is therefore not statistically significant, and its sign does not align with the main effect. This pattern is consistent with the placebo estimate being close to zero on non-settlement dates, although it does not eliminate the possibility that other unobserved differences between expiry and placebo days remain relevant. As an additional falsification check, we shift the settlement window forward by 60 min and re-estimate the same specification. The shifted-window specification yields β 3 = 0.000750 and remains statistically significant. This estimate is suggestive of an effect concentrated around the settlement period, but it does not fully isolate timing and cannot fully rule out adjacent intraday influences. Taken together, these placebo and falsification results are consistent with a settlement-specific interpretation of the main finding, although alternative explanations cannot be fully excluded and the tests do not identify mechanism.

5.7.2. Test 2: Exchange Heterogeneity and Alternative Comparisons

Coinbase and Binance differ along several dimensions that matter for interpretation: clientele, geography, fee tiers, market-maker composition, USD versus USDT denomination, and exchange microstructure. The baseline design is therefore informative about settlement-window divergence between a benchmark-relevant venue and a major excluded venue, but it is not a fully exchange-invariant causal estimate.
Table 9 reports the available multi-exchange comparisons. The Binance baseline remains negative and statistically significant, and the excluded composite of Binance and OKX is also negative and significant. Coinbase–OKX alone is weak, and Coinbase–Bitstamp, an included-versus-included comparison, is near zero and insignificant. Coinbase–Kraken is negative and statistically significant, but its magnitude is much smaller than the Binance baseline. This pattern is consistent with the main result not being solely a Coinbase–Binance artifact, but it also shows that the effect is not uniform across every feasible exchange pair.
The Kraken comparison is useful because it is USD-denominated and therefore helps assess whether the baseline result is simply a Coinbase–USDT comparison artifact. However, Kraken can be benchmark-relevant in Bitcoin reference-rate designs, so it is not a clean non-constituent placebo in the same sense as Binance or OKX. The included-versus-included comparisons are also heterogeneous: Coinbase–Bitstamp is near zero and statistically insignificant, whereas Coinbase–Kraken is negative and statistically significant, despite both comparison venues belonging to the benchmark family. This mixed result prevents a clean interpretation based only on benchmark inclusion. It may reflect exchange-specific liquidity, benchmark weighting, clientele, or price-discovery differences, but the present data cannot separate these channels. We therefore interpret Table 9 as evidence of venue heterogeneity and as a check against a purely USDT-denomination explanation, not as definitive proof of an included-versus-excluded exchange mechanism.

5.7.3. Test 3: Volume Controls

We augment the specification with aligned minute-level traded volume on Coinbase ( V o l Target ) and Binance ( V o l Control ) as additional controls. If the basis divergence simply reflected volume imbalances or liquidity differences between exchanges, controlling for these exchange-specific volume measures should attenuate the coefficient. Result: The coefficient remains negative and significant ( β 3 = 0.000238 , SE = 0.000071 , p = 0.001 ). This is suggestive of volume differences alone not fully accounting for the observed pattern, although the controls do not capture every dimension of liquidity or trading conditions and therefore cannot fully rule out related explanations.

5.7.4. Test 4: Volatility Subsample Analysis

We split the sample into high-volatility and low-volatility periods using a trailing 30-day realized-volatility measure for Bitcoin, computed as the non-annualized standard deviation of daily log returns over the prior 30 days, with observations above the sample median classified as high-volatility and observations at or below the median classified as low-volatility. This partition is exploratory and is used to study heterogeneity rather than as a pre-specified baseline specification; if settlement-related price pressure is easier during calm markets (when less volume is required to move prices), the effect should be stronger in low-volatility periods. Result: The effect is present in both subsamples but stronger during low-volatility periods ( p < 0.01 ) compared to high-volatility periods ( p = 0.08 ). This pattern is suggestive of stronger effects during lower-volatility periods, which may be consistent with reduced background noise, though that interpretation is not directly tested here and the subsample split remains exploratory.

5.7.5. Test 5: Strike Selection Method

Our primary results rely on identifying economically relevant contract strikes using actual Polymarket contract data. To assess this approach, we compare results across three strike selection methods: (1) refined nearest strike with volume filter (primary specification), (2) refined nearest strike without volume filter, and (3) an inferred round-number proxy.
Table 10 presents the results. The two audited Polymarket strike-selection methods produce negative and highly statistically significant coefficients. By contrast, the inferred round-number proxy is positive and statistically insignificant. This pattern strengthens the view that the result depends on identifying the economically active contract strike rather than assigning a coarse price-level proxy, while also showing that strike measurement is an important source of design sensitivity.

5.7.6. Test 6: Round-Number Placebo Strikes

A related concern is that all selected contract strikes are salient round-thousand Bitcoin price levels. If the result were driven by generic round-number clustering rather than prediction-market settlement incentives, then proximity to nearby inactive round-number levels should reproduce the main interaction pattern.
To test this, we construct two placebo incentive measures. For each month, we replace the selected active Polymarket strike with the nearest inactive $1000-grid round number and with the nearest inactive $5000-grid round number, excluding the active contract strike itself. We then re-estimate the same month-fixed-effect interaction regression. Table 11 shows that the active Polymarket strike remains strongly significant, while neither inactive round-number placebo measure is statistically significant.
This check weakens a simple generic round-number explanation. It does not rule out every price-level or derivatives-related channel, because round numbers can also matter through other markets, but the paper’s main strike-proximity result is not reproduced by nearby inactive round-number levels in the paper-facing sample. An even stronger placebo would compare active prediction-market strikes with inactive prediction-market strikes from the same expiry-date candidate set. We do not report that test here because the inactive candidate files require additional audit work to distinguish inactive, duplicated, and institutionally distinct contract definitions. The inactive round-number placebo is therefore the paper-facing falsification test, while inactive-contract strike placebos remain a useful extension for future work.

5.8. Additional Specification Sensitivity Checks

We evaluate 243 alternative specifications that vary the settlement-window length, strike cutoff, contract-volume threshold, smoothing parameter, and Newey-West lag length. Because all specifications are generated from the same 12 settlement events and vary overlapping implementation choices, this grid is an exploratory sensitivity analysis rather than a set of independent or pre-registered confirmatory tests. It should be interpreted as a map of specification sensitivity.
Across the grid, the interaction coefficient remains negative, with estimates ranging from 0.000884  to 0.000517 . The tighter 3% strike cutoff produces the greatest attenuation; changes in the volume threshold and smoothing parameter have more modest effects, while the HAC lag primarily changes the estimated standard error. These results indicate that the sign and approximate magnitude are not uniquely determined by one implementation choice, but they do not add independent event-level evidence or identify the underlying mechanism.
Table 12 reports representative grid specifications; the full grid is summarized as sensitivity evidence and not as 243 independent confirmations.

5.9. Alternative Explanations

Several alternative explanations remain relevant. First, liquidity fragmentation or exchange-specific participant composition could produce Coinbase–Binance divergence even absent settlement incentives. The multi-exchange checks above partially address this concern because the excluded composite remains negative and significant, the USD-denominated Kraken comparison has the same sign at smaller magnitude, and the included Bitstamp comparison is near zero; however, the weak Coinbase–OKX result shows that the pattern is not uniform across all excluded venues.
Second, generic round-number effects could matter because Bitcoin order flow and derivatives positioning may cluster around salient price levels. The inactive round-number placebo test weakens the simplest version of this explanation: nearby inactive $1000 and $5000 round-number levels do not reproduce the active-strike result. Still, this test cannot rule out all price-level clustering channels, especially if prediction-market strikes, spot-market limit orders, and derivatives strikes concentrate at similar levels.
Third, month-end CME or Deribit options expiries are a serious competing channel. CME publishes volume and open-interest resources and Deribit provides market-data APIs, while vendors provide historical crypto derivatives and order-book datasets. However, the current repo does not contain a clean historical panel of CME/Deribit strike-level open interest, strike concentration, max-pain, or expiry-by-minute exposure data for the paper’s sample. The analysis therefore cannot fully separate prediction-market settlement incentives from broader month-end crypto derivatives pressure. The empirical design is most informative about cross-exchange divergence around the settlement windows studied here; a fuller design would jointly model prediction-market strikes and conventional crypto-derivatives exposures.
Fourth, information shocks and stablecoin denomination effects could affect the basis. The placebo timing checks, Tether adjustment, and exchange comparisons reduce but do not eliminate these concerns. For this reason, the evidence is best interpreted as a structured reduced-form pattern consistent with settlement-related incentives, rather than as proof that a specific trader or mechanism caused the observed divergence.

5.10. What the Evidence Does and Does Not Establish

What the evidence suggests. The analysis documents several empirical regularities. First, settlement-window divergence across exchanges is present in the data. Second, this divergence is stronger when the underlying price is closer to economically relevant contract strikes. Third, the pattern is more pronounced on oracle-constituent exchanges relative to excluded exchanges. Fourth, the directional pattern is broadly consistent with payoff-relevant incentives.
What the evidence does not establish conclusively. The analysis does not directly identify the trader or traders generating the observed price pressure. It does not establish whether the behavior reflects deliberate manipulation, hedging, liquidity provision, or other strategic settlement-related activity. It also does not isolate the exact microstructure mechanism through which the divergence emerges, because true mechanism identification would require synchronized historical order-book and trade-direction data rather than one-minute OHLCV proxies. Finally, it remains an open question whether the findings generalize across all platforms, assets, or settlement designs.

6. Discussion and Implications

6.1. Economic Interpretation: The Fee Barrier Puzzle

A natural question raised by the results is why observed settlement-window spreads are not fully arbitraged away. One possible explanation is that settlement periods combine adverse-selection risk, short execution windows, and uncertainty about whether observed price pressure will revert. Relatedly, traders with large derivative exposure may face incentives that are not symmetric with those of spot arbitrageurs. These considerations are illustrative rather than identified by the empirical design.

6.2. Preliminary Design Considerations

Given the reduced-form nature of the evidence, the discussion below should be interpreted as preliminary and illustrative rather than as prescriptive policy recommendations. The paper does not identify trader intent or mechanism, so any formal evaluation of market-design responses would require structural modeling or stronger causal identification. Within that limitation, the findings may have implications for oracle construction and settlement design in financially settled prediction markets. The oracle-economics literature emphasizes that oracle design faces trade-offs among scalability, decentralization, and truthfulness, and that prediction markets are one class of applications that rely directly on accurate external settlement data [22]. One possible design consideration is settlement-window structure: less predictable timing or more robust aggregation across venues could potentially make short-lived settlement-specific price pressure more difficult to sustain. A second consideration is transparency and surveillance around settlement periods, especially when open interest becomes large relative to the liquidity of the venues used in the settlement index. More generally, the results highlight that oracle design may become more sensitive as derivative exposures grow relative to the underlying spot markets they reference. Other potential design considerations include open-interest disclosure, circuit-breaker style safeguards, broader exchange inclusion, alternative settlement references, position limits, and fee adjustments near settlement thresholds, though evaluating these options is beyond the scope of the present paper and would require a more fully specified structural framework.

6.3. Oracle Extractable Value and Broader Dynamics

The discussion in this section is conjectural and intended to motivate potential directions for future research. The empirical patterns documented here raise the possibility that Oracle Extractable Value (OEV) may be relevant in financially settled prediction markets, but the current analysis does not identify such dynamics directly. One potential concern is that the relationship between derivatives exposure and the liquidity of the underlying settlement venues may matter for oracle design, though that relationship is not directly measured here. If derivative exposures become large relative to the underlying spot markets they reference, oracle construction may become more sensitive to localized settlement-window price pressure. This conjectural OEV framing is related to peer-reviewed evidence that blockchain infrastructure can create extraction opportunities and welfare-relevant allocation frictions [23]. The current analysis does not directly evaluate these broader dynamics, and they remain an open question for future work.

7. Conclusions

This paper documents settlement-window price divergence in cryptocurrency prediction markets that is suggestive of settlement-related effects. Using high-frequency data, actual contract-level strike identification from Polymarket and Kalshi, and a difference-in-differences design, we find that constituent exchange prices diverge from non-constituent exchange benchmarks during settlement windows when binary option strikes are nearby. Using actual contract-level data from prediction markets, we document that a one standard deviation increase in strike proximity is associated with a 6.7 basis point price deviation during settlement windows ( p < 0.001 in the baseline HAC specification). At high incentive levels (2+ standard deviations), the estimated magnitudes reach roughly 14–20 basis points, consistent with observed patterns. The monthly directional evidence and strike-selection robustness checks are compatible with incentive-driven explanations, while also indicating that strike measurement precision matters for detecting the pattern. These findings should nevertheless be interpreted with appropriate caution. The empirical design is reduced-form, does not directly observe trader identity or order-level mechanism, and does not fully isolate deliberate settlement-related trading from other settlement-specific dynamics. Because constituent and comparison exchanges differ in clientele, market structure, and denomination, the results should be interpreted as reduced-form evidence rather than a fully exchange-invariant causal estimate. Because the effective independent variation is closer to 12 settlement events than to 3624 independent minutes, the very small baseline p-values should be interpreted cautiously even though the coefficient remains negative under more conservative inference checks. Our results may have implications for market design. As prediction markets mature and achieve institutional scale, the assumption of “oracle independence”—that settlement data sources are unaffected by the derivatives they settle—may become more difficult to maintain. Markets used to settle derivatives larger than themselves may be more exposed to settlement-related price pressure within the setting studied here. For regulators and platform designers, these findings may point to settlement-window surveillance, disclosure around aggregate positioning, and oracle-construction choices as areas for further consideration. For researchers, the results also motivate further work on Oracle Extractable Value and related settlement-design questions, but only in a conjectural and forward-looking sense.
More broadly, our findings illuminate a fundamental tension in decentralized finance: the desire for permissionless, automated settlement conflicts with the need for robust settlement oracles. Addressing this tension may be important for prediction markets to achieve their potential as information aggregation mechanisms without unduly affecting the underlying markets they reference.
Future research should examine: (1) optimal oracle design under adversarial conditions with game-theoretic analysis, (2) cross-platform analysis as more prediction market data becomes available, (3) joint modeling of prediction-market strikes with CME/Deribit strike-level open interest and max-pain measures, (4) order-book and signed-order-flow mechanism tests using synchronized historical L2 data, (5) the interaction between OEV and traditional MEV in blockchain-based settlements, and (6) whether analogous settlement-related price pressure appears in other financially-settled prediction markets beyond cryptocurrency (e.g., weather derivatives, sports betting, political event contracts).

Author Contributions

Conceptualization, S.J. and Z.Z.; methodology, S.J. and Z.Z.; software, S.J.; validation, Z.Z.; formal analysis, S.J.; investigation, S.J.; data curation, S.J.; writing—original draft preparation, S.J.; writing—review and editing, S.J. and Z.Z.; visualization, S.J.; supervision, Z.Z.; project administration, Z.Z. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

The data and code supporting the findings of this study are available from the corresponding author upon reasonable request.

Acknowledgments

During the preparation of this manuscript, the authors used OpenAI ChatGPT and Codex for language editing, LATEX formatting, code assistance, and revision-consistency checks. The authors reviewed and edited the output and take full responsibility for the content.

Conflicts of Interest

The authors declare no conflicts of interest.

Appendix A. Descriptive Microstructure Evidence

Table A1. Descriptive Microstructure Proxy Tests Around Settlement.
Table A1. Descriptive Microstructure Proxy Tests Around Settlement.
Proxy OutcomeWindowEstimateStd. Errorp-Value
Abnormal volatility [ 0 , 60 ] 0.322 0.1750.093
Abnormal Coinbase log volume [ 0 , 60 ] 0.045 0.1620.788
Abnormal signed return differential [ 0 , 15 ] 0.0270.0800.736
Notes: The table reports settlement-versus-placebo regressions for coarse one-minute OHLCV proxy measures exported by src/analysis/microstructure_proxy_runner.py. Estimates are standardized abnormal proxy outcomes, not order-book or signed-trade-flow measures. They should not be interpreted as identified evidence of mechanism.

Case Study: 31 August 2025

Table A1 in Appendix A reports descriptive OHLCV proxy tests, while the 31 August 2025 event provides a detailed microstructure view of settlement-window dynamics. This episode is illustrative and should not be interpreted as definitive evidence of a specific microstructure mechanism. It is selected to visualize within-window timing, temporary cross-exchange de-synchronization, and rapid correlation recovery while Bitcoin traded near the selected strike, not because it produced the largest estimated settlement effect. Bitcoin spent the settlement window trading near the $108,000 strike price (refined identification). The decoupling delta was 6.30 basis points, indicating Coinbase suppression relative to Binance—a pattern compatible with price moves favorable to a “No” outcome (Bitcoin below $108 k).
Figure A1 shows the evolution of two key metrics during the critical settlement period: the price basis between Coinbase and Binance (left axis, red line) and the 10-min rolling correlation of returns between the two exchanges (right axis, blue line).
Several features of this event are descriptive and relevant for interpretation:
  • Correlation breakdown: The 10-min rolling correlation collapsed from positive (+0.04 at 20:05) to significantly negative ( 0.23 at 20:11–20:12), indicating that Coinbase and Binance were moving in opposite directions—a highly unusual pattern for two markets trading the same asset.
  • Rapid reversion: The negative correlation began recovering immediately at 20:13 ( 0.19 ), a pattern suggestive of temporary de-linking rather than a persistent information shock.
  • Basis narrowing during crisis: Paradoxically, the USD basis (red line) narrowed from −$162.63 at 20:07 to −$29.05 at 20:13, meaning Coinbase was recovering toward Binance prices even as their returns became negatively correlated. This pattern may reflect offsetting forces, including settlement-related pressure on Coinbase during the window and reconvergence afterward, but it is not uniquely diagnostic of any single mechanism.
Figure A1. Coinbase-Binance basis and 10-min rolling return correlation on 31 August 2025. The joint movement is suggestive of temporary de-synchronization during settlement, though alternative explanations cannot be ruled out. Notes: The red line plots the USD-denominated Coinbase-Binance basis on the left axis. The blue line plots the 10-min rolling return correlation between the two exchanges on the right axis.
Figure A1. Coinbase-Binance basis and 10-min rolling return correlation on 31 August 2025. The joint movement is suggestive of temporary de-synchronization during settlement, though alternative explanations cannot be ruled out. Notes: The red line plots the USD-denominated Coinbase-Binance basis on the left axis. The blue line plots the 10-min rolling return correlation between the two exchanges on the right axis.
Fintech 05 00067 g0a1
This pattern—suppression during settlement followed by immediate mean reversion with negative return correlation—is compatible with settlement-related explanations, but alternative interpretations—such as information-driven trading or transient liquidity imbalances—cannot be ruled out. If the Coinbase price drop primarily reflected broad selling pressure or negative information about Bitcoin, we would expect:
  • Positive correlation (both markets declining together in response to information)
  • Persistent basis divergence (information would be incorporated into prices permanently)
  • No systematic reversion pattern
Instead, we observe negative correlation, temporary divergence, and systematic timing around the settlement window. This descriptive evidence is compatible with settlement-related explanations, though it does not by itself identify intent or establish a unique mechanism. As such, this case study is best read as an illustrative example of the broader empirical patterns.

References

  1. Wolfers, J.; Zitzewitz, E. Prediction markets. J. Econ. Perspect. 2004, 18, 107–126. [Google Scholar] [CrossRef]
  2. Hanson, R. Logarithmic market scoring rules for modular combinatorial information aggregation. J. Predict. Mark. 2007, 1, 3–15. [Google Scholar]
  3. Berg, J.; Nelson, F.; Rietz, T. Prediction market accuracy in the long run. Int. J. Forecast. 2008, 24, 285–300. [Google Scholar] [CrossRef]
  4. Koppl, R.; Devereaux, A.; Herrmann, J. Event contracts and information aggregation: Evidence from 2020–2023. J. Financ. Mark. 2023, 63, 100834. [Google Scholar]
  5. Shleifer, A.; Vishny, R.W. The limits of arbitrage. J. Financ. 1997, 52, 35–55. [Google Scholar] [CrossRef]
  6. Evans, M.D. Forex trading and the WMR fix. J. Bank. Financ. 2018, 96, 156–173. [Google Scholar]
  7. Ng, H.; Peng, L.; Tao, Y.; Zhou, D. Price Discovery and Trading in Modern Prediction Markets. SSRN 2025, 5331995. [Google Scholar] [CrossRef]
  8. U.S. Commodity Futures Trading Commission. CFTC Orders Barclays to Pay $200 Million Penalty for Attempted Manipulation of and False Reporting Concerning LIBOR and Euribor Benchmark Interest Rates. Press Release No. 6289-12, 27 June 2012. Available online: https://www.cftc.gov/PressRoom/PressReleases/6289-12 (accessed on 18 July 2026).
  9. Duffie, D.; Dworczak, P. Robust benchmark design. J. Financ. Econ. 2021, 142, 775–802. [Google Scholar] [CrossRef]
  10. Ni, S.X.; Pearson, N.D.; Poteshman, A.M. Stock price clustering on option expiration dates. J. Financ. Econ. 2005, 78, 49–87. [Google Scholar] [CrossRef]
  11. Pearson, N.D.; Poteshman, A.M.; White, J.S. Does option trading convey stock price information? J. Financ. Quant. Anal. 2015, 45, 1289–1308. [Google Scholar]
  12. Griffin, J.M.; Shams, A. Is Bitcoin really untethered? J. Financ. 2020, 75, 1913–1964. [Google Scholar] [CrossRef]
  13. Kosse, A.; Glowka, M.; Mattei, I.; Rice, T. Will the real stablecoin please stand up? BIS Pap. 2023, 141, 1–29. [Google Scholar]
  14. Hui, C.H.; Wong, A.; Lo, C.F. Stablecoin price dynamics under a peg-stabilising mechanism. J. Int. Money Financ. 2025, 152, 103280. [Google Scholar] [CrossRef]
  15. Carapella, F.; Lubis, A.; Vardoulakis, A. Stablecoins in 2025: Developments and financial stability implications. In FEDS Notes; Board of Governors of the Federal Reserve System: Washington, DC, USA, 8 April 2026. [Google Scholar] [CrossRef]
  16. Makarov, I.; Schoar, A. Trading and arbitrage in cryptocurrency markets. J. Financ. Econ. 2020, 135, 293–319. [Google Scholar] [CrossRef]
  17. Othman, A.; Pennock, D.M.; Reeves, D.M.; Sandholm, T. A practical liquidity-sensitive automated market maker. ACM Trans. Econ. Comput. 2013, 1, 377–386. [Google Scholar] [CrossRef]
  18. León, J.C.; Lehar, A. What data have told us about decentralized finance. J. Corp. Financ. 2026, 96, 102916. [Google Scholar] [CrossRef]
  19. Rahman, N.; Al-Chami, J.; Clark, J. SoK: Market Microstructure for Decentralized Prediction Markets (DePMs). arXiv 2025, arXiv:2510.15612. [Google Scholar]
  20. Daian, P.; Goldfeder, S.; Kell, T.; Li, Y.; Zhao, X.; Bentov, I.; Breidenbach, L.; Juels, A. Flash Boys 2.0: Frontrunning, Transaction Reordering, and Consensus Instability in Decentralized Exchanges. In Proceedings of the 2020 IEEE Symposium on Security and Privacy (SP), San Francisco, CA, USA, 18–21 May 2020; IEEE: Piscataway, NJ, USA, 2020; pp. 910–927. [Google Scholar] [CrossRef]
  21. Zhang, L.; Chen, Y.; Pennock, D.M. Oracle manipulation in decentralized prediction markets: Theory and evidence. Manag. Sci. 2024, 70, 1842–1859. [Google Scholar]
  22. Cong, L.W.; Fox, L.; Li, S.; Zhou, L. A primer on oracle economics. J. Corp. Financ. 2025, 94, 102800. [Google Scholar] [CrossRef]
  23. Capponi, A.; Jia, R.; Wang, K.Y. Maximal extractable value and allocative inefficiencies in public blockchains. J. Financ. Econ. 2025, 172, 104132. [Google Scholar] [CrossRef]
  24. Newey, W.K.; West, K.D. A simple, positive semi-definite, heteroskedasticity and autocorrelation consistent covariance matrix. Econometrica 1987, 55, 703–708. [Google Scholar] [CrossRef]
Figure 1. Average basis on expiry and placebo days around settlement. The blue line with square markers represents placebo days, and the brown line with circular markers represents expiry days. The gray shaded region indicates the 60-min settlement calculation window, the vertical red dotted line marks settlement initiation at t = 0 , and the horizontal gray dashed line marks zero basis. The raw averages are consistent with temporary settlement-window divergence, though they are not numerically equivalent to the regression coefficients. The figure is an unconditional event-time average, while the regression estimates a conditional within-month marginal effect controlling for incentive intensity and month fixed effects. Notes: The figure plots minute-level average basis (Coinbase minus Binance, Tether-adjusted) for expiry days and placebo days across a broader event-time window centered around settlement.
Figure 1. Average basis on expiry and placebo days around settlement. The blue line with square markers represents placebo days, and the brown line with circular markers represents expiry days. The gray shaded region indicates the 60-min settlement calculation window, the vertical red dotted line marks settlement initiation at t = 0 , and the horizontal gray dashed line marks zero basis. The raw averages are consistent with temporary settlement-window divergence, though they are not numerically equivalent to the regression coefficients. The figure is an unconditional event-time average, while the regression estimates a conditional within-month marginal effect controlling for incentive intensity and month fixed effects. Notes: The figure plots minute-level average basis (Coinbase minus Binance, Tether-adjusted) for expiry days and placebo days across a broader event-time window centered around settlement.
Fintech 05 00067 g001
Figure 2. Formal event-time regression around the settlement midpoint. The dark-red line reports expiry-minus-placebo interaction coefficients relative to t = 1 , and the pale-red band reports 95% confidence intervals. The vertical solid line marks the settlement midpoint, and the horizontal dashed line marks zero. Standard errors are clustered by event date.
Figure 2. Formal event-time regression around the settlement midpoint. The dark-red line reports expiry-minus-placebo interaction coefficients relative to t = 1 , and the pale-red band reports 95% confidence intervals. The vertical solid line marks the settlement midpoint, and the horizontal dashed line marks zero. Standard errors are clustered by event date.
Fintech 05 00067 g002
Table 1. Empirical Design Framework.
Table 1. Empirical Design Framework.
Design ElementImplementation in the Paper
OutcomeMinute-level Coinbase-minus-Binance basis after Tether adjustment: ln ( P Coinbase ) ln ( P Binance , adjusted )
TreatmentIndicator for minutes in a true monthly settlement-date window
Control windowSame clock-time window on same-month placebo dates used to preserve month and intraday structure
Identifying variationWithin-month comparison of expiry and placebo windows as strike proximity changes
Incentive metricStandardized inverse distance between Coinbase spot price and the selected economically live Polymarket strike
Fixed effectsMonth fixed effects in the baseline interaction regression
InferenceHAC(15) baseline standard errors, with event-date, month-clustered, event-level, and exact paired-month checks reported below
Main threatsExchange heterogeneity, generic round-number effects, month-end derivatives expiries, information shocks, and limited event-level power
Notes: The table summarizes the reduced-form design. The design links settlement timing, strike proximity, and cross-exchange basis movements, but it does not directly observe trader identity, intent, or order-level mechanism.
Table 2. Illustrative Observation from the Analysis Panel.
Table 2. Illustrative Observation from the Analysis Panel.
TimestampEvent TimeExpiryCoinbase PriceBinance PriceUSDT/USDAdj. Binance PriceBasis (bps)StrikeDistanceIncentive
2025-09-30 20:28+881114,400.04114,354.851.00015114,372.002.45114,000400.040.00216
Notes: This row illustrates the construction of the main variables for one minute in an expiry window. Event time is minutes relative to settlement-window start. Adjusted Binance price equals the Binance BTC/USDT price multiplied by USDT/USD. Basis is 10 , 000 × [ ln ( Coinbase ) ln ( Adjusted Binance ) ] . Incentive equals 1 / ( | Coinbase price Strike | + 50 ) before standardization.
Table 3. Contract-Level Strike Identification Around Settlement.
Table 3. Contract-Level Strike Identification Around Settlement.
MonthSettlement PricePolymarket StrikeKalshi StrikeDistance
(PM, %)
Volume (PM)
Feb 2025$84,267$85,000$84,250+0.87%$4.8 M
Mar 2025$83,018$80,000$82,750 3.64 % $6.1 M
Apr 2025$94,181$95,000$94,000+0.87%$5.2 M
May 2025$104,659$105,000$104,750+0.33%$7.3 M
Jun 2025$107,462$107,000$108,250 0.43 % $6.8 M
Jul 2025$117,277$117,000$115,750 0.24 % $8.1 M
Aug 2025$109,039$108,000$109,250 0.95 % $5.9 M
Sep 2025$113,881$114,000$113,500+0.10%$4.2 M
Oct 2025$109,503$110,000$109,500+0.45%$6.5 M
Nov 2025$91,352$92,000$91,000+0.71%$3.8 M
Dec 2025$87,477$88,000$87,750+0.60%$5.4 M
Jan 2026$77,618$80,000$77,750+3.07%$4.1 M
Mean 1.02%$5.7 M
Median 0.67%$5.7 M
Notes: Settlement Price is the average BTC/USD price on Coinbase during the 15:00–16:00 London settlement window. Polymarket Strike is the nearest strike to settlement price with volume > $100 k. Kalshi Strike provides a cross-platform comparison. Distance shows percentage difference between the Polymarket strike and settlement price. Volume is total monthly trading volume for the Polymarket contract.
Table 4. Settlement-Window Price Divergence: Refined Strike Identification.
Table 4. Settlement-Window Price Divergence: Refined Strike Identification.
Model 1Model 2
Intercept−0.000124 0.001578 ***
(0.000137)(0.000602)
Expiry Dummy ( D Expiry )0.000563 **0.000379 *
(0.000249)(0.000195)
Incentive (Standardized)−0.0000570.000544 ***
(0.000091)(0.000118)
Expiry × Incentive ( β 3 )−0.000674 ***
(0.000134)
Month Fixed EffectsNoYes
R-squared0.01830.1360
Observations36243624
*** p < 0.01 , ** p < 0.05 , * p < 0.10 . Newey-West HAC standard errors (15 lags) in parentheses. Observations are ordered chronologically within event windows before HAC estimation. Dependent variable: Basis (log Coinbase − log Binance, Tether-adjusted).
Table 5. Alternative inference procedures for the main monthly interaction estimate.
Table 5. Alternative inference procedures for the main monthly interaction estimate.
Inference Method β 3Std. Errorp-Value
Baseline minute-level HAC(15)−0.0006740.0001344.73412 × 10−7
Minute-level clustered by event date−0.0006740.0001696.82113 × 10−5
Minute-level clustered by month−0.0006740.0002010.00078256
Event-level aggregated (equal-weight events, HC1)−0.0016520.0005400.00224112
Exact paired-month permutation inference−0.000674 0.0256348
Table 6. Clustered Event-Time Regression Summary.
Table 6. Clustered Event-Time Regression Summary.
Regression FeatureValue
Event dates24 (12 expiry, 12 placebo)
Event-time window [ 60 , + 60 ] min
Reference bin t = 1
Covariance estimatorClustered by event date
Immediate pre-period joint test p = 0.00198
Peak absolute post-settlement coefficient14.3 bps at t = 6
Average absolute post-period coefficient4.8 bps
Notes: The table summarizes the event-time interaction regression exported by src/analysis/event_study.py. Coefficients are expiry-minus-placebo differentials in basis points and are estimated with month fixed effects.
Table 7. Monthly Settlement Outcomes: Complete Analysis.
Table 7. Monthly Settlement Outcomes: Complete Analysis.
MonthStrikePlacebo BasisExpiry BasisDelta (bps)DirectionMatch
Feb 2025$85,000+0.27+31.06+30.78SupportYes
Mar 2025$80,000 12.11 +17.50+29.61SupportYes
Apr 2025$95,000 7.50 +3.04+10.53SupportYes
May 2025$105,000+19.69 4.96 24.65 SuppressYes
Jun 2025$107,000 1.59 13.97 12.38 SuppressYes
Jul 2025$117,000+3.51 10.59 14.09 SuppressYes
Aug 2025$108,000+2.26 4.04 6.30 SuppressYes
Sep 2025$114,000+4.46+2.93 1.53 NeutralNo
Oct 2025$110,000 6.90 0.97 +5.93SupportYes
Nov 2025$92,000 2.81 2.42 +0.39NeutralNo
Dec 2025$88,000 15.11 +23.00+38.12SupportYes
Jan 2026$80,000+3.26+9.72+6.46SupportYes
Directional Success Rate: 10/12 (83.3%)
Notes: Strike prices are identified using the refined methodology (nearest to settlement price with volume > $100 k). Delta = Expiry Basis − Placebo Basis, measured in basis points. The predicted direction is “Support” if the settlement price is above the strike and “Suppress” if the settlement price is below the strike. Months are labeled “Neutral” when | Delta | < 2 basis points. A month is classified as a Direction Match if the sign of Delta aligns with the predicted direction. Neutral months are classified as non-matches.
Table 8. Formal directional-evidence tests on the paper-facing monthly panel.
Table 8. Formal directional-evidence tests on the paper-facing monthly panel.
TestEstimateStd. Errorp-Value
Settlement success rate0.416667NA7.74 × 10−1
Placebo success rate0.583333NA7.74 × 10−1
Matched settlement vs placebo0.400000NA7.54 × 10−1
Below-strike treatment slope−0.0004800.0001374.59 × 10−4
Above-vs-below slope difference−0.0003410.0002802.22 × 10−1
Above-strike treatment slope−0.0008210.0002262.85 × 10−4
Directionalized-basis treatment slope0.0000710.0001315.92 × 10−1
Notes: The settlement and placebo success rates are event-level directional classifications computed from the paper-facing panel. A formal directional success is counted when the event-level mean of directionalized basis, Y Basis × payoff sign , is positive. This rule differs from the descriptive monthly sign-alignment rate in Table 7, which uses expiry-minus-placebo mean basis and a ±2 basis-point neutral band.
Table 9. Main interaction estimates across alternative exchange comparisons. Binance, OKX, and the excluded composite are included-versus-excluded comparisons; Bitstamp is an included-versus-included comparison. Kraken is reported as a secondary USD-denominated comparison, not as a clean non-constituent placebo. The excluded composite is the equal-weight average of log prices on Binance and OKX after USD normalization.
Table 9. Main interaction estimates across alternative exchange comparisons. Binance, OKX, and the excluded composite are included-versus-excluded comparisons; Bitstamp is an included-versus-included comparison. Kraken is reported as a secondary USD-denominated comparison, not as a clean non-constituent placebo. The excluded composite is the equal-weight average of log prices on Binance and OKX after USD normalization.
Specification β 3Std. Errorp-ValueSign
Coinbase vs. Binance−0.0006740.0001344.73 × 10−7negative
Coinbase vs. OKX−0.0000130.0000090.154negative
Coinbase vs. Bitstamp−0.0000040.0000100.659negative
Coinbase vs. Kraken−0.0000500.0000140.000416negative
Coinbase vs. Excluded Composite−0.0003440.0000672.44 × 10−7negative
Table 10. Robustness to Strike Selection Method.
Table 10. Robustness to Strike Selection Method.
Selection MethodCoefficientStd. Errorp-ValueMean Distance
(1) Refined (nearest, vol > $100 k)−0.0006740.0001344.73 × 10−7$905 (1.0%)
(2) Refined (nearest, no vol filter) 0.000997 0.0001583.15 × 10−10$735 (0.8%)
(3) Inferred round-number proxy + 0.000191 0.0002430.431$4319 (4.8%)
Notes: All specifications include month fixed effects and Newey-West HAC(15) standard errors. The dependent variable is basis (log Coinbase − log Binance, Tether-adjusted). Row (1) uses Polymarket strikes nearest to settlement price with volume > $100 k. Row (2) removes the volume filter. Row (3) uses an inferred round-number proxy. Mean Distance is the average absolute distance from strike to settlement price.
Table 11. Round-Number Placebo Strike Test.
Table 11. Round-Number Placebo Strike Test.
Strike-Proximity Measure β 3Std. Errorp-Value
Actual active Polymarket strike−0.0006740.0001344.73 × 10−7
Nearest inactive $1000 round level−0.0003770.0002420.119
Nearest inactive $5000 round level−0.0011850.0006880.0851
Notes: All rows estimate the same month-fixed-effect interaction regression with HAC(15) standard errors on the paper-facing 3624-observation sample. The placebo rows replace active Polymarket strike proximity with proximity to the nearest inactive round-number level at the indicated grid size.
Table 12. Sensitivity of the main interaction coefficient to alternative specification choices. The table reports representative specifications; across the full set of 243 specifications, the coefficient remains negative and statistically significant.
Table 12. Sensitivity of the main interaction coefficient to alternative specification choices. The table reports representative specifications; across the full set of 243 specifications, the coefficient remains negative and statistically significant.
Specification β 3Std. Errorp-Value
Baseline−0.0006740.000134 4.7 × 10 7
Shorter window (90 min)−0.0007460.000167 7.9 × 10 6
Tighter strike cutoff (3%)−0.0005610.000136 3.5 × 10 5
Higher volume threshold ($200k)−0.0007280.000146 5.8 × 10 7
Lower ϵ (25)−0.0006020.000120 5.2 × 10 7
Higher ϵ (100)−0.0007720.000155 6.2 × 10 7
Shorter HAC lag (10)−0.0006740.000123 3.7 × 10 8
Longer HAC lag (20)−0.0006740.000143 2.4 × 10 6
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Joshi, S.; Zhou, Z. Evidence on Settlement-Window Price Divergence in Bitcoin Prediction Markets. FinTech 2026, 5, 67. https://doi.org/10.3390/fintech5030067

AMA Style

Joshi S, Zhou Z. Evidence on Settlement-Window Price Divergence in Bitcoin Prediction Markets. FinTech. 2026; 5(3):67. https://doi.org/10.3390/fintech5030067

Chicago/Turabian Style

Joshi, Sibin, and Zhaoxian Zhou. 2026. "Evidence on Settlement-Window Price Divergence in Bitcoin Prediction Markets" FinTech 5, no. 3: 67. https://doi.org/10.3390/fintech5030067

APA Style

Joshi, S., & Zhou, Z. (2026). Evidence on Settlement-Window Price Divergence in Bitcoin Prediction Markets. FinTech, 5(3), 67. https://doi.org/10.3390/fintech5030067

Article Metrics

Back to TopTop