Abstract
We formalise aircraft stand pre-assignment as a constrained Markov decision process whose hard operational predicates (stand validity, size compatibility, buffered temporal non-overlap and adjacency exclusion) define a feasible action mask; we prove that any scoring policy composed with this mask assigns only feasible stands, and we derive in closed form the expected conflict minutes (ECM) accumulated by a schedule when occupancy durations follow calibrated piecewise-linear predictive distributions. In a preregistered one-shot replay of all 516 test days at San Francisco International Airport, a greedy policy that consumes this risk model cuts model-based expected conflict under the calibrated occupancy-risk model from 1175.5 to 615.9 min per 100 reservations relative to the strongest rule baseline: a 48% reduction (Holm-corrected ), with zero constraint violations and no deferrals. Three further preregistered hypotheses are reported not confirmed: graph-embedding and blended rankers reverse significantly, and masked discrete conservative Q-learning beats behaviour cloning but not the rule.
Keywords:
stand assignment; gate assignment; constrained Markov decision process; action masking; quantile regression; uncertainty-aware allocation; expected conflict; conservative Q-learning; offline reinforcement learning; preregistration; airport operations; open data MSC:
90C40; 90B36; 90B06; 62G08; 68T05
1. Introduction
An aircraft arriving at a large airport must be given a parking stand before it lands, and the assignment must be held for as long as it actually stays. Two properties make this hard. The constraints are not soft: a stand holds one aircraft at a time, small stands cannot take wide-body aircraft, and a pad cannot be used simultaneously as a whole and through its lettered sub-positions. Violating any of these is an operational failure rather than a cost, so a policy that merely pays a penalty for infeasible choices is unusable. And the quantity that drives conflict, how long an aircraft will occupy the stand, is unknown at decision time: planners work with reserved durations that are systematically wrong, and the cost is paid later by the aircraft assigned to the same stand.
These properties pull in different directions. Guaranteeing feasibility invites a hard-constraint formulation solved by mixed-integer programming; consuming uncertainty invites a probabilistic forecaster and a policy that scores candidates by risk. Published work on gate and stand assignment has largely followed one route or the other. Exact and heuristic optimisation delivers feasibility and strong objective values on a fixed instance [1,2,3]. A growing body of learning-based literature applies reinforcement and imitation learning to the allocation decision itself [4,5], a strand documented by a recent systematic review [6]. Robust and predictive-prescriptive formulations bring uncertainty into the plan, through a robustness-optimising multi-commodity flow model [7] or predictive-prescriptive planning under uncertainty [8]. Rare is the method that does all four at once: guarantee hard feasibility by construction rather than by penalty; let a calibrated predictive distribution enter the decision rule that the policy optimises at decision time; define that risk as an analytical functional rather than a sampled surrogate; and report its evaluation under a protocol frozen before the test data were read. Table 1 places this paper against the closest work on those axes.
Table 1.
Positioning against representative learning-based and optimisation-based stand/gate assignment work, selected for proximity on the four axes below rather than by systematic search. Classifications reflect our reading of what each paper reports and are offered as a positioning aid, not as a survey finding. Among these representative studies we are not aware of one that satisfies all four axes simultaneously.
We take the composition route. A masking operator turns the four hard predicates into a state-dependent feasible action set, and any scoring function composed with it inherits feasibility (Proposition 1); the scorer is then free to be whatever consumes uncertainty best, in our case a greedy rule penalising each candidate stand by the predicted overlap it would create. Measuring what such a policy buys needs a risk quantity well defined on a schedule of predicted occupancies, and our main mathematical result supplies one: under piecewise-linear predictive duration distributions, the expected buffered overlap of two same-stand occupancies has an exact closed form, obtained by integrating a piecewise-quadratic integrand segment by segment (Theorem 1). We call the daily aggregate the expected conflict minutes per 100 reservations, ECM. It is not a convenience: on the public San Francisco International Airport (SFO) parking log the recorded end of a reservation equals its requested end, so reservation-based overlap is identically zero for every policy respecting the mask, and a degenerate label channel cannot rank policies. Every claim we derive from ECM is stated as an expected conflict under the calibrated occupancy-risk model, never as realised minutes saved.
1.1. Contributions
- Closed-form expected-conflict functional (Theorem 1): for piecewise-linear predictive duration CDFs with bounded linear tails, the expected buffered overlap of a same-stand pair integrates exactly, with no quadrature error. Two corollaries connect it to the decision problem: it is monotone in the effective separation of a pair and under stochastic-order shifts of either duration (Corollary 1), and it collapses to the deterministic buffered overlap, hence to zero on any mask-feasible schedule, when the predictive distributions degenerate to point masses (Corollary 2). A Monte-Carlo gate confirms the implementation on 30 of 30 validation days, at an aggregate relative difference of 0.00159.
- Feasibility by construction (Proposition 1): the masked policy class assigns only stands satisfying all four predicates, for any scorer; across 516 one-shot test days and every policy evaluated, the measured violation count is 0.
- Preregistered headline result: a greedy policy consuming a overlap-risk term attains 615.9 ECM against 1175.5 for the strongest rule baseline B1, a 48% reduction (day-paired , 95% CI , Holm-corrected ), at a defer rate of 0 and with churn below B1’s. B1 never sees the occupancy model, so this is a like-for-like decision-quality contrast; a second, larger contrast against an uncertainty-blind twin (, Holm ) is a mechanism ablation isolating the information channel.
- Calibrated occupancy-duration model: a LightGBM quantile model with conformal interval calibration reaches test pinball loss 205.08 min against 243.79 for a point- model and 209.50 for the historical expanding-quantile baseline, at 80% interval coverage 0.782 (day-paired against the historical baseline min/day, Holm ).
- Masked discrete conservative Q-learning, motivated and measured: we adapt the CQL support-restriction argument to the masked simplex, obtaining a conservative lower bound for the policy-evaluation operator evaluated at its own softmax (Proposition 2); this motivates the conservatism mechanism of a deployed control-CQL learner, which is not itself the object of the proposition. The learner beats the behaviour-cloning ensemble on the composite reward by per day (Holm ) and does not reach the rule baseline (), so under the preregistered conjunction rule that hypothesis is not confirmed and we report it as such.
- Measured boundary conditions, and an optimality reference: gradient boosting beats every attempted upgrade (deep, graph-attention, graph-hybrid and blended rankers) on 516 replayed days with paired statistics, two of them preregistered to help and significantly reversed on test; a per-day mixed-integer program solved to proven optimality on 20 sampled test days puts the historical plan’s headroom at 66.4 objective units per day (62.2%).
- Reproducible, preregistered pipeline: hypotheses, endpoints, comparator sets, decision rule and the ECM definition were frozen before any development iteration; the frozen test split was read for the preregistered evaluation exactly once, and every other test-set contact, before and after that read, is enumerated in Section 4.7.
1.2. Scope of the Claims
Three of the four preregistered hypotheses are reported as not confirmed, with numbers, in Section 5: the graph-embedding hybrid ranker (H-A), the score blend (H-B) and the offline-RL comparison against the rule (H-D). Only the uncertainty-aware allocation hypothesis (H-C) is confirmed. That confirmed claim is model-based by construction: it says that consuming calibrated occupancy uncertainty reduces expected conflict under the calibrated occupancy-risk model of Theorem 1, on one airport’s log, under the pairwise conventions of Section 2. It does not say that realised conflict minutes fall, because on this data snapshot realised conflict cannot be measured at all. We lead with a comparison against a rule that never consumes the occupancy model, which makes the result more than a policy scoring well on its own surrogate. The sign of the advantage is established beyond this single calibrated model: Section 5.8 re-runs the comparison under a second, independently conformalised duration model, both as the policy’s input and as the yardstick that judges it, and under correlated durations induced by a day-level common shock. The ranking holds throughout, and the margin widens under the alternative model.
2. Mathematical Formulation
2.1. Decision Problem
We model one operating day as an episode of a constrained Markov decision process , in which F is a state-dependent admissible-action map rather than a family of expected-cost constraints (Figure 1). Let be the stand inventory and . Requests are processed in order of reserved start, ties broken by event identifier; the data carry no booking timestamps, so this is the decision order. A state pairs the request under decision with the committed schedule , , and exogenous context (weather, load, calendar); a request carries a deterministic reserved start , a reserved duration, an aircraft model, an operating company and a reservation-chain label , the pair (source type, source identifier) grouping the recurring blocks of one standing reservation. The schedule holds both the boundary occupancies inherited from the previous day and the decisions already taken. Transitions are deterministic: yields , and ∅ leaves . The reward is a weighted sum of four costs,
with weights locked in the project configuration. The composite endpoint of Section 5 takes and excludes the tie-break-tier size-mismatch term, an exclusion fixed in the preregistration; that term is reported descriptively. Where the classical constrained MDP of Altman [9] restricts policies through expected-cost inequalities, our constraints admit no expectation: a double-booking is a failure on the realised path, not on average. Enforcing them pathwise through F is strictly stronger and is what makes Proposition 1 exact rather than a bound; action masking of this kind is standard in policy-gradient practice [10], but here it also carries the feasibility guarantee.
Figure 1.
Framework. Predicates C1–C4 define the feasible action mask; a calibrated quantile model of occupancy duration feeds the closed-form ECM functional (Theorem 1); three policies are replayed over 516 preregistered one-shot test days; the blind twin is dashed because it bypasses the uncertainty channel. Day-mean model-based expected conflict under the calibrated occupancy-risk model per 100 reservations: 615.9, B1 1175.5, blind twin 7142.6. Prop., Proposition; Thm, Theorem; BC, behaviour cloning; P10/P50/P90, 0.1/0.5/0.9 quantile forecasts; q90, 0.9 quantile.
2.2. Hard-Constraint Predicates
Write for the reserved interval of i and min for the buffer.
Definition 1
(Constraint predicates). For a request i, a stand and a committed schedule H:
The feasible action set in state is
Each predicate carries a measured design decision. C1 reads the validity window, not present status, because 34,195 historical events sit on stands inactive today (only 48 predate a stand’s valid_from and 12 postdate its valid_to). C2 is an expanding, leakage-safe count of model-by-category co-occurrence with a flagged global fallback. C3 exempts same-chain occupancies, since standing reservations expand into back-to-back blocks meeting at midnight and counting those continuations as conflicts would make 6.59% of historical choices self-conflicting instead of 2.91%. C4 excludes a base pad against its own lettered sub-positions but not sub-positions against each other, since excluding all dependent pairs would render 7.20% of historical choices infeasible against 1.26%. Two consequences are used repeatedly: C1 and C2 are unary and C3 and C4 pairwise, so none couples three or more occupancies; and C3 and C4 are monotone in H, since enlarging the schedule can only remove stands from .
2.3. Masked Policy Class
Definition 2
(Masked policy class). Let be any measurable scoring function. The masking operator m produces the policy
with ties broken by a fixed deterministic rule. The masked policy class is .
Every policy compared here lies in , differing only in ; for the value-based agents of Section 4.5 the mask assigns logits to infeasible actions, reproducing (7) exactly.
2.4. Feasibility by Construction
Call H admissible if every decision-created pair in it satisfies C3 and C4 and every such occupancy satisfies C1 and C2 for its own request.
Proposition 1
(Feasibility by construction). Let and let be the inherited boundary schedule of an episode. Then for every epoch t the assignment produced by π satisfies C1–C4 with respect to , and the terminal schedule is admissible. In particular, π creates no violation of C1–C4 that was not already present in ; and by (7) a policy of returns ∅ only when .
Proof.
Induct on t. For the claim on decisions is vacuous, no decision having been taken, and is inherited. Assume every decision up to epoch satisfies C1–C4 against the schedule prevailing when it was taken. At epoch t, if then by (7) and , so the hypothesis carries over; otherwise satisfies C1–C4 with respect to by (6), for any , because the maximisation in (7) ranges over and nothing else. A pair certified at epoch t also stays certified: requests are decided in order of reserved start, so when is added every occupancy it can conflict with is already in and was checked, and each later pair is checked at the epoch of its later-starting member against the full committed schedule. Because C3 and C4 are pairwise, a pair’s certification never depends on occupancies added afterwards, and because C1 and C2 are unary they cannot be invalidated by later decisions. Hence every decision pair in is certified, which is admissibility. □
Remark 1.
The guarantee is relative to the four modelled predicates and the current inventory snapshot. It is exercised rather than assumed: the argument above is what the replay engine asserts at run time, and across 516 test days and every masked policy, including both offline-RL agents, the measured violation count is 0 (Section 5.4).
2.5. Conservative Value Bound Under Masking
The offline-RL instantiation of Section 4.5 learns by conservative Q-learning [11] restricted to . The restriction is not cosmetic: CQL’s log-sum-exp regulariser ranges over the action set, so shrinking it changes the penalty.
Proposition 2
(Conservatism under masking). Let be the behaviour policy generating the logged dataset and suppose for every and every s in the support of . Consider the masked CQL(H) fixed-point iteration in which the regularising maximisation is restricted to the masked simplex ; let denote its fixed point and let be the softmax over induced by . This μ is both the entropy-regularised maximiser and the policy the bound certifies, and satisfies the masked policy-evaluation fixed-point equation for it. Then, evaluated under μ itself,
provided α exceeds the (policy-dependent) threshold of [11], Theorem 3.2, instantiated at μ on ; in the exact-Bellman case any suffices. The two policies in play are kept distinct: is the data-generating behaviour and μ the induced softmax the bound certifies.
Proof sketch.
The support-restriction argument of [11] is pointwise in the action set, so replacing by leaves the per-iteration update in the form on , with the entropy-regularised softmax over . The strengthened support assumption keeps for every feasible action, so the correction is finite wherever places mass; under it is nonnegative, by Cauchy–Schwarz, whereas the same sum under an arbitrary can be negative, which is why the bound is stated at only. Appendix A.3 gives the full argument. □
Remark 2
(Scope: the deployed learner is not the object of Proposition 2). Two scope restrictions travel with Proposition 2. First, it concerns the policy-evaluation CQL(H) fixed point under an exact Bellman operator, whereas the learner of Section 4.5 is a control variant: Bellman-optimality target , exponential-moving-average target network, Huber temporal-difference loss, finite-epoch stochastic optimisation, and a behaviour-cloning-distilled warm start. Second, the strengthened support assumption is demanding and the logged policy does not meet it: the recorded behaviour is a near-point-mass, placing one action among roughly 109 feasible stands. The warm start smooths the initial Q-estimate, not the data-generating distribution entering the penalty ratio , so it does not make the full-support assumption hold on the logged data and we make no such claim. Proposition 2 is therefore about an idealised masked policy-evaluation operator whose reference behaviour has full support on : it isolates and formalises the conservatism mechanism, that restricting the log-sum-exp penalty to preserves the lower-bound structure the deployed control learner inherits informally. It motivates that mechanism rather than certifying the agent, whose behaviour is reported, as a non-confirmed hypothesis, in Section 5.5. Training uses only reachable events, those whose logged action lies inside the mask, which is 95.53% of test events.
2.6. The Expected-Conflict Functional
Because the recorded end of a reservation equals its requested end here (Section 3.1), a schedule satisfying C3 has exactly zero reservation-based overlap minutes, so that channel cannot separate masked policies. We replace it with expected conflict under the calibrated predictive model of Section 4.1.
Definition 3
(Predictive duration distribution). For event i let be its calibrated predicted duration quantiles, clipped below at one minute and sorted if the raw predictions invert, and put , . The duration has the CDF that is piecewise linear through the knots
equal to 0 below and to 1 above . Coinciding knots merge, retaining the largest CDF value, so is right-continuous and may carry atoms. Write for the survival function and for the random end time.
The support is bounded by construction, and the linear tails, fixed in the preregistration, let non-meeting pairs be pruned exactly.
Assumption 1.
Throughout Section 2.6 and Section 5:
- (A1)
- Independence: and are independent for .
- (A2)
- Deterministic starts: each reserved start is known and not random.
- (A3)
- Pairwise accounting: the risk of a schedule is the sum of its same-stand pair terms; k-fold simultaneous occupancy is counted times.
- (A4)
- One-sided buffer: the earlier occupancy’s end is inflated by β and the later occupancy’s start is not, so a pair contributes the expected overlap of with .
Assumption 1 (A1) is the strongest: durations at one airport on one day share shocks (weather, an inbound delay wave), and positive dependence would raise joint tail risk, so the functional understates it, while (A3) inflates the absolute level on busy stands. Because these conventions are identical for every policy compared, a paired difference in ECM is a fair comparison within the calibrated occupancy-risk model that defines the functional. It is not model-free, and neither the absolute level nor, strictly, the sign is guaranteed to survive a materially different duration model or a strong positive-dependence correction (the sensitivity analysis is reported in Section 5.8).
Theorem 1
(Expected pairwise conflict in closed form). Let i and j be committed to the same stand and ordered so that i precedes j in the total order (so , with the event identifier breaking equal-start ties), and let follow Definition 3 under Assumption 1. Then the expected buffered overlap
satisfies
and the integral evaluates exactly. Let be the ordered union of with the knots of and of that fall in . On each segment write , , and let and be the values and slopes of the two survival factors at . Then
If either or the buffered supports do not meet, then .
Proof.
We prove (11) here and (12) in Appendix A.1. Because we have , so
the integrand being the indicator of , which is empty exactly when the positive part vanishes. Both factors are nonnegative and measurable, so Tonelli’s theorem allows exchanging expectation and integral, and (A1) factorises the expectation of the product,
the second equality using (A2) to treat as constants. By Definition 3, the supports are bounded, so for and for ; truncating the upper limit at therefore changes nothing, which is (11). □
Equation (12) is exact, not a quadrature rule, and the implementation evaluates it verbatim; Appendix A.1 carries the segment algebra.
Remark 3
(Ordering at equal starts). Under the one-sided buffer (A4) depends on which member is treated as earlier when , since only the earlier end is inflated. The total order fixes this tie deterministically, matching the code (risk.py sorts by (start_min, event_id) before pairing), and the same order fixes the later-starting member used for day attribution below. This specifies the convention the implementation already uses; no reported number changes.
Definition 4
(Daily risk functional). Let be the counted pairs of day d: same stand, distinct reservation chains, later-starting member a decision of day d. With the number of requests the policy scored on day d,
We report the mean of over evaluation days and call it expected conflict minutes per 100 reservations, ECM.
These conventions mirror the mask and are applied identically to all policies: C3’s same-chain exemption carries over, and day attribution keeps inherited boundary-versus-boundary pairs out of every policy’s account.
Corollary 1
(Monotonicity). Write for the effective separation of the pair. Then
is nonincreasing in δ; hence is nonincreasing in the start separation and nondecreasing in the buffer β. If in addition in the usual stochastic order, then replacing by does not increase , and symmetrically for j.
Proof.
Substituting in (11) gives the displayed form; survival functions are nonincreasing, so is nonincreasing for every u and the integral inherits it. For the stochastic-order claim, means pointwise, and the integrand is monotone in each nonnegative factor. □
Note the direction: is a fixed convention of the metric, not a lever a policy pulls, and what lowers ECM is separating a pair in time.
Corollary 2
(Degeneracy and consistency with the mask). Suppose and are point masses at and . Then
the deterministic one-sided buffered overlap. In particular, for a distinct-chain pair () that satisfies C3 with reserved durations, ; consequently every mask-feasible schedule has when evaluated with reserved durations in place of the predictive distributions.
Proof.
Point masses make , so the integrand in (11) is the indicator of and its integral is the length of that interval, or zero if empty. For the second claim, restrict to a distinct-chain pair, over which C3’s quantifier ranges (Definition 1); then makes unattainable, so C3 reduces to , under which . The distinct-chain restriction is exactly the counted pair set of (Definition 4), so the schedule-level conclusion is unaffected: a same-chain pair may carry positive but is never counted. □
So ECM is a strict refinement of the mask, and the corollary also forces the identically zero risk term of the uncertainty-blind twin of Section 4.4, registered in advance and verified on test. Because (12) is exact only if implemented correctly, we validated it against Monte-Carlo simulation before any policy was tuned: on 30 validation days with 2000 inverse-CDF draws each, all 30 fell inside the preregistered tolerance, aggregate relative difference 0.00159 (Appendix A.2).
2.7. Reference Mixed-Integer Program
To quantify how much room the historical plan leaves, we also solve each sampled day jointly. Binary variables over stands passing C1 and C2 carry C3 and C4 as in-model clique constraints, the objective being the cost form of (1) at the same locked weights; pairs whose buffered overlap is below 60 min keep a purchasable hard constraint priced at the reward’s conflict charge, heavier pairs get pure clique constraints, and boundary occupancies are fixed at their historical stands. It is a hindsight bound, not a method: it sees the whole day at once and knows every reserved duration in advance (Section 5.6).
3. Data
3.1. The Central Caveat, Stated First
The main event log records one interval per reservation, whose recorded end is its requested end; no separate realised end exists. Four duration quantities recur below and are kept distinct throughout. The requested duration is the interval length on the reservation record, fixed at booking and available at the decision epoch. The recorded duration is the interval length as stored in the operator log; on this snapshot it is identical to the requested duration by construction. The realised duration is the time the aircraft actually occupied the stand; it is not observable in this source and is never used. The predictive occupancy duration is the distribution produced by the calibrated quantile model of Section 4.1 from operational context that excludes the requested duration, and it is the object the risk functional integrates. Every downstream consequence in this paper follows from that single fact. Reservation-based overlap minutes are structurally zero for any schedule satisfying predicate C3 (Corollary 2), so that channel cannot rank masked policies. The airport also publishes a table of realised gate events, but its rolling two-week window (17 June 2026 to 30 June 2026) does not intersect the parking log, which ends 31 May 2026: zero rows join, and the planned-versus-actual check is not merely noisy but impossible on the current snapshots. This is why the conflict channel of Section 5 is the model-based functional of Theorem 1, and why every conflict claim in this paper is phrased as expected conflict under the calibrated occupancy-risk model.
3.2. Sources
Six public tables enter the pipeline, all open data. The main event log is the SFO parking-activity table (5rkh-waic), 194,503 rows over 213 consecutive calendar months, split into 157,273 call-in visits and 37,230 standing blocks, the latter with recurring daily expansions of one long-lived reservation, which is why predicate C3 needs its same-chain exemption; median occupancy is 8.25 h, first and 99th percentiles 1.25 and 98 h. The stand inventory (2ymc-znns) lists 287 stands with date-effective validity windows, a size category and a configuration flag: 220 hardstands and 67 gates, of which 182 are active today and 105 are not, and 186 are configured independent against 101 dependent, multi-large or multi-dependent. Because 34,195 historical events sit on stands inactive today, C1 reads the validity window and ignores the present status, and 53 historical stands are absent from the snapshot altogether. A tail-to-model reference (u7dr-xm3v, 8999 rows), the realised gate-event table (chfu-j7tc, 14,718 rows), NOAA Local Climatological Data v2 for station USW00023234 (261,352 hourly records) and the Bureau of Transportation Statistics On-Time Reporting Carrier database restricted to SFO (1,926,950 rows) complete the set. Licences and sources are given in the Data Availability Statement.
Linkage was measured at audit time: tail numbers resolve to an aircraft model for 98.23% of the rows carrying a tail, though 27.81% carry none at all, and on matched rows the recorded model agrees exactly for 96.96%; stand names resolve to the current inventory for 97.69% of rows and weather joins within three hours for 99.92% of feature rows. The BTS join is the weakest link by design, covering US domestic reporting carriers from 2019 onward, so only 56.09% of activity rows find a same-day flight record (60.98% within one day).
Quality checks are unremarkable in scale but were acted on: 1008 exactly duplicated business rows and 7 rows whose end precedes their start are dropped. A further 1596 rows (0.82%) start before the running maximum end on the same stand; these are genuine historical double-use, kept as recorded, and the dominant reason the mask does not reach 100% recall (Section 4.2).
3.3. Splits
Splits are rolling-origin and time-extrapolating; nothing is shuffled. After de-duplication, the windows hold 5776 warm-up events (1 September 2008 to 31 December 2009, excluded from training by convention), 138,132 training events (1 January 2010 to 31 December 2022), 28,153 validation events (1 January 2023 to 31 December 2024) and 21,434 test events over the 516 days from 1 January 2025 to 31 May 2026. Every calendar month in train, validation and test contains data. Volume in 2020 falls to 51% of 2019 but never breaks; we keep the period in training and expose it through a period indicator rather than excising it. The inventory a split sees shifts materially across windows, 132 distinct stands appearing in test against 248 in train, which is one source of the time-extrapolation decay reported in Section 5.2.
The test window was read for the preregistered evaluation exactly once, after every configuration had been locked on validation. This is a claim about the preregistered round, not about the split’s whole history: Section 4.7 states the protocol and tabulates every test-set contact across all rounds.
4. Methods
4.1. Calibrated Occupancy-Duration Model
The predictive distributions of Definition 3 come from a LightGBM [12] quantile model of occupancy duration, trained at under the pinball loss [13] against the recorded duration in minutes. The following is what the predictive distribution represents. Because the recorded end of a reservation equals its requested end on this snapshot (Section 3.1), the label is the requested duration, and that value is on the record at the decision epoch. The model is therefore not forecasting a quantity that is unknown when the stand is chosen. What it estimates is the conditional distribution of occupancy duration across comparable movements, given aircraft category, operating company, calendar and time-of-day position, weather and leakage-safe expanding history, with the requested duration itself withheld from the feature set. The spread it reports is between-movement variability for operations of this kind; it is not the deviation of an individual aircraft’s actual occupancy from its own request, which this data source cannot observe. The risk term consumes that spread as a planning hedge: a candidate stand is scored by the overlap expected if this movement behaves like comparable movements, rather than by treating the single requested value as exact. Whether between-movement spread is a good proxy for within-movement uncertainty about actual occupancy is not testable here, and Section 6.5 records that as a scope condition rather than a demonstrated property.
Feature construction is leakage-safe by exclusion: the requested end and duration are dropped because they equal the label here, the assigned stand, its area prefix and the area load at start because they are outcomes of the decision the model informs, and the raw tail number for cardinality. What remains is aircraft model, operating company, calendar and time-of-day structure, weather, BTS flight context, and stacked expanding company-by-model duration quantiles computed strictly from events completed before the decision date, which is the leakage rule the rule baseline also obeys. Twelve configurations were compared on validation pinball loss, crossing an absolute against a residual parameterisation, raw against target scaling and two capacity settings; the winner fits residuals around the expanding pair-quantile anchor supplied as an initial score, with 127 leaves and a minimum of 100 samples per leaf, five seeds combined by the per-level median and rearranged to restore monotonicity, needed for 7772 events across all splits. Intervals are calibrated by split-conformal conformalized quantile regression [14] on validation only (, fitted correction on the scale), so validation coverage is partly by construction and test coverage the honest generalisation check; we report both, with the pinball loss, a strictly consistent scoring function for the quantile functional [15].
4.2. Mask Construction and Candidate Sets
Predicates C1–C4 are materialised once into a candidate table of 21,161,217 request-by-stand rows with per-predicate flags and decision-time features. Feasible sets are large (mean 109.3 stands overall, 131.4 on test, 10th and 90th percentiles 85 and 134), so the mask constrains without starving the scorer, and C2 falls back to its flagged global compatibility window on 4.844% of feasible rows. Empty feasible sets occur for 363 events (0.19%), all inside the excluded warm-up window, and no train, validation or test request ever has one, which is why the deferral rate is zero for every policy in Section 5. Mask recall, the share of requests whose historical stand lies inside the feasible set, is 93.53% overall, 96.94% on validation and 95.53% on test; the dominant residual cause is C3, history having placed different reservation chains on one stand less than a buffer apart 5548 times. Those are precisely the conflicts the policy is asked to avoid, so counting them infeasible is intended behaviour rather than mask error, at the price of an upper bound below one on any agreement-with-history metric.
4.3. Imitation Rankers
Four scoring functions imitate the historical planner, each trained on the feasible candidate sets with the historical choice as the positive. B1 is a rule rather than a learner, preferring the area of the reservation’s recorded stand and, within it, the feasible stand with the largest temporal buffer; B2 is an XGBoost [16] ranker trained with the LambdaMART objective [17,18] on 40 sampled negatives per request; B3 a listwise multilayer perceptron over the same features; B4 a relational graph attention network [19,20] over a heterogeneous request-stand graph with same-area, pad-variant and request-stand relations. Every learned ranker uses five seeds combined into a mean-score ensemble, with a frequency rule over company-by-stand counts as a floor. The preregistered round added two attempts to improve on B2: the hybrid ranker concatenates the seed-matched 64-dimensional penultimate representation of the frozen GAT to the B2 feature stack, stored as in half precision, and the blend normalises each model’s ensemble score to a per-request ordinal rank and takes a convex combination, weights chosen on validation from a simplex grid of step 0.05 (0.70 on XGBoost, 0.30 on the GAT, zero on the MLP).
4.4. Uncertainty-Aware Policy and Its Blind Twin
The uncertainty-aware policy is a masked greedy policy in the sense of Definition 2, whose scorer subtracts a risk term from the frozen imitation score:
Here, is the frozen five-seed B2 ensemble score (no ranker is retrained) and the calibrated 0.9 knot of Definition 3, so the risk term is a point summary of the predictive distribution evaluated against the occupancies already committed to the candidate stand, including the inherited boundary. The weight was selected on validation from the preregistered grid , the whole grid counting as one development iteration; the argument minimum was , at the grid edge, though the curve is flat above and the improvement from 8 to 16 is 0.04 ECM.
The blind twin is the same functional form with reserved durations substituted for the predicted quantiles. Corollary 2 predicts what happens: C3 already forbids reserved-time overlap on the feasible set, so the reserved-duration risk term is identically zero and collapses to plain greedy scoring on , which is exactly the B2 policy. Registered as a structural observation before the evaluation, this was verified decision by decision: 0 mismatches out of 28,153 validation and 0 out of 21,434 test decisions. The twin is therefore not an independent baseline but a mechanism ablation, holding ranker, mask, replay and tie-breaks fixed and removing only the information channel; the like-for-like decision-quality comparison is against B1, which never sees the occupancy model at all.
4.5. Masked Offline Reinforcement Learning
The offline agent is a discrete conservative Q-learner [11] over the masked candidate set. Infeasible actions receive logits, so both the greedy policy and the CQL log-sum-exp regulariser range over only, the algorithmic mirror of Proposition 1, unit-tested against perturbed masks. Training and evaluation use the identical objective, the three-term composite reward ( expected conflict, preference, churn) with the size-mismatch term of (1) excluded from both per the preregistration and reported only descriptively, so there is no train/evaluation mismatch. The conflict term is the marginal expected-conflict contribution of a decision, the sum of over pairs it creates; attribution is exact, since summing contributions over a day reproduces the schedule-level total, an identity we verify directly. Training uses 91,164 logged transitions from reachable training events from 2015 onward. One asymmetry is structural, shared by every method’s training data: on logged transitions the behaviour policy is the history, so the churn term is identically zero there and acts only at evaluation time on counterfactual trajectories, leaving the agent to learn from an expected-conflict and preference signal alone. That is a property of the data, not a difference in objective, and it shapes the result in Section 5.5.
The network is warm-started for two epochs by distilling the frozen B2 ensemble score, then trained for six CQL epochs with learning rate , batches of 256 requests, target-network exponential moving average 0.005 and a Huber loss. Only two knobs were tuned, over three validation iterations: the conservatism weight and the discount , ending at , . Three seeds are combined by averaging Q-values and the policy is the masked greedy argument maximum of that mean. An implicit Q-learning agent [21] with expectile parameter 0.7 and every shared hyperparameter copied from the locked CQL configuration serves as a non-confirmatory ablation; nothing was selected on it.
4.6. Replay Protocol and Metrics
Evaluation replays each historical day as one episode: requests arrive in reserved-start order, the episode is initialised with the boundary occupancies inherited from previous days at their historical stands, and the committed schedule updates after every decision, so a policy sees the consequences of its own earlier choices. All policies replay the same days from the same initial condition, making the day the natural resampling unit for the paired statistics below. Recorded per day are the model-based ECM of Definition 4; reservation-based overlap per 100 requests, retained for continuity with the historical plan though Corollary 2 makes it zero for every masked policy; agreement, the share of requests placed on the historical stand, and churn its complement, both imitation-fidelity diagnostics; defer rate; and mask violations, which Proposition 1 predicts to be zero.
4.7. Statistics and Preregistration
All comparisons are day-paired: we form the per-day endpoint difference and estimate its mean with a percentile bootstrap over 10,000 resamples [22], resampling whole days so that within-day dependence is preserved. Days are not independent of one another, however: the per-day difference series is serially correlated (Section 5.4), so the primary interval is reported alongside a moving-block bootstrap [23] across block lengths of 1 to 28 days, and we treat the block-length sweep rather than any single length as the interval statement. We test each difference with a two-sided Wilcoxon signed-rank test [24], reporting the number of non-zero days, and quantify effect size with Cliff’s [25], computed over the cross product of the two per-day samples and therefore measuring unpaired stochastic dominance. Because the design is day-paired we also report the matched-pairs rank-biserial correlation, formed from the signed ranks of the per-day differences; the two answer different questions and we give both. The preregistration labels the whole-day procedure the “day-block bootstrap”, after the block resampling of [23]. Every interval quoted in the results is the whole-day percentile interval; where the block-length sweep is relevant it is reported explicitly alongside it (Section 5.4). Multiplicity is controlled by Holm’s step-down procedure [26] within the preregistered family.
Five things were frozen before any iteration ran: hypotheses H-A to H-D, their endpoints and comparator sets, the decision rule, the and weight grids, and the ECM definition of Section 2.6. The decision rule requires Holm-corrected and a 95% confidence interval excluding zero and a point estimate in the registered direction, over a family of comparisons. The protocol was fixed before the first iteration; development used train and validation only, and each track had a budget of at most three validation iterations. For this preregistered round, the frozen test split of 516 days was read exactly once, with every endpoint produced from that single reading, and no configuration reported here was selected on it. Three deviations, all definitional clarifications, were recorded before that evaluation.
That “exactly once” is specific to the preregistered round and does not imply the split had never been touched: Table 2 enumerates every contact across all rounds. Of the contacts that occurred before and after it, none selected a reported configuration, and only this round was preregistered.
Table 2.
Every contact with the frozen test split across the project’s history. No contact selected a configuration reported here; all selection decisions were made on validation, and only this round was preregistered. Reads are enumerated, not summarised.
All computations were run in Python 3.13.9 (Python Software Foundation, Beaverton, OR, USA) with LightGBM 4.6.0 (Microsoft Corporation, Redmond, WA, USA) for the occupancy-duration model; XGBoost 3.2.0 (https://xgboost.ai), PyTorch 2.12.0 (https://pytorch.org) and PyTorch Geometric 2.8.0 (https://pytorch-geometric.readthedocs.io) for the rankers and offline-RL agents; PuLP 3.3.0 with the COIN-OR CBC solver (https://github.com/coin-or/pulp) for the reference program; and SciPy 1.16.0 (https://scipy.org) for the statistical tests.
5. Results
Every number here is transcribed from the recorded results: the preregistered quantities from the single one-shot reading of the 516-day test split, and the comparison quantities from the earlier full-pipeline evaluation on the same split. Section 5.4 collects all six preregistered comparisons, with verdicts and paired statistics, in one table.
5.1. The Occupancy-Duration Model Is Calibrated and Beats Its Baselines
The forecaster passes its preregistered gate on both splits: on test its mean pinball loss is 205.08 min against 243.79 for a point- model emitting the same value at all three levels and 209.50 for the historical company-by-model expanding-quantile baseline (Table 3). Day-paired over the 516 test days the improvement over the historical baseline is min/day, 95% CI , Holm-corrected , ; over the point model, min/day , Holm , . The margin over the historical baseline is small and we say so: 2.1% of the loss, comparable to the across-seed standard deviation, and at the 0.9 level the ensemble and the historical baseline are essentially tied on test (310.03 vs. 310.74).
Table 3.
Occupancy-duration forecaster comparison on validation and test. Test is never used for selection.
Coverage of the nominal 80% central interval is 0.800 on validation and 0.782 on test (Figure 2, left), the validation figure partly by construction and the test figure the honest check, inside the preregistered tolerance band. A controlled ablation holding configuration, seeds and calibration fixed shows what the BTS flight context is worth here: min/day of pinball loss , . The same features are worth nothing measurable to the imitation layer (Section 5.7), a clean separation of where an external signal helps.
Figure 2.
(Left) Calibration of the occupancy-duration forecaster: quantile reliability on (a) validation () and (b) test () for the calibrated LightGBM model against the historical expanding-quantile baseline, and (c) coverage of the nominal 80% interval; numbers in Table 3. (Right) Preregistered negative result H-A: frozen graph-attention embeddings do not help the XGBoost ranker. The day-paired difference in reachable hit@1 is pp , Holm over days, significant opposite to the registered direction, so H-A is not confirmed. Event-pooled hit@1 by year puts the deficit in 2025 (plain 0.4456 vs. hybrid 0.4278) and is roughly flat in 2026, where the hybrid is marginally higher (0.3036 vs. 0.3063). LGBM, LightGBM; CQR, conformalized quantile regression; hist., historical baseline; val, validation; XGB, XGBoost; GAT, graph attention network; pp, percentage points.
5.2. Imitation Boundary Conditions
Four attempts to beat gradient boosting at imitating the historical planner all fail, two under preregistration.
5.2.1. Architecture
On test the XGBoost ensemble reaches hit@1 0.3971 against 0.3583 for the graph attention network, 0.3575 for the listwise perceptron and 0.2715 for the frequency rule (Table 4); day-paired, the graph model loses to gradient boosting by , Holm , , but ties the perceptron at an equal five-seed budget, , Holm . The architecture-decisive statement is therefore narrow: graph attention does not beat gradient boosting on these tabular features, and we claim nothing about it against the perceptron.
Table 4.
The imitation channel on the 516 test days. Left columns: top-k feasible hit rates over the reachable denominator. Right columns: the replay summary of the earlier full-pipeline evaluation, as day means.
5.2.2. Graph Structure as a Feature Extractor (H-A, Not Confirmed)
Concatenating the frozen graph model’s penultimate representation to the boosted ranker was preregistered to raise hit@1; it lowers it, by , Holm , over 516 days, significant opposite to the registered direction. The deficit sits in 2025 (event-pooled hit@1 0.4456 plain against 0.4278 hybrid) and essentially vanishes in 2026 (0.3036 against 0.3063; Figure 2, right). The mechanism is visible in development: the embeddings are in-sample on training rows, so the booster over-trusts them and early stopping truncates before the ordinary features are learned, two seeds stopping at a single tree. All three variants tried on validation (raw 64-dimensional, scalar score, 16 principal components) were negative, so none survived even the development split.
5.2.3. Score Blending (H-B, Not Confirmed)
A rank-normalised convex blend, weights chosen on validation, was preregistered to beat the best single ranker. On validation it gained , not significant (); on test it loses , Holm , so the validation-selected gain reversed and the reversal is itself significant. A gain of that size, selected over a 231-point grid, is indistinguishable from selection noise until a frozen split says otherwise.
5.3. Replay Agreement
Under replay the same ordering holds on agreement with the historical plan: 0.327 for the boosted ranker, 0.294 for the perceptron, 0.287 for the graph model, 0.210 for the frequency rule and 0.129 for the same-area rule (Table 4), the boosted ranker beating the rule by , Holm , . Agreement decays from 0.364 in 2025 to 0.238 in 2026, the expected cost of extrapolating across a changing inventory, and is capped below one by the 95.53% mask recall. Reservation-based overlap is 0.00 per 100 requests for every masked policy and 373.25 for the unmasked historical plan, which is exactly what Corollary 2 predicts.
5.4. Uncertainty-Aware Allocation: The Headline Result
The uncertainty-aware policy attains a day-mean ECM of 615.9 against 1175.5 for the B1 same-area rule: a 48% reduction in model-based expected conflict under the calibrated occupancy-risk model, day-paired , 95% CI , Wilcoxon , Holm-corrected , Cliff’s , all 516 days contributing a non-zero difference (Table 5). The per-day difference series is serially dependent (lag-1 autocorrelation ; Ljung–Box on 14 lags ), so we also report the interval under a moving-block bootstrap: it excludes zero at every block length from 1 to 28 days, widening from at one day to at 28. The matched-pairs rank-biserial effect size is . We lead with this comparison because B1 never sees the occupancy model. The ratio-of-sums view agrees, ECM per unit of exposure, as do the medians, 558.7 against 788.7; dispersion also collapses, the day-level standard deviation being 311.1 for against 1255.7 for B1.
Table 5.
The six preregistered comparisons of family F3, evaluated in a single one-shot reading of the test split.
Against its uncertainty-blind twin the gap is an order of magnitude larger, , Holm , , with lower on all 516 of 516 days. That contrast is a mechanism ablation, not a co-equal baseline: as Section 4.4 predicted from Corollary 2, the blind twin’s risk term is identically zero and it coincides with the B2 imitation policy, verified with 0 mismatches over 21,434 test decisions. Removing only the information channel therefore isolates what consuming the uncertainty estimates is worth: nearly all of the gap between imitating history and managing risk (Figure 3).
Figure 3.
Expected conflict by policy over the 516 preregistered one-shot test days. Every quantity is model-based expected conflict under the calibrated occupancy-risk model of Theorem 1, not reservation-based overlap, which is structurally zero for every masked policy on this snapshot. The headline is 615.9 against B1 1175.5, a 48% reduction, because B1 never sees the occupancy model; the uncertainty-blind twin (7142.6) is a mechanism ablation, not a co-equal baseline, since it provably coincides with the B2 policy (0/21,434 mismatches) and so isolates the information channel rather than the ranker. CQL (3728.9) and the IQL ablation (7599.5) are shown for reference. Bars are day means with percentile bootstrap 95% intervals resampling whole days (10,000 resamples), logarithmic ordinate; statistics in Table 5 and Table 6. CQL, conservative Q-learning; IQL, implicit Q-learning; RL, reinforcement learning; XGB-BC, XGBoost behaviour cloning.
Table 6.
One-shot test policy summary: model-based expected conflict, composite reward, and behavioural diagnostics.
5.4.1. The Gain Is Not Bought by Refusing to Decide
A risk-penalising policy could trivially lower expected conflict by deferring hard requests or scattering aircraft across the field; neither happens. The deferral rate is exactly 0.0 for every policy, no request ever reaching an empty feasible set (Section 4.2), and assigns a stand to all 21,434 test requests. Its churn rate, 0.729, sits below B1’s 0.871 and above the blind twin’s 0.673: the risk term trades some imitation fidelity (agreement 0.271 against 0.327) for a large reduction in expected conflict while departing from the historical plan less often than the rule it beats. Counted conflict pairs per day fall to 13.7 for from 33.7 for the blind twin, with B1 at 8.3: B1 achieves few pairs by spreading aircraft at a heavy agreement cost, while reaches lower expected conflict with more pairs, because the pairs it creates are well separated (Table 6).
5.4.2. What Validation Said First
On validation, the same policy reached 634.7 ECM against 955.6 for B1 (the level implied by the paired mean difference ; B1’s tabulated day-mean is 955.5), but the day-level effect size was slightly positive (): the mean favoured while the median validation day favoured B1. On test both agree (). We report both rather than the more flattering one.
5.5. Safe Offline Reinforcement Learning: A Partial Result
Masked discrete conservative Q-learning was preregistered to beat both the imitation ensemble and the rule on the composite reward; it beats one. Per day the ordering is , the rule , CQL , the imitation ensemble . Against the ensemble CQL gains , Holm , , satisfying every criterion; against the rule it loses , Holm , in the wrong direction. The preregistered rule required both, so H-D is not confirmed and no claim of reinforcement-learning superiority appears anywhere in this paper.
On the conflict channel CQL lands between the two at 3728.9 ECM, below the imitation ensemble and above the rule. The offline agent does learn to avoid expected conflict relative to the behaviour it was distilled from, but stops well short of a rule that simply maximises temporal buffer, and much further short of , which reaches 615.9 using the very same base scores (Section 6 argues why). Feasibility held throughout: zero mask violations for the ensemble and for each of the three seeds.
The conservatism parameter behaves as the mechanism formalised by Proposition 2 would suggest, though that proposition concerns the policy-evaluation fixed point, not this control learner (Remark 2). Lowering from 1.0 to 0.1 to 0.01 on validation released the policy from the behaviour prior monotonically (agreement , ECM ), improving the reward at every step and never producing a violation. The sweep is qualitatively consistent with the conservatism interpretation; it is three points, not a proof.
The implicit Q-learning ablation, run once with every shared hyperparameter copied from the locked CQL configuration and nothing selected on it, collapses: reward per day, ECM 7599.5, agreement 0.011, churn 0.989, and 52% of its assignments sent to gates against 1.4% for the imitation ensemble. Learning Q only on logged actions, without the conservative term ranging over the whole feasible set, discards the behavioural anchoring the warm start provided; that the anchoring term is load-bearing is a useful negative for anyone instantiating offline RL under a mask.
5.6. Optimality Reference
On 20 test days sampled across the volume range, the day-joint mixed-integer program of Section 2.7 solved to proven optimality in 20 of 20 cases. Mean objective headroom over the historical plan is 66.41 units per day, 62.2% of the historical objective or 157.9 per 100 requests; the median is 22.16, 6 of the 20 days have no headroom at all, and the optimal plans buy it cheaply, moving 1.45 requests per day off their historical stand. This is a hindsight bound, not a method we propose: the program decides the whole day at once with every reserved duration known, where the planner decided sequentially under uncertainty. Its use is as a scale, showing that the sequential policies operate in a problem with a large objective gap in principle, so the differences between them are not squeezed against a ceiling.
5.7. Extreme Days and Feature Ablations
A subset of 52 test days (top decile by reservation volume crossed with weather severity, defined before this round and outside the preregistered family) preserves every sign of the headline. There, reaches 609.8 ECM against 1640.1 for the rule, , and against the blind twin ; the rule’s risk rises by 46% on extreme days (1640.1 against 1123.4 on ordinary days) while barely moves (609.8 against 616.6). The unmasked historical plan accrues 545.9 reservation-based overlap minutes per 100 requests on extreme days against 353.9 on the rest, and agreement falls from 0.331 to 0.291. These are descriptive, without Holm correction.
Three feature ablations round out the picture. Replacing the quantile occupancy features by point predictions costs of replay agreement (Holm ); removing weather moves hit@1 by with a raw that does not reach significance after correction (), so we make no claim for the weather signal; and removing the BTS features changes nothing measurable for the imitator (, ), in sharp contrast to their clear value to the forecaster (Section 5.1).
5.8. Robustness of the Conflict Ranking
The headline compares two policies inside one calibrated occupancy model, under an independence convention, through a pairwise accounting of overlap. Each of those three choices could, in principle, carry the result. None of them does, and we establish this directly rather than by argument. The analyses in this subsection sit outside the preregistered family, were run as a single batch on the frozen split, and are recorded as such in Table 2; no locked configuration was altered and no preregistered endpoint was recomputed. Table 7 collects them, and Table 8 adds, on validation, the sensitivity of the ranking to the risk weight over the whole preregistered grid.
Table 7.
Robustness of the –B1 expected-conflict ranking on the 516 frozen test days. Every row is a day-paired comparison of the same two policies under a different stress applied to the evaluation; ECM is in expected conflict minutes per 100 reservations. The first row is the preregistered one-shot, reproduced exactly.
Table 8.
Validation sensitivity of the risk weight over the whole preregistered grid (731 validation days, day-mean ECM per 100 reservations). The rule baseline B1 attains 955.5 on the same days. The argument minimum lies at the grid edge, but the curve is flat above , the gain from 8 to 16 is 0.04, and the ranking against B1 is preserved at every grid point, so the reported configuration reflects a broad plateau rather than a tuned optimum.
5.8.1. The Ranking Does Not Depend on One Calibrated Model
Re-running the comparison on a second, independently conformalised duration model (trained without the flight-context features, so it is a genuinely different forecaster rather than a perturbation) preserves the sign and widens the margin, from 47.6% to 57.4% (). The harder variant is the third row, where the locked is graded by a model it never consumed: this severs the shared-model channel entirely, and the advantage survives at 22.9% (). Whether the second model informs the policy or only judges it, is ahead.
5.8.2. Positive Dependence Strengthens the Case Rather than Weakening It
Assumption (A1) is what lets the joint tail factor, and same-day durations plausibly share shocks, so we couple them through a day-level common factor and sweep the copula correlation. Absolute expected conflict rises for both policies (by roughly half at , confirming that the independent functional understates joint tail mass) but the paired advantage is preserved at every level and grows at the top of the range, with Cliff’s strengthening monotonically from to . The mechanism is visible in the pair counts: B1 concentrates its exposure in fewer but far larger overlaps, and it is precisely those that a common shock inflates.
5.8.3. What the Functional Counts, Measured
Because ECM sums pairwise overlaps, a k-fold simultaneous occupancy enters through terms, and the magnitude should be read with that convention in mind. Resolving the counted pair set into connected clusters of simultaneously conflicting occupancies puts a number on it: under , 95.3% of expected conflict minutes arise in isolated two-occupancy overlaps and only 4.7% in clusters of three or more, so ECM and a simultaneity-based reading of the same schedule almost coincide. Under B1 the figure is 35.5%. ECM is a policy-ranking functional over predicted occupancy risk; it is not a forecast of realised delay minutes, and the convention that inflates it inflates the rule baseline roughly seven times more than the proposed policy.
5.8.4. Where the Advantage Comes from
A comparator that consumes the same predictive quantiles through a different decision rule separates the two candidate explanations. B1-pred keeps B1’s mechanism exactly (same-area preference, then the largest buffer behind the arriving reservation) and changes only the information, taking that buffer from the predicted end rather than the reserved end. It reaches 824.3 ECM against B1’s 1175.5 (), and reaches 615.9. Predictive uncertainty is therefore valuable through more than one decision rule, which is the stronger claim for the information channel; and the composed policy converts it further, at 2.3 times B1-pred’s replay agreement (0.271 against 0.117), so the scoring rule buys operational continuity that the heuristic does not.
6. Discussion
6.1. Feasibility and Risk Are Separable, and They Compose
The architecture is a composition: a mask that owns feasibility and a scorer that owns everything else. Proposition 1 makes the split rigorous, since the guarantee holds for any measurable scorer, and the record confirms it is no artefact of well-behaved scorers: two offline-RL agents, whose scorers share no structure with the rule or the ranker, still produce zero violations across 516 days. The split frees the scorer entirely, be it a boosted ranker, a rule, a Q-network or a hand-written penalty, and the two halves compose in a specific direction. The mask alone does not manage risk: the B2 imitation policy is fully feasible yet carries 7142.6 ECM, six times the same-area rule’s. Nor does the forecaster alone: it is well calibrated (Section 5.1), yet the blind twin, identical ranker and no access to the forecast, is exactly the high-risk policy. Uncertainty must be consumed by the decision rule to matter, and demonstrates it: same mask, same ranker, same replay, one extra term, 615.9 against 7142.6. The B1 comparison shows the effect is not merely relative to a weak imitator, also beating the strongest rule baseline at lower churn and with no deferrals.
6.2. Why Online Model-Based Risk Beat Offline Value Learning Here
The most instructive contrast is between two policies sharing a base ranker: reaches 615.9 ECM while masked CQL, trained on a reward whose dominant term is the same expected-conflict quantity, reaches 3728.9. The offline agent optimises the right objective and still lands six times higher, for informational rather than algorithmic reasons. ’s risk term is evaluated online against the episode’s own committed schedule: at the moment of decision it reads which occupancies already sit on each candidate stand and integrates their predicted overlap with the request in hand. The Q-network has no such access; it sees a static feature vector summarising the logged state, and whatever conflict structure that vector fails to encode is invisible to it however much conservative training it receives. This is consistent with the offline-RL literature’s emphasis on distributional shift and representation error under function approximation [27]. Two structural facts push the same way: the training reward carries no churn signal, because on logged transitions the behaviour is the history; and the mask was a filter applied on top, not a learning target.
The reading we take is not that offline RL is unsuited to stand assignment, but that when a calibrated model of the uncertain quantity exists and can be queried at decision time, a policy that queries it directly is a strong baseline value learning must earn its way past.
6.3. The Imitation Channel Has a Ceiling, and It Is Not the Architecture
It would be easy to read the four failures of Section 5.2 as a verdict on graph learning, but the data do not support that: the graph model was trained under a disclosed budget, so only the head-to-head loss against gradient boosting on identical features is architecture-decisive, and even that is a statement about these tabular features.
A more likely ceiling is the target itself. Agreement with a historical planner is bounded above by 95.53% mask recall, and the residual disagreement is dominated by real historical double-use of stands, exactly the behaviour a conflict-avoiding policy should not imitate; a ranker that imitates perfectly would inherit the historical packing, which is what the ECM channel shows the imitation ensemble doing. Better imitation is therefore not obviously the direction of travel, and the H-A result, which traces to in-sample embeddings, is a symptom of pushing harder on a target with little left to give. What would change the picture is a different label: realised occupancy times.
6.4. A Measured Case for Preregistration
Hypothesis H-B is the paper’s most transferable methodological finding. Had the blend’s grid search and the test evaluation happened in the same loop, the natural narrative would have been an ensemble improvement, wrong in sign. Three of four preregistered hypotheses did not confirm and two reversed significantly; absent the frozen protocol and the single test reading, at least one would plausibly have been reported as positive. That is what the protocol bought on this data, and it is why we report preregistration as a methodological safeguard that underwrites the evidence rather than as a scientific finding in its own right: it converts the negative results from an absence of evidence into measured evidence about what does not work here.
6.5. Assumptions and Scope Conditions
The guarantees and the measured advantage each hold under stated conditions, and it is worth setting out that envelope precisely, both because it delimits what the evidence supports and because several of its edges are now measured rather than assumed.
6.5.1. What the Conflict Channel Is a Claim About
This operator log records reservations: the recorded end of an interval is its requested end, and the published realised-gate feed does not intersect the parking log on this snapshot. Every conflict quantity here is therefore a statement about expected overlap between reservations under the calibrated occupancy-risk model of Theorem 1, expected conflict computed from predictive duration distributions, not observed blockage at the stand. The 48% headline says that consuming calibrated occupancy uncertainty lowers model-based expected conflict against a rule that never sees the model, and Section 5.8 establishes that this survives a second calibrated forecaster and correlated durations. It is not a measurement of realised delay minutes, and we do not read it as one.
6.5.2. The Functional’s Conventions and Their Measured Effect
Assumption 1 (A1)–(A4) fix independent durations, deterministic starts, pairwise accounting and a one-sided buffer. All four apply identically to every policy compared, which is what makes the paired differences fair, and the two that could plausibly carry the ranking are now quantified rather than argued: relaxing independence through a day-level common shock preserves and slightly widens the advantage (Table 7), and resolving the pairwise sum into simultaneity clusters shows 95.3% of ’s expected conflict minutes sitting in isolated two-occupancy overlaps. Deterministic starts remain a modelling choice; relaxing them changes the mathematics rather than only the numbers, since a random turns the outer integral into a convolution that no longer factorises segment by segment, and that is the natural next extension of the functional.
6.5.3. What “Feasibility-Guaranteed” Guarantees
Proposition 1 certifies feasibility with respect to predicates C1–C4 and the stand inventory as published, for any scorer, pathwise; and the certificate is exercised, at zero violations across 516 test days and every policy evaluated, including two offline-RL agents whose scorers share no structure with the rule or the ranker. Two of the predicates are operational reconstructions: C4 infers physical sharing from stand naming, a base pad against its lettered sub-positions, and C1 reads the published validity windows of a current inventory snapshot. Requirements outside these four (ground-handling resourcing or contractual stand rights, for instance) enter as additional predicates supplied by the deploying operator rather than through F as specified here. Because the guarantee is a property of the policy class and holds for any measurable scorer, it composes with such predicates without re-derivation, which is precisely the property that makes the construction portable.
6.5.4. The Envelope This Evidence Covers
The results come from one airport’s operator log over 213 consecutive months, evaluated on 516 held-out days under time-extrapolating splits. That design tests generalisation forward in time, which is the axis a deployed allocator actually faces, and the year-by-year view shows where re-fitting earns its keep: as the stand inventory drifts, the imitation components lose ground while the risk channel does not, so a deployment would recalibrate the ranker on a roughly annual cycle and the occupancy model whenever the published inventory changes materially. Transfer to a second airport requires the same four predicates to be instantiated against that airport’s layout and inventory, and we make no claim beyond this operator’s log.
6.5.5. What the Behavioural Metrics Report
Replay agreement and churn are behavioural diagnostics, not quality measures: a policy raises agreement by reproducing historical packing, which the conflict channel shows to be the higher-risk choice. Agreement is additionally bounded by construction, since the historical stand lies inside the feasible set for 95.53% of test requests, the residual being genuine historical double-use by distinct reservation chains. We therefore read agreement as a measure of operational continuity, that is, of how much of the existing plan a policy leaves undisturbed, and report it alongside the conflict endpoint rather than as a substitute for it.
7. Conclusions
Stand pre-assignment needs two guarantees usually traded against each other: constraints that hold on every realised path, and a decision rule that takes occupancy uncertainty seriously. This paper obtains both by composition. A masking operator built from four hard predicates makes feasibility a property of the policy class rather than of the learned scorer (Proposition 1), with zero violations across 516 one-shot test days and every policy evaluated, including two offline-RL agents. Onto that class we graft a risk term from the main mathematical result: under piecewise-linear predictive duration distributions with bounded linear tails the expected buffered overlap of a same-stand pair has an exact closed form (Theorem 1), degenerating precisely to the constraint it generalises when those distributions collapse to point masses (Corollary 2).
Empirically, in a preregistered one-shot replay of all 516 test days at SFO, a greedy policy consuming this risk model reduces model-based expected conflict under the calibrated occupancy-risk model from 1175.5 to 615.9 min per 100 reservations against the strongest rule baseline, a 48% reduction (Holm-corrected ), while deferring no request and churning less than that rule. A mechanism ablation against an uncertainty-blind twin of the same policy, which provably coincides with plain imitation greedy, isolates the information channel and accounts for the bulk of the effect. The result is model-based by construction, because the public snapshot cannot measure realised conflict at all, and we state it that way throughout.
We report the negatives with the same precision. Graph-attention embeddings hurt the imitation ranker, a validation-selected score blend reversed significantly on test, and masked discrete conservative Q-learning beat behaviour cloning but not the rule, so its hypothesis is not confirmed; for that agent we prove the conservative value bound survives masking for the policy-evaluation operator (Proposition 2), a result that motivates rather than certifies the deployed control learner. Three of four preregistered hypotheses did not confirm, under a protocol frozen before the preregistered evaluation and read exactly once in that round with every other test-set contact enumerated (Table 2). Reporting them under a protocol fixed in advance is what allows the confirmed result to be read as evidence rather than as selection.
Future Work
The most valuable next step is data, not method: weekly snapshots of the realised gate-event feed plus a re-download of the parking log once the windows overlap would supply realised occupancy times and convert every model-based conflict claim here into a realised-conflict claim, giving the forecaster a genuine target and making ECM checkable against outcomes rather than only against Monte-Carlo simulation of its own model. The multi-model and dependence sensitivity analyses are already reported in Section 5.8: the H-C ranking is preserved under a second calibrated duration model and under a day-level common shock, so the headline’s sign is not an artefact of the single model that defines the functional. Four method extensions follow from the limitations. First, start-time jitter remains open: Assumption 1 (A2) is the one convention whose relaxation changes the mathematics rather than only the numbers, since a random turns the outer integral into a convolution that no longer factorises segment by segment. Second, alternative tail families for the predictive distribution are a natural companion to it. Third, cross-fitted, out-of-fold graph embeddings would test whether the H-A negative is about graph structure or in-sample leakage. Fourth, the exploratory four-component blend, which reached on validation with but lay outside the registered functional form, can now be tested under its own preregistration.
Author Contributions
Conceptualization, T.Z. and J.G.; methodology, T.Z. and J.G.; software, T.Z.; validation, T.Z. and J.G.; formal analysis, T.Z. and J.G.; investigation, T.Z. and J.G.; resources, J.G.; data curation, T.Z.; writing—original draft, T.Z. and J.G.; writing—review and editing, T.Z. and J.G.; visualization, T.Z. and J.G.; supervision, J.G.; project administration, J.G. All authors have read and agreed to the published version of the manuscript.
Funding
This research received no external funding.
Data Availability Statement
All primary data are public. The four San Francisco International Airport tables (parking activity 5rkh-waic; stand inventory 2ymc-znns; tail-to-model reference u7dr-xm3v; realised gate events chfu-j7tc) are published on the DataSF open-data portal at https://data.sfgov.org (accessed on 23 July 2026) under the Open Data Commons Public Domain Dedication and Licence; hourly weather is NOAA Local Climatological Data v2 for station USW00023234, available at https://www.ncei.noaa.gov/ (accessed on 23 July 2026), and flight-level context is the Bureau of Transportation Statistics On-Time Reporting Carrier database, available at https://www.transtats.bts.gov/ (accessed on 23 July 2026); both are US public domain. The analysis code and the study protocol are available with the paper.
Conflicts of Interest
The authors declare no conflicts of interest.
Appendix A. Derivations and Proofs
Appendix A.1. The Expected-Conflict Functional in Closed Form
Notation follows Definition 3 and Assumption 1. Step 1. The proof in Section 2.6 establishes, by the layer-cake identity, Tonelli’s theorem and (A1) and (A2),
for a pair on the same stand with . Definition 3 gives for and for , so the integrand vanishes for and the upper limit may be truncated there; if the integral is empty and . That is the exact pruning rule used by the implementation, exact because the tails of Definition 3 are bounded rather than asymptotic.
Step 2: both factors are affine between consecutive breakpoints. Let be the knot abscissae of after merging coinciding knots (Definition 3), and similarly for j. On each open interval the CDF is affine, hence so is ; outside it is constant, one below and zero above, and merged knots carry atoms at which jumps, the knot value being post-jump since is right-continuous. The affine shifts of (A1) move those breakpoints to and . Let
so that and . On each open segment both factors are affine, and every jump point of either is an endpoint of a segment, never an interior point; in particular the midpoint is always a point of affinity of both, which makes the expansion below legitimate even in the presence of atoms.
Step 3: exact integration of each segment. Write and expand each factor about the midpoint,
where are the survival values at and the slopes, both nonpositive as negatives of piecewise-linear CDF slopes. Their product is the quadratic
Integrating over the segment, which is symmetric about , and using
the cross term vanishes by symmetry and
Summing over gives (12). Every step is an identity, so the result carries no discretisation error: refining the partition beyond (A2) changes nothing, and coarsening it below the breakpoint set is the only way to introduce error.
Remark A1
(Fidelity to the implementation). Four details make the theorem statement and the code describe the same object. Quantile preprocessing: predicted quantiles are clipped below at one minute and sorted ascending when the raw predictions invert, before the knots of (9) are formed, and coinciding knots merge keeping the larger CDF value, turning that knot into an atom; the derivation is unaffected, since merged knots remain elements of (A2). Fallback distributions: an event lacking a finite predicted quantile triple gets a point mass at its reserved duration, contributing the deterministic buffered overlap by Corollary 2; no such event occurs on validation or test. Pair set and attribution: the per-decision marginal used as policy risk term and offline-RL reward drops the day-attribution filter of Definition 4, because every pair a decision forms is caused by it; summing the marginals over a day reproduces the schedule-level total exactly, an identity we verify directly. Two buffer conventions: the reservation-based overlap metric retained from the earlier round inflates both intervals by , whereas Theorem 1 inflates only the earlier end by β (A4); the two are not comparable and no comparison between them appears here.
Appendix A.2. Monte-Carlo Validation of the Closed Form
The closed form was validated against simulation before any policy iteration was logged, as a correctness gate. Thirty validation days were sampled with a fixed generator from the 731 eligible days (those with at least one candidate pair under the historical schedule), each simulated with inverse-CDF draws from the distributions of Definition 3. Each draw evaluates the realised one-sided buffered overlap over exactly the pair set the analytical evaluation integrates, so the sample mean is an unbiased estimator of (A1) and any systematic gap indicates an implementation error, not a modelling one. Writing A for a day’s analytical value and M for its Monte-Carlo mean, the preregistered tolerance required on at least 28 of the 30 days, plus an aggregate relative difference of at most 0.01. The gate passed on its first run: 30 of 30 days inside tolerance, aggregate relative difference 0.00159. Worst of the thirty was 29 January 2023, analytical 11,101.48 against simulated 10,732.67 (Monte-Carlo standard error 161.18), a gap equal to 0.76 of its own tolerance.
What this gate does and does not establish deserves stating plainly. It is a validation of the numerical implementation: it confirms that the closed-form segment sum of Theorem 1 evaluates the same integral that direct inverse-CDF sampling from the distributions of Definition 3 produces, to within the stated tolerance and with no quadrature error. It is not an empirical validation of the model. It says nothing about whether those predictive distributions describe actual stand occupancy, because the sampler draws from the same distributions the closed form integrates; agreement between them is a statement about arithmetic, not about the world. Empirical validation of the predictive family would require a realised-occupancy feed that this data source does not provide, and Section 6.5 records it as such.
Appendix A.3. Conservatism Under Masking: Proof of Proposition 2
We adapt the support-restriction argument of Kumar et al. [11], Theorem 3.2, replacing by the state-dependent feasible set throughout. The adaptation is not merely notational, since the CQL(H) regulariser is a log-sum-exp over the action set and restricting that set changes the maximiser and hence the penalty; we verify that the lower-bound structure survives and, following the scope of that theorem, that it holds when the resulting is evaluated under its own induced softmax rather than under an arbitrary masked policy.
- Masked objective.
Let be the logged dataset with empirical behaviour policy , and assume for every and every s appearing in (the strengthened support condition of Proposition 2; see Remark 2 for its applicability under the behaviour-cloning-distilled prior). Constrain for and consider the policy-evaluation iteration for the target
The first bracket is the entropy-regularised maximisation , whose maximiser is the softmax of over ; let be the policy-evaluation Bellman operator for . Differentiating (A7) in for and setting the derivative to zero gives, for every state-action pair in the data,
the update of [11] with in place of . The strengthened assumption makes for every feasible a, so is finite throughout and the update is well defined; for the constraint forces .
- The penalty is nonnegative in μ-expectation.
Fix s. Evaluating the bracket of (A8) in expectation under gives
because by the Cauchy–Schwarz inequality and both marginal sums equal one on under the support assumption, with equality if and only if there. This nonnegativity is what makes the term conservative: it lowers on actions favours relative to . The step is specific to : for a general masked the analogous sum can be negative, for instance if concentrates on an action under-weights relative to , so the bound is claimed at only.
- Fixed point and the bound.
Iterating (A8) and using that is a -contraction on the masked value functions, the limit satisfies
for , with the state-action transition operator induced by , restricted to feasible actions. We take the -expectation of both sides explicitly, the subtlety being the passage from the per-state (A9) to the value bound. Define the state-level penalty
nonnegativity being (A9). Averaging a state-action function under intertwines with the state-to-state kernel : for any h,
with . Applying this to the Neumann series term by term gives . Taking the -expectation of (A10) and using ,
which is (8): the inequality holds because is a nonnegative operator, a discounted sum of stochastic kernels, applied to the pointwise-nonnegative , so the subtracted term is at every state. Actions outside carry and are excluded because is supported on ; a masked greedy policy never selects them, so the bound is exactly the quantity such a policy acts on. Without sampling error the argument holds for every ; with finite data must exceed the policy-dependent threshold of [11], Theorem 3.2, instantiated at with its concentrability coefficient over rather than . It is not a universal constant, and is vacuous at , where .
- Scope.
Proposition 2 concerns the masked CQL(H) fixed point evaluated at its own induced softmax ; it is not a theorem about the deployed control learner and not a claim about the empirical policy’s quality (Remark 2). Training uses the reachable subset only, 95.53% of test events having their logged action inside the mask. The full-support condition is an idealisation the empirical data does not satisfy, the logged behaviour being a near-point-mass; the behaviour-cloning warm start smooths the initial Q-estimate, not the data-generating , so it does not repair the support defect and we do not claim it does. The proposition certifies the idealised operator and formalises the conservatism mechanism; it does not certify the deployed learner, whose achievement, including the comparison it loses, is reported in Section 5.5.
References
- Li, Y.; Clarke, J.P.; Dey, S.S. Using Submodularity within Column Generation to Solve the Flight-to-Gate Assignment Problem. Transp. Res. Part Emerg. Technol. 2021, 129, 103217. [Google Scholar] [CrossRef] [Scilit]
- Daş, G.S.; Gzara, F.; Stützle, T. A Review on Airport Gate Assignment Problems: Single Versus Multi Objective Approaches. Omega 2020, 92, 102146. [Google Scholar] [CrossRef] [Scilit]
- Bouras, A.; Ghaleb, M.A.; Suryahatmaja, U.S.; Salem, A.M. The Airport Gate Assignment Problem: A Survey. Sci. World J. 2014, 2014, 923859. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Li, H.; Wu, X.; Ribeiro, M.; Santos, B.; Zheng, P. Deep Reinforcement Learning Approach for Real-Time Airport Gate Assignment. Oper. Res. Perspect. 2025, 14, 100338. [Google Scholar] [CrossRef] [Scilit]
- Ding, C.; Bi, J.; Wang, Y. A Hybrid Genetic Algorithm Based on Imitation Learning for the Airport Gate Assignment Problem. Entropy 2023, 25, 565. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Ali, H.; Dönmez, K.; Lim, W.L.; Alam, S. Machine Learning Algorithms and Models for Airport Gate Assignment Problem: A Systematic Literature Review. Transp. Res. Part Logist. Transp. Rev. 2026, 209, 104734. [Google Scholar] [CrossRef] [Scilit]
- Wang, R.; Allignol, C.; Barnier, N.; Gondran, A.; Gotteland, J.B.; Mancel, C. A New Multi-Commodity Flow Model to Optimize the Robustness of the Gate Allocation Problem. Transp. Res. Part Emerg. Technol. 2022, 136, 103491. [Google Scholar] [CrossRef] [Scilit]
- Zhang, C.; Jin, Z.; Ng, K.K.H.; Tang, T.Q.; Zhang, F.; Liu, W. Predictive and Prescriptive Analytics for Robust Airport Gate Assignment Planning in Airside Operations under Uncertainty. Transp. Res. Part Logist. Transp. Rev. 2025, 195, 103963. [Google Scholar] [CrossRef] [Scilit]
- Altman, E. Constrained Markov Decision Processes: Stochastic Modeling; Routledge: London, UK, 2021. [Google Scholar] [CrossRef] [Scilit]
- Huang, S.; Ontañón, S. A Closer Look at Invalid Action Masking in Policy Gradient Algorithms. In Proceedings of the Thirty-Fifth International Florida Artificial Intelligence Research Society Conference (FLAIRS), Jensen Beach, FL, USA, 15–18 May 2022. [Google Scholar] [CrossRef] [Scilit]
- Kumar, A.; Zhou, A.; Tucker, G.; Levine, S. Conservative Q-Learning for Offline Reinforcement Learning. In Proceedings of the Advances in Neural Information Processing Systems 33 (NeurIPS), Virtual, 6–12 December 2020. [Google Scholar]
- Ke, G.; Meng, Q.; Finley, T.; Wang, T.; Chen, W.; Ma, W.; Ye, Q.; Liu, T.Y. LightGBM: A Highly Efficient Gradient Boosting Decision Tree. In Proceedings of the Advances in Neural Information Processing Systems 30 (NIPS), Long Beach, CA, USA, 4–9 December 2017. [Google Scholar]
- Koenker, R.; Bassett, G., Jr. Regression Quantiles. Econometrica 1978, 46, 33–50. [Google Scholar] [CrossRef] [Scilit]
- Romano, Y.; Patterson, E.; Candès, E.J. Conformalized Quantile Regression. In Proceedings of the Advances in Neural Information Processing Systems 32 (NeurIPS), Vancouver, BC, Canada, 8–14 December 2019. [Google Scholar]
- Gneiting, T.; Raftery, A.E. Strictly Proper Scoring Rules, Prediction, and Estimation. J. Am. Stat. Assoc. 2007, 102, 359–378. [Google Scholar] [CrossRef] [Scilit]
- Chen, T.; Guestrin, C. XGBoost: A Scalable Tree Boosting System. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, San Francisco, CA, USA, 13–17 August 2016; pp. 785–794. [Google Scholar] [CrossRef] [Scilit]
- Wu, Q.; Burges, C.J.C.; Svore, K.M.; Gao, J. Adapting Boosting for Information Retrieval Measures. Inf. Retr. 2010, 13, 254–270. [Google Scholar] [CrossRef] [Scilit]
- Burges, C.J.C. From RankNet to LambdaRank to LambdaMART: An Overview; Technical Report MSR-TR-2010-82; Microsoft Research: Redmond, WA, USA, 2010. [Google Scholar]
- Veličković, P.; Cucurull, G.; Casanova, A.; Romero, A.; Liò, P.; Bengio, Y. Graph Attention Networks. In Proceedings of the 6th International Conference on Learning Representations (ICLR), Vancouver, BC, Canada, 30 April–3 May 2018. [Google Scholar]
- Busbridge, D.; Sherburn, D.; Cavallo, P.; Hammerla, N.Y. Relational Graph Attention Networks. arXiv 2019, arXiv:1904.05811. [Google Scholar]
- Kostrikov, I.; Nair, A.; Levine, S. Offline Reinforcement Learning with Implicit Q-Learning. In Proceedings of the Tenth International Conference on Learning Representations (ICLR), Virtual, 25–29 April 2022. [Google Scholar]
- Efron, B.; Tibshirani, R.J. An Introduction to the Bootstrap; Chapman & Hall/CRC: New York, NY, USA, 1994. [Google Scholar] [CrossRef] [Scilit]
- Künsch, H.R. The Jackknife and the Bootstrap for General Stationary Observations. Ann. Stat. 1989, 17, 1217–1241. [Google Scholar] [CrossRef] [Scilit]
- Wilcoxon, F. Individual Comparisons by Ranking Methods. Biom. Bull. 1945, 1, 80–83. [Google Scholar] [CrossRef] [Scilit]
- Cliff, N. Dominance Statistics: Ordinal Analyses to Answer Ordinal Questions. Psychol. Bull. 1993, 114, 494–509. [Google Scholar] [CrossRef]
- Holm, S. A Simple Sequentially Rejective Multiple Test Procedure. Scand. J. Stat. 1979, 6, 65–70. [Google Scholar]
- Levine, S.; Kumar, A.; Tucker, G.; Fu, J. Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems. arXiv 2020, arXiv:2005.01643. [Google Scholar]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.


