1. Introduction
Tourism-demand forecasting has long been recognized as essential for destination planning, resource allocation, and strategic decision-making [
1,
2]. Accurate forecasts enable destination marketing organizations (DMOs) to optimize promotional spending across source markets, allow hospitality businesses to adjust pricing, staffing, and service design, and help destination analysts monitor shifts in market mix [
3,
4]. While the tourism-forecasting literature is extensive, the vast majority of studies focus on predicting aggregate arrivals or total expenditure [
5,
6]. Considerably less attention has been paid to forecasting the
composition of demand, that is, how the mix of visitors from different origin markets evolves over time.
To see why origin composition matters independently of aggregate volume, consider a DMO setting its source-market campaign budgets for the coming year. The relevant question is not only how many visitors may ultimately arrive, but what the evolving booking mix suggests about where they are likely to come from, because the allocation decision must be made months before stays materialize. A forecast that the German share of current booking flows is trending upward while the UK share is declining for a Mediterranean destination justifies redirecting spend toward German-market campaigns, travel fair partnerships, and airline co-marketing before the season opens, not after it closes.
Two further decisions have the same structure. Origin composition is a measure of strategic risk, because a destination whose booking mix concentrates in a single source market accumulates fragility that aggregate volume forecasts conceal, as the collapse of Chinese-origin bookings into the Asia-Pacific made concrete. It also drives operational planning, because staffing, amenity, and service decisions are difficult to reverse mid-season.
Section 6 develops all three in detail against the empirical results.
In this paper, the forecast target is the monthly composition of reservations by guest origin indexed by booking date, not the composition of realized arrivals or stays indexed by check-in month. That distinction matters. Booking-month composition can serve as a leading indicator for destination managers because it is observed earlier than realized arrivals, but it also reflects changes in booking timing and platform usage. We therefore interpret the series as booking demand on Airbnb, not as a census of realized inbound tourism demand.
The composition of booking demand also presents a distinct methodological challenge. Visitors from different source markets exhibit distinct spending patterns, budget allocations, and demand sensitivities [
7,
8], and geopolitical events, exchange rate fluctuations, and public health crises affect origin markets asymmetrically [
9,
10]. Capturing these dynamics requires methods that respect the constraints inherent in compositional data: origin market shares must sum to unity and each share is bounded between zero and one. Standard time-series methods that ignore these constraints can produce incoherent forecasts, for example, predicted shares that sum to more than 100% or take negative values [
11,
12].
Two families of methods respect the constraint, and the difference between them is where the coherence comes from. The compositional data analysis literature, pioneered by Aitchison [
13], maps the simplex to unconstrained Euclidean space through log-ratio transformations, applies standard multivariate machinery there, and maps back. This works, and two of our benchmarks are built exactly this way. The approach has costs. Because the back-transformation is nonlinear, a symmetric interval in log-ratio coordinates is asymmetric on the simplex, interpretation in terms of shares is indirect, and performance can degrade when shares approach zero [
14]. In the standard formulations used as benchmarks below, predictive dispersion follows from the transformation and the fitted error variance rather than from a separately specified process, though richer transformed families can parameterize covariance directly.
The alternative places a distribution directly on the simplex. The Dirichlet is the canonical choice, and three of its properties motivate its use here. Its support and the sample space of the data coincide, so every posterior predictive draw is a valid composition and no post hoc normalization is required. Its concentration parameter is a free quantity, which means predictive dispersion can be modeled rather than implied, and it is precisely this feature that permits the seasonal precision specification we introduce in
Section 3. And because the model is specified on the compositions themselves, its parameters retain an interpretation in terms of shares. We apply the Bayesian Dirichlet Autoregressive Moving Average (BDARMA) framework of Katz et al. [
15] to this setting;
Section 2 reviews the relevant methodological background.
This paper contributes to both the tourism-forecasting and compositional time-series literatures by developing and applying BDARMA models to forecast guest origin market shares in large-scale platform booking data. We analyze Airbnb reservations from 2017 to 2025 across four major destination regions (EMEA, North America, Asia-Pacific, and Latin America), providing what is, to our knowledge, the first application of Bayesian compositional time-series methods to large-volume platform booking-origin shares across multiple global destination regions.
Our empirical setting offers several advantages. First, Airbnb operates globally with standardized data collection, enabling consistent measurement across diverse markets. Second, platform booking data capture actual reservations rather than survey-based intentions or aggregate border-crossing statistics, providing a direct measure of revealed booking demand. Because reservations are indexed by booking date rather than check-in date, the series provide an earlier signal of changing market mix, even though they are not identical to realized stay-month demand. Third, our sample period spans the COVID-19 pandemic and subsequent recovery, allowing us to examine how compositional dynamics shifted during and after this unprecedented disruption. Prior work using Airbnb data has examined other dimensions of pandemic-era booking behavior, but the evolution of origin market composition has not been systematically analyzed.
We find that BDARMA models achieve the lowest forecast error for EMEA, the most compositionally diverse destination region, while simpler benchmarks remain difficult to beat in the other three regions, a pattern consistent with a long-standing finding in the tourism-forecasting literature. The pandemic caused dramatic compositional shifts in booking flows, most notably a surge in within-region booking shares at the expense of long-haul markets, with recovery trajectories that varied markedly across destination regions. Our models capture these dynamics through autoregressive and moving average terms that allow past compositions and shocks to influence current shares, while the Dirichlet likelihood ensures forecasts remain valid probability distributions.
Seasonal variation in the precision parameter improves point accuracy in every region, which we establish in
Section 5.4 through a direct comparison against an otherwise identical constant-precision model.
The remainder of this paper is organized as follows.
Section 2 reviews related work on tourism-demand forecasting, compositional data analysis, and peer-to-peer accommodation research.
Section 3 presents the BDARMA modeling framework and estimation approach.
Section 4 describes our Airbnb booking data and the construction of origin market compositions.
Section 5 reports empirical findings, including model comparisons and forecast accuracy assessments.
Section 6 discusses implications for destination management and directions for future research.
Section 7 concludes.
4. Data
4.1. Airbnb Booking Data and Forecast Target
Our analysis uses reservation data from Airbnb’s global platform spanning January 2017 through December 2025. Airbnb is one of the world’s largest peer-to-peer accommodation marketplaces, operating in over 220 countries and regions with more than 9 million active listings [
48]. The platform’s standardized booking infrastructure enables consistent measurement of guest origin and destination across diverse markets.
Reservations are extracted from Airbnb’s internal bookings data warehouse, which records one row per reservation event with the booking date, the guest’s registered country of residence, and the listing’s destination country and region. We restrict to confirmed reservation events, excluding cancellations, and aggregate to monthly frequency by booking month, destination region, and guest origin country. The full sample comprises 108 monthly observations per destination region. Guests are assigned to origin countries by registered country of residence at the time of booking, which is the platform’s own account-level field rather than an inferred location.
Cancellation status is applied ex post: a reservation booked in month t and subsequently canceled is excluded from month t’s composition, so the historical series reflect final booking status rather than the exact information set available in real time, and the leading-indicator interpretation developed below should be read accordingly.
The time index in our analysis is the month in which the reservation is
booked, not the month of check-in or realized stay. Accordingly, each composition in the paper should be read as the share of bookings made in month
t that comes from each origin market. This makes the series useful as a forward-looking indicator of demand mix because bookings are observed before stays occur, but it also implies that the object of interest is booking composition rather than realized arrival-month tourism demand. When booking lead times shift, as they did during and after COVID-19 [
43], booking-month and stay-month compositions need not move one-for-one.
The reservation-level data are proprietary to Airbnb and cannot be released. We return to the consequences for reproducibility in
Section 6.4, and the Data and Code Availability statement sets out what is available.
4.2. Geographic Scope
We analyze bookings into four major destination regions defined by Airbnb’s financial planning and analysis (FPA) taxonomy:
EMEA: Europe, Middle East, and Africa, Airbnb’s largest region by booking volume, encompassing major destinations such as France, Spain, Italy, the United Kingdom, and Germany.
NAMER: North America, comprising the United States and Canada, with the U.S. representing the platform’s founding market.
APAC: Asia-Pacific (excluding mainland China), including Australia, Japan, South Korea, and southeast Asian markets.
LATAM: Latin America, spanning Mexico, Brazil, Argentina, and other Central and South American countries.
Three considerations motivate working at this level of aggregation rather than at country or city level. The regional taxonomy is the unit at which the platform’s own demand planning and financial reporting operate, so the compositions we forecast correspond to a level at which allocation decisions are actually taken, and a forecast that does not map onto a decision unit is of limited operational use. Disaggregation also multiplies the number of series while thinning the counts within each: monthly booking volumes for a given origin into a single city are small enough that the strictly positive shares required by both the Dirichlet likelihood and the ILR transformation are no longer guaranteed, whereas at the regional level, the minimum observed share is 0.8% with no exact zeros (
Section 4.4). Finally, the regional level supports the cross-regional comparison that is central to our findings, since it is the contrast between compositionally diverse EMEA and concentrated NAMER that identifies where direct compositional modeling helps.
The cost of this choice is real, and we treat it as a limitation in
Section 6.4: regional aggregation obscures within-region heterogeneity that matters to local destination managers.
Section 6.5 discusses finer granularity as an extension, where hierarchical structure would be needed to borrow strength across sparse series.
For each destination region, we compute the monthly composition of guest origin markets.
4.3. Origin Market Classification
To obtain interpretable compositions, we consolidate origin countries into a manageable number of categories. For each destination region, we identify the top seven origin markets by average booking share over the sample period. Markets outside the top seven are aggregated into an “Other” category. This yields eight distinct origin categories per destination region. The same seven markets are obtained in every region when selection is restricted to the pre-evaluation window (January 2017 through December 2021), so the category definitions do not import information from the evaluation period.
The “Other” category is material rather than residual, averaging above 20% in several regions, and it pools origin markets whose dynamics need not resemble one another.
Section 6.4 treats the consequences.
For interpretive convenience, we define a booking as within-region if the guest’s origin country falls within the same destination region (e.g., a French guest booking in Spain, both within EMEA), and outside-region otherwise. This distinction serves as a proxy for short-haul versus long-haul travel patterns in our subsequent analysis.
Table 2 presents descriptive statistics for each destination region, including the number of observations, origin market components, and summary statistics for compositional variability.
4.4. Descriptive Patterns
Figure 2 displays the evolution of origin market compositions over our sample period for each destination region. Several patterns emerge. First, compositions exhibit substantial temporal variation, with visible seasonal patterns (e.g., increased European-origin booking share into EMEA during summer months) and trend shifts. Second, the COVID-19 pandemic (beginning March 2020) caused dramatic compositional changes: within-region origin shares surged as long-haul booking activity collapsed. Third, the recovery from pandemic lows has been uneven across origin markets, with some shares returning to pre-pandemic levels while others show persistent deviations. Because the figure displays shares rather than counts, these movements describe changes in the booking mix; a rising share can reflect growth in that origin, contraction elsewhere, or both.
The high lag-1 autocorrelations of the raw share series (0.87–0.94 across regions;
Table 2) confirm substantial persistence in compositional shares, motivating autoregressive modeling. APAC exhibits the most dramatic compositional variation, with the Chinese-origin booking share dropping from approximately 25% pre-pandemic to under 5% during restrictions, with gradual recovery thereafter. NAMER shows the least compositional variation, with U.S.-origin bookings consistently dominating at 75–90% of the total.
Both the Dirichlet likelihood and ILR transformation require strictly positive compositional components. In our data, the minimum observed share across all region-month-origin combinations is 0.8%, occurring for the “Other” category in NAMER during peak U.S.-origin months. No exact zeros appear in the data, a consequence of aggregating large booking volumes (tens of thousands of reservations per month) where even small origin markets contribute positive counts. We therefore apply no zero-replacement or smoothing procedures.
4.5. Compositional Dynamics
To motivate our modeling choices, we examine temporal patterns in compositional variability.
Figure 3 plots the Herfindahl–Hirschman Index (HHI) of origin market concentration over time for each destination region. NAMER exhibits consistently high concentration (HHI ≈ 0.60–0.75), reflecting U.S. dominance, with a pandemic-induced spike to nearly 0.90 as Canadian cross-border travel collapsed. In contrast, EMEA, APAC, and LATAM maintain diverse origin portfolios (HHI ≈ 0.20) with only modest pandemic disruption.
Throughout, the index is computed over the eight modeled origin buckets. Because the “Other” category pools many small origins, the resulting HHI overstates concentration relative to a fully disaggregated country-level calculation: for positive shares
a and
b,
. The index should therefore be read as concentration in the modeled composition rather than as a country-level concentration measure. This matters for the concentration-risk application discussed in
Section 6.2, and we return to it in
Section 6.4. The especially high concentration of NAMER is robust to this consideration, but finer comparisons among the remaining three regions should be interpreted cautiously.
This heterogeneity in market structure explains why forecast method performance varies across regions: BDARMA’s ability to model compositional dynamics provides greater value where multiple origin markets compete.
Figure 4 displays average autocorrelation functions of CLR-transformed shares by destination. All regions exhibit substantial persistence at short lags, with NAMER showing the slowest decay (lag-1 ACF ≈ 0.90) and EMEA the fastest (lag-1 ACF ≈ 0.70). APAC and LATAM display a secondary peak at lag 12, indicating seasonal patterns in composition. These autocorrelation structures motivate the AR(1) specification in our BDARMA models.
Critically, compositional variability itself exhibits seasonal patterns.
Figure 5 shows boxplots of the Aitchison distance from the mean composition by calendar month for EMEA. Spring and early summer months (March–June) display substantially higher compositional dispersion than autumn months (September–October), with median Aitchison distances approximately 25% larger. This pattern reflects greater uncertainty in origin mix during shoulder seasons when booking patterns are less predictable compared to peak summer months when established seasonal flows dominate.
This empirical pattern motivates the seasonal precision specification of Equation (
4): rather than assuming constant compositional volatility, we allow the Dirichlet precision parameter
to vary with Fourier seasonal terms.
Section 5.4 evaluates whether that structure earns its place.
6. Discussion
6.1. Summary of Findings
Our analysis demonstrates that BDARMA models achieve the lowest forecast error for EMEA and competitive performance across other destination regions, with particular strength where multiple-origin markets exhibit rich compositional dynamics. For EMEA, BDARMA achieves 27% lower forecast error than naïve methods, with highly significant improvement (). Averaged across all destinations, ETS slightly edges BDARMA in MAE, 0.0086 versus 0.0089, an average achieved under a protocol that permits the benchmarks to re-select their structure at every forecast origin while the BDARMA orders remain fixed. The model has lower error than seasonal naïve in EMEA, NAMER, and APAC, reflecting its ability to adapt to regime changes while still capturing underlying temporal patterns. All three comparisons reach the unadjusted 5% threshold, and the NAMER and APAC comparisons remain significant after Holm adjustment across the full family of tests. Throughout, these forecasts concern booking-month composition rather than realized stay-month arrivals, so they are best interpreted as leading indicators of market mix.
Seasonal precision earns its place in the specification. The ablation in
Section 5.4 compares the selected model against an otherwise identical constant-precision version and finds lower point error in all four regions, ranging from 13% in LATAM to 35% in APAC, with seasonal precision producing the lower error at between 10 and 13 of the 14 forecast origins in each region. The mechanism is the one the descriptive evidence suggested: compositional volatility differs systematically across months, spring and early summer showing greater dispersion than autumn, and a model with constant precision has no way to represent that. The comparison establishes point accuracy rather than calibration; a full probabilistic assessment of the two specifications remains open and is identified as a near-term priority in
Section 6.5.
The BDARMA(1,1) specification achieves the best in-sample fit for three of four destination regions (BDARMA(2,1) is preferred for EMEA), as measured by LOO-CV, indicating that both autoregressive and moving average components capture important features of compositional dynamics. The relative advantage of BDARMA varies systematically with the nature of compositional variation: it excels where multiple origin markets compete with shares in the 5–25% range (EMEA), performs comparably to simpler methods where one market dominates (NAMER), is outperformed by ETS in settings with smooth trend dynamics (LATAM), and is outperformed by naïve extrapolation where a single persistent recovery path dominates the signal (APAC). As shown in
Section 4, the HHI concentration analysis (
Figure 3) helps explain this pattern: BDARMA provides greatest value for destinations with diverse origin portfolios where compositional dynamics are meaningful.
The convergence with Gunter et al. [
35] deserves emphasis. Working with occupancy levels rather than origin composition, across 43 European countries, and with a seasonal Markov-switching autoregression benchmarked against panel-data counterfactuals, they arrive at a conclusion our pandemic-era shift analysis supports from the compositional side: forecasting through the pandemic rewarded specifications that pair regime adaptation with explicit seasonal structure. Two independent routes to the same requirement make it more credible than either alone. In our framework, the adaptation is smooth, carried by the autoregressive terms. Abrupt shifts can instead be handled through gated intervention structure [
49], and a Markov-switching Dirichlet ARMA in which latent regimes govern the ILR-level mean or the precision would be the direct compositional counterpart of their approach.
6.2. Implications for Destination Management
Our findings have three concrete implications for destination marketing organizations and tourism planners, corresponding to the business use cases motivating this study. We develop the first as a concise implementation workflow that links the forecast to the timing and information requirements of a decision, while treating the forecast as one input to that decision rather than as an automatic allocation rule.
Marketing budget allocation across source markets.
Consider a regional tourism-planning team responsible for source-market strategy across EMEA. At a regular planning date, the team would first generate one- to six-month forecasts of origin market composition and compare the projected changes with historical forecast error and posterior predictive uncertainty. A sufficiently large and persistent projected change would trigger managerial review rather than automatic reallocation. The team would then combine the composition forecast with campaign costs, expected incremental returns, operational constraints, and a separate forecast of aggregate booking volume before adjusting source-market allocations. The volume forecast is essential because a rising origin share may reflect growth in that origin, contraction elsewhere, or both.
The EMEA results provide the clearest empirical basis for this workflow: BDARMA reduces MAE by 27% relative to the naïve benchmark (
Table 4). In NAMER, APAC, and LATAM, where simpler methods are competitive or superior, the same process should use the method with the strongest rolling-origin performance rather than assume that greater model complexity improves the decision. Predictive intervals can inform the review, but because their calibration is not formally assessed here, they should not serve as mechanical commitment thresholds. Because the forecast target is booking composition rather than stay-month arrivals, evaluation against realized visitor mix must also account for booking-to-stay lead-time drift [
43].
Concentration risk monitoring.
The pandemic provided a stark illustration of the fragility that builds when origin portfolios become concentrated. As documented in
Section 5.8, the Chinese-origin share of bookings into APAC fell by more than 20 percentage points within a few months. Aggregate volume forecasts alone do not reveal this form of exposure because they do not describe how demand is distributed across origin markets. BDARMA forecasts can be used alongside HHI trajectories to monitor baseline concentration exposure and identify persistent increases in dependence on a small number of source markets. This is a monitoring application rather than scenario-based stress testing, since the current model contains no explicit shock, intervention, or policy mechanism.
The caveat noted in
Section 4 applies operationally as well as descriptively. The HHI is computed over the modeled origin buckets, and the aggregation of many countries into the “Other” category means that its level differs from a fully disaggregated country-level index. Its trajectory is therefore more informative for monitoring than any fixed concentration threshold.
Operational planning at the property level.
The regional forecasts estimated here are not property-specific. After extension to country- or city-level series, likely using hierarchical structure to borrow strength across sparse destinations, the same framework could support staffing, language capability, amenity, and service-design decisions. For example, a projected increase in Korean-origin share, interpreted jointly with expected total booking volume, could inform Korean-language staffing or service preparation before the travel period begins.
The seasonal-precision ablation in
Section 5.4 shows that allowing compositional dispersion to vary across months improves point accuracy in every region. It does not, however, establish that the posterior predictive intervals attain their nominal coverage. The intervals should therefore be treated as indicative measures of uncertainty rather than as formal rules governing operational commitments.
These applications are most directly supported at the one- to six-month horizons evaluated in this study; longer-horizon strategic uses require separate validation. The heterogeneity in forecast performance also provides practical guidance about model complexity. BDARMA appears most useful in compositionally diverse markets where several origins hold material shares, whereas simpler methods remain competitive or superior where one origin dominates or where smooth trend dynamics carry most of the forecast signal.
6.3. Methodological Contributions
This study makes three contributions to the tourism-forecasting literature, corresponding to those stated in the abstract. We give each as a claim followed by the reason it holds, rather than resting on measured performance alone.
Contribution 1: Bayesian compositional time-series methods applied to booking-origin composition. The Dirichlet likelihood is appropriate to this problem for a structural reason rather than a convenient one: the support of the distribution and the sample space of the data coincide. Origin shares are non-negative and sum to one, which is exactly the unit simplex. Although the mean dynamics are parameterized in ILR coordinates, the observation likelihood and the posterior predictive distribution are supported directly on the simplex, so no post hoc normalization is required and every predictive draw is a valid composition. This differs from the ETS and SARIMA benchmarks used here, which fit Euclidean forecasting models to transformed coordinates and recover composition-valued forecasts only through inverse transformation.
The empirical case for this contribution rests on rolling-origin comparison against operational benchmarks, including log-ratio ETS and SARIMA with automatic order selection refitted at every origin. We report that comparison in full rather than selectively: BDARMA achieves the lowest error for EMEA and significantly outperforms naïve and ETS there, while its advantage over seasonal naïve reaches the unadjusted 5% threshold but not the Holm-adjusted threshold, and ETS attains the lowest error averaged across regions. Reporting the cases where simpler methods win is part of the point, because it identifies where compositional modeling is warranted and where it is not.
Contribution 2: Seasonal structure in the precision parameter, with its contribution isolated. The Dirichlet concentration parameter governs how tightly realizations cluster around the mean composition, and there is no principled reason to hold it constant when the dispersion of the data itself varies across the calendar.
Section 4 documents that variation: the median Aitchison distance from the mean composition is roughly 25% larger in spring and early summer than in autumn. A constant-precision model must absorb that heterogeneity somewhere, and
Section 5.4 measures what it costs when it cannot, finding higher point error in all four regions. The Dirichlet precision supplies a parsimonious and explicitly modeled dispersion process that the ETS and SARIMA benchmarks used here do not contain, since their predictive spread follows from the transformation and the fitted error variance rather than from a separately specified process. We note that other transformed families, logistic-normal state-space models among them, can parameterize time-varying covariance directly, so the contrast is with the specific benchmarks employed rather than with log-ratio methods in general.
Contribution 3: Booking-date indexing as a substantive choice. Indexing compositions by booking month rather than stay month makes the series observable ahead of realized arrivals, which is what allows the forecasts to inform decisions with long lead times. This is not a data-availability convenience. It changes what the forecast is for, and it introduces a linkage problem, since booking-month and stay-month compositions need not move together when lead times shift [
43].
6.4. Limitations
Several limitations should be acknowledged, and we group them by what they constrain.
Data access and reproducibility. The reservation-level data underlying this study are proprietary to Airbnb and cannot be released. We mitigate it where we can: the estimation software is publicly available, the methodology is fully specified in
Section 3, and the analysis code operates on any compositional series in the same layout, so the method is reproducible even where the empirics are not.
External validity. Our data capture Airbnb bookings, which represent a subset of total accommodation demand. Origin compositions may differ in both level and dynamics for hotel guests, for visitors staying with friends and relatives, and for other booking channels. Conclusions about which forecasting method suits which market structure should generalize more readily than the specific compositional trajectories we document.
Forecast horizon. We evaluate at horizons of one to six months, chosen to match the lead times of the marketing allocation decisions that motivate the study. Many tourism planning decisions operate on annual or multi-year cycles: market development strategy, language capability investment, and partnership structure among them. Longer-horizon performance is unknown, and extending the evaluation would require attention to whether the Dirichlet precision remains adequately specified as predictive uncertainty grows, since the concentration parameter governs how quickly the predictive distribution flattens.
Exogenous drivers not included. The specification estimated in this study uses lagged compositions and seasonal terms but does not include external predictors such as exchange rates, airline capacity, visa policy, fuel prices, or geopolitical indicators. Consequently, the analysis forecasts compositional dynamics without attributing them to particular external drivers or evaluating scenarios based on alternative paths of those variables. This reflects the covariate set used in the present application, not a limitation of the BDARMA framework, which can incorporate exogenous regressors through the covariate matrix.
Pandemic dominance of the sample. The 2017–2025 window is dominated by COVID-19 and its aftermath, and the evaluation period beginning in January 2022 falls entirely within the recovery. The advantages we document are therefore advantages under regime adaptation. Whether direct compositional modeling retains an edge over simple benchmarks in a stable tourism environment cannot be established from these results. The convergence with Gunter et al. [
35], whose Markov-switching specification performed well precisely because the period rewarded regime adaptation, is consistent with this reading.
The “Other” aggregation. Consolidating markets outside the top seven into a single category simplifies estimation but pools origins whose dynamics need not resemble one another and may offset. Movement within the Other bucket is invisible to the model, and a stable Other share can conceal substantial reallocation among its constituents. This has a specific consequence for concentration-risk monitoring: the HHI reported in
Section 4 is computed over the eight modeled buckets and therefore overstates concentration relative to a full country-level calculation. Since the Other category averages above 20% in several regions, the discrepancy is not negligible.
Aggregation to regions. Regional aggregation obscures within-region heterogeneity that may matter to local destination managers. A forecast for EMEA composition is of limited direct use to a manager in Lisbon.
Evaluation scope. The empirical evaluation focuses on point accuracy throughout, including the precision ablation of
Section 5.4. Prediction intervals are shown illustratively rather than assessed through a calibration analysis using log scores, coverage, or Aitchison-based scoring rules. The benchmark set covers strong operational baselines but does not exhaust the space of multivariate or compositional alternatives.
6.5. Future Research
We group the extensions by what they require, since several could be undertaken immediately with the existing data and infrastructure while others need new linkage or substantial methodological development.
6.5.1. Near-Term Extensions
Probabilistic validation and calibration. The evaluation reports point accuracy throughout. A full assessment using log predictive density, interval coverage, and Aitchison-based scoring rules, applied to both the benchmark comparison and the precision ablation, would establish whether the point-accuracy gains carry over to the predictive distribution. This is the most immediate priority and requires no new data.
Exogenous predictors. Incorporating exchange rates, airline capacity, and search trends through the covariate matrix would improve accuracy and enable scenario analysis. The framework already accommodates this.
Broader comparator set. Extending the comparison to logistic-normal state-space models, multivariate log-ratio dynamics, and Markov-switching specifications would isolate which elements of the BDARMA specification drive its performance. The Markov-switching case is the most directly motivated: Gunter et al. [
35] show that regime-switching structure carries much of the forecasting burden for pandemic-era Airbnb occupancy, and a Markov-switching Dirichlet ARMA would test whether the same holds when the target is a composition. Machine learning comparators including recurrent architectures belong here as well, though their evaluation would be more informative at the finer granularity discussed below, where the panel of series is large enough to support the parameter counts involved.
6.5.2. Longer-Term Programs
Booking-to-stay linkage. Linking booking-month compositions to check-in-month outcomes or official arrivals data would clarify when and how booking composition serves as a reliable leading indicator of realized visitor mix. This requires data integration beyond the present scope and is the most substantively important of the open questions, since it determines how the forecasts should be interpreted.
Finer geographic granularity. Country-level or city-level analysis would support targeted marketing decisions, but data sparsity at that resolution threatens the strict positivity that both the Dirichlet likelihood and the ILR transformation require. Hierarchical modeling that borrows strength across destinations would be necessary.
Combined volume and composition forecasts. Integrating compositional forecasts with aggregate volume forecasts would yield complete predictions of arrivals by origin market. Doing this coherently requires attention to how uncertainty in the two components combines.
Real-time implementation. Deploying BDARMA models in operational forecasting systems with automated updating would maximize practical value, and would require addressing the estimation cost of refitting at each origin.
7. Conclusions
This paper developed and applied Bayesian Dirichlet autoregressive moving average models to forecast the evolving composition of guest origin markets in Airbnb bookings. The target throughout is booking-month composition, so the results should be read as forecasts of platform booking demand and of the booking pipeline, not as direct forecasts of stay-month arrivals. Our analysis of 108 months of reservations across four major destination regions demonstrates that BDARMA achieves 27% lower forecast error than naïve methods for EMEA, with strong evidence under both the pooled and origin-level comparisons ().
BDARMA also has lower error than seasonal naïve in EMEA, NAMER, and APAC. The pooled comparisons provide Holm-adjusted evidence in NAMER and APAC, while the more conservative origin-level sensitivity analysis retains clear statistical evidence only for APAC. Averaged across destinations, ETS attains the lowest error, with BDARMA close behind; BDARMA’s advantage is concentrated in compositionally diverse markets where multiple origin markets compete for shares in the 5–25% range.
The COVID-19 pandemic caused dramatic compositional shifts, most notably the collapse of the Chinese-origin booking share into APAC and the surge in within-region bookings across all destinations, with recovery trajectories that varied markedly across regions. Our models capture these dynamics through autoregressive and moving average terms while the Dirichlet likelihood ensures forecasts remain valid probability distributions.
Allowing seasonal variation in the precision parameter improves point accuracy in every destination region, by between 13% and 35% relative to an otherwise identical constant-precision specification, conditional on each region’s selected autoregressive and moving-average structure. The comparison concerns point accuracy; the corresponding effect on interval calibration is not established here.
The methodological contribution lies in bringing recent advances in Bayesian compositional time series to tourism forecasting with an explicit focus on booking-date origin shares. The empirical contribution documents substantial heterogeneity in how origin market compositions evolved during and after the pandemic, with direct implications for destination marketing strategy.
We opened this paper with three decisions that require knowing not just how many visitors may eventually arrive, but what the booking pipeline suggests about where they will come from: marketing budget allocation across source markets, concentration risk monitoring, and operational planning at the property level. These decisions share a common structure. They must be made months in advance, they are difficult to reverse once committed, and they depend critically on the composition of demand rather than its aggregate volume. A destination that watched its Chinese-origin booking share grow from 15% to 25% over three years before the pandemic was accumulating fragility that no aggregate volume forecast would have revealed. A DMO setting its Q3 campaign budget in January needs a distributional view of where its booking demand is coming from, not just how much there is. The BDARMA framework, with its coherent simplex-valued forecasts and uncertainty quantification, is designed precisely for this class of decision. As tourism markets continue to evolve in response to changing travel preferences, economic conditions, and potential future disruptions, the ability to forecast demand composition in forward-looking booking data will remain a practical imperative for destination stakeholders.