1. Introduction
The film industry has long occupied a distinct position in cultural economics research as an investment environment in which standard competitive market conditions—predictable demand, stable cost–return relationships, and symmetric information—break down in systematic ways.
De Vany and Walls (
1999) demonstrated that box office revenues exhibit effectively infinite variance, suggesting that the notion of a “typical” film carries little financial meaning.
Rosen (
1981) showed that in certain economic activities, output tends to concentrate among a small number of producers, income distributions exhibit pronounced right-skewness, and rewards at the top grow disproportionately large. This superstar economics framework finds its counterpart in the motion picture industry, where a handful of exceptional productions appear to capture a disproportionate share of both audience attention and financial returns.
Caves (
2003) argued that cultural industries are shaped by the contractual arrangements between creative talent and ordinary inputs, and that demand uncertainty is a structural rather than incidental feature of these markets.
Hadida (
2009) similarly found that success in the motion picture industry operates simultaneously across commercial, artistic, and sociological dimensions, complicating financial prediction. When these distributional properties are combined with large-scale sunk costs and irreversible production expenditures, film investment emerges as one of the more analytically demanding domains within cultural economics (
Eliashberg et al., 2006;
Filson, 2026).
One question that remains incompletely resolved in film economics research concerns the relationship between financial scale and investment efficiency. The existing literature has leaned heavily on box office revenue as the primary measure of commercial success, yet revenue taken in isolation does not account for the production expenditure that generated it. A film with substantial gross receipts may still deliver poor returns to its investors if the underlying budget was proportionally large.
Ravid (
1999), in one of the few studies to examine ROI directly, found that the signals producers and investors typically rely on, budget and star power, failed to reliably predict investment returns.
Yoong Hon and Yen (
2023) extended this line of inquiry to ancillary markets, with comparable findings on budget and investment efficiency.
Patrocínio et al. (
2024) raised a related methodological concern, showing that conclusions about financial performance can shift depending on how ROI is defined.
A second unresolved question concerns the relative weight of audience engagement metrics compared to content characteristics in shaping investment outcomes. The spread of digital platforms has made audience interaction data (ratings, vote counts, engagement scores) observable at a scale that was previously unavailable, yet the relationship between these metrics and financial returns in cultural markets remains incompletely understood.
Liu (
2006) identified word-of-mouth volume, independent of tone, as a key audience-driven signal for box office performance.
Duan et al. (
2008) identified a similar self-reinforcing dynamic between engagement volume and sales performance.
Castillo et al. (
2021) confirmed similar engagement effects across different national market contexts.
A review of the existing literature points to several gaps that motivate the present study. First, the majority of studies evaluate commercial success through box office revenue, with limited attention paid to return on investment as a performance measure that is arguably more meaningful from an investor’s perspective. Second, while the associations between film performance and variables such as production budget, critical reviews, online engagement, and social media indicators have been examined in various combinations, no clear consensus has emerged on which factors most reliably account for investment success. Third, existing work has concentrated largely on linear relationships and predictive accuracy, leaving non-linear effects and potential threshold values in investment outcomes relatively underexplored. Finally, despite the growing use of machine learning methods in film performance research, the economic mechanisms underlying model predictions and their theoretical implications for cultural markets have received limited attention. The question of whether investment success in cultural product markets is more closely tied to production scale or to audience attention and engagement therefore remains open.
This study addresses these questions by examining the determinants of film-level ROI using a sample of 3153 films spanning the period 1916–2017. Rather than treating financial prediction as an end in itself, the analysis is framed as a means of identifying which observable characteristics are associated with above- or below-median investment returns, and how these associations vary across different value ranges. Because several of the variables examined, most notably audience engagement metrics, become observable only after theatrical release, the study is explicitly explanatory rather than pre-release predictive in orientation: the objective is to identify the characteristics that distinguish films with high and low realised investment efficiency, not to forecast returns at the greenlighting stage. Ensemble learning algorithms and SHAP-based variable decomposition are employed as analytical instruments for uncovering the economic mechanisms at work; the contribution of this study lies not in the methodological tools themselves but in what they reveal about the economics of film investment.
The study makes four contributions to the cultural economics literature. First, it shifts the outcome measure from box office revenue, which has dominated film success research, to return on investment, and adopts a median-based binary classification framework that directly addresses the extreme distributional skewness documented by
De Vany and Walls (
1999), thereby extending
Ravid’s (
1999) foundational ROI analysis with substantially larger data and more recent analytical tools. Second, it establishes that audience engagement volume stands alongside production budget as a leading signal of investment efficiency on a cost-normalised outcome measure, ranking first in the classification framework and second in the continuous specification. This is a stronger test than the revenue-based associations reported in the word-of-mouth literature, since attention is itself partly financed by expenditure and its explanatory standing would be expected to attenuate once returns are normalised by cost. Third, it identifies non-linear threshold effects that linear specifications cannot recover: in the fitted model, the estimated contribution of production budget declines across the upper part of the budget range and becomes negative in the region corresponding to roughly 25 million USD, while audience engagement contributes sharply above a comparatively modest volume of accumulated attention, so that scale and attention operate in opposite directions across their respective ranges.
2. Literature Review
The financial structure of the film industry exhibits characteristics that sit uneasily with standard investment theory.
De Vany and Walls (
1999), drawing on a sample of over 2000 motion pictures, found that box office revenues follow an asymptotic Pareto distribution with effectively infinite variance. This distributional form, in which the mean is dominated by a small number of exceptional productions, suggests that constructing a meaningful expectation of typical performance in film investment is statistically untenable.
Eliashberg et al. (
2006), in a broad review of research in the field, noted that the motion picture industry encompasses a wide range of unresolved analytical problems spanning demand forecasting, distribution strategy, and beyond, while also observing that the richness of available data makes this industry an unusually fertile ground for financial analysis at the individual project level.
Hadida (
2009), synthesising film performance research published between 1977 and 2006, found that success in the motion picture industry operates simultaneously across commercial, artistic, and sociological dimensions, and that the production structure theorised by
Caves (
2003) through the concept of “humdrum inputs” appears to deepen rather than resolve this underlying uncertainty.
The limitations of revenue-based analysis in this uncertain environment have received growing recognition.
Basuroy et al. (
2003), analysing eight weeks of box office data, found that critics played both an influencing and a predictive role, that negative reviews carried stronger effects than positive ones, and that high production budgets appeared to serve a protective function only for films receiving unfavourable critical responses. This suggests that the signalling argument, whereby budget size functions as a quality guarantee, may hold only under specific conditions.
Ravid (
1999), examining ROI directly using a sample of films produced in the 1990s, found that neither production budget nor star power reliably predicted investment returns, and that the informational content of the signals on which producers and investors typically rely appeared to be considerably more variable than commonly assumed.
Lauria and Phillips (
2021) argued that a return-oriented rather than revenue-oriented framework is necessary for film investment to be evaluated in terms comparable to other asset classes. Along similar lines,
Kim et al. (
2020) positioned ROI as a more appropriate performance measure for film investment decisions, while
Patrocínio et al. (
2024) showed that conclusions about financial performance can shift substantially depending on how ROI is defined and computed, underscoring the importance of definitional consistency in comparative analyses.
Yoong Hon and Yen (
2023), focusing on ancillary markets, found that higher-budget productions tended to underperform lower-budget counterparts on investment efficiency metrics despite their advantage in absolute sales volumes. Given the pronounced right-skewness of the ROI distribution and the asymmetric influence of a small number of exceptional performers, a binary classification framework may offer a more statistically tractable approach than continuous dependent variable modelling (
Filson, 2026;
e Souza et al., 2023).
Early systematic attempts to forecast film financial performance relied primarily on regression-based statistical methods, though the identification of areas where these approaches fell short contributed to a gradual shift towards machine learning classification frameworks. A significant turning point in this transition was the work of
Sharda and Delen (
2006), who moved away from the point estimation approach that had dominated film research and instead classified films into nine categories ranging from flop to blockbuster based on box office revenue ranges. Testing their neural network model using 10-fold cross-validation, they both established a methodological benchmark and demonstrated the applicability of classification approaches to film performance prediction.
Ghiassi et al. (
2015), employing a dynamic artificial neural network in the pre-release period, incorporated variables including budget, advertising expenditure, runtime, and seasonality, and achieved meaningful improvements in predictive accuracy relative to existing methods.
Mahmud et al. (
2020) noted that only 36% of films released in the United States between 2000 and 2010 recovered their production costs through box office revenue, and compared logistic regression, support vector machines, and multilayer perceptron models across success categories.
Lash and Zhao (
2016) structured profitability prediction around the questions of who, what, and when, examining how different types of pre-release information, including cast, subject matter, and release timing, contribute to early-stage investment decisions.
The advantages of ensemble learning methods in this context have received growing empirical support.
Ahmad et al. (
2020b), in a systematic review covering 36 relevant studies, found that regression and classification approaches dominated the literature, with multiple linear regression and support vector machines among the most frequently used techniques.
Lee et al. (
2020), using a sample of 1439 films, found that decision tree ensembles—including random forests, bagging, and boosting variants—outperformed both k-nearest neighbour and linear regression-based ensemble methods in predicting box office revenue across multiple post-release windows.
e Souza et al. (
2023), using inflation-adjusted profit as the dependent variable across a sample of 3167 films, compared Random Forest, support vector machines, and neural networks, and found that Random Forest consistently produced the strongest results across different sampling strategies.
Ni et al. (
2022) evaluated multi-model ensemble approaches, while
Tang (
2024) assessed an optimised XGBoost model in the context of box office forecasting, with both studies suggesting that these algorithms can achieve competitive performance on large and complex datasets.
Leukel et al. (
2026) argued that the primary challenge in this domain lies not in data availability but in constructing variables with sufficient predictive content to drive model performance.
The role of social media engagement and online discourse in shaping film financial performance has attracted considerable research attention.
Liu (
2006), analysing more than 12,000 messages from Yahoo Movies discussion boards, found that word-of-mouth volume carried meaningful explanatory power for both weekly and cumulative box office revenue, while the valence of that communication, whether positive or negative, contributed comparatively little additional explanatory value.
Moul (
2007) estimated the effect of word-of-mouth on theatrical admissions econometrically, finding that approximately 10% of the variation in consumer expectations could be attributed, directly or indirectly, to the spread of information, and that this information reached consumers relatively quickly.
Dellarocas et al. (
2007) developed diffusion models integrating online review metrics into film revenue forecasting and found that these metrics improved predictive accuracy when used alongside pre-release marketing variables and professional critical assessments, suggesting that online audience engagement may carry informational content beyond what traditional variables capture.
Duan et al. (
2008) modelled the relationship between word-of-mouth volume and box office revenues using a dynamic simultaneous equations system, finding evidence of a self-reinforcing dynamic whereby engagement volume and sales performance appeared to mutually reinforce one another, independently of content quality assessments.
Castillo et al. (
2021), using a sample of 966 films released in the United Kingdom and Spain, found that both personal and interactive social media engagement were positively associated with opening-week revenues, and that these effects appeared to amplify one another.
Elberse and Eliashberg (
2003) showed that film attributes and advertising expenditure tended to influence revenues indirectly through exhibitors’ screen allocation decisions rather than directly, pointing to the potential endogeneity of audience engagement variables in this setting.
Fan et al. (
2021), drawing on signal theory and a sample of 80 Chinese films, found that internal signals, star and director power, appeared more influential during the opening week, while eWOM volume tended to be more determinative in subsequent weeks.
Eliashberg and Shugan (
1997) found that critical reviews were not meaningfully associated with early box office performance but were strongly related to later and cumulative revenues, suggesting that critics may function more as leading indicators of underlying quality than as active shapers of audience behaviour, and that quality signals may offer more limited immediate predictive relevance compared to engagement volume. Uncertainty regarding the non-linear and time-dependent patterns through which these variables relate to film revenue nonetheless persists in the literature (
McKenzie, 2023;
Ahmad et al., 2020a).
As the predictive accuracy of machine learning models has improved, the interpretability of their decision-making mechanisms has become an increasingly prominent concern in the research agenda.
Arrieta et al. (
2020), in a comprehensive examination of the explainable artificial intelligence field spanning concepts, taxonomies, and applications, systematically categorised model-agnostic and post hoc interpretability techniques and argued that these techniques constitute a prerequisite for responsible and trustworthy AI deployment. The SHAP framework (
Lundberg & Lee, 2017), grounded in Shapley values from cooperative game theory, decomposes each variable’s contribution to an individual prediction in terms of both direction and magnitude, thereby supporting both global and local interpretability.
Mammadov et al. (
2025) noted that while machine learning models have improved predictive accuracy in film profitability research, the interpretability gap remains a substantial unresolved limitation in this area, and proposed a multi-layered explainability framework combining XGBoost, CatBoost, and artificial neural network models with SHAP, permutation feature importance, and feature importance ranking measure techniques.
Behrens et al. (
2021), drawing on conceptual frameworks spanning script analytics, talent analytics, and audience analytics, observed that the integration of data-driven approaches into film production decisions remains an emerging rather than mature area of practice. In the present study, SHAP analysis is applied to interpret the predictive mechanism of the Random Forest model at the variable level, decompose the direction and magnitude of individual effects, and identify potential threshold values in the relationships between predictors and investment outcomes.
4. Results
This section presents the findings on the ROI classification performance of the Random Forest, XGBoost, and CatBoost algorithms. Descriptive statistics for the dataset are presented first, followed by the model performance comparison, variable importance analysis, and SHAP-based interpretation findings.
Table 2 reports descriptive statistics for the variables used in the study. Despite a median ROI of 1.30, the maximum value reaches 8,499,999, and the standard deviation is exceptionally large (SD = 152,535.30), consistent with the extreme right-skewness (skewness coefficient: 54.998) that motivated the median-based binary classification approach described in
Section 3.2. Descriptive statistics for all variables used in the study are presented in
Table 2.
4.1. Model Performance Comparison
The performance of Random Forest, XGBoost, and CatBoost on the ROI classification task was compared using AUC, accuracy, recall (sensitivity), specificity, precision, F1 score, and balanced accuracy. Results are presented in
Table 3.
Random Forest demonstrated superior overall classification performance compared with the other algorithms. Balanced results across different performance metrics indicate that the model does not favour a particular class but instead distinguishes between high- and low-ROI films with comparable success. This finding suggests that Random Forest provides a more balanced and reliable prediction framework for ROI classification.
While XGBoost exhibited strong discriminative performance, it demonstrated a relatively balanced ability to classify both high- and low-ROI films. Although its overall performance remained below that of Random Forest, it achieved a stable trade-off between sensitivity and specificity.
CatBoost produced a relatively balanced performance profile but remained consistently below Random Forest across all evaluation metrics.
The ROC curves for all three models are presented comparatively in
Figure 1. All models lie above the diagonal reference line (represented by the grey dashed line) representing random classification, suggesting meaningful discriminative capacity with respect to the ROI classes. Random Forest’s ROC curve maintains a more consistent distance from the diagonal across different threshold values, pointing to a more stable sensitivity–specificity balance across classification scenarios.
Following the comparative evaluation of the three algorithms, Random Forest demonstrated the strongest overall classification performance and was selected for further stability assessment. Repeated 10-fold cross-validation with five repetitions was subsequently conducted on the training set to assess the robustness of the selected model. AUC served as the primary optimisation metric, with sensitivity and specificity reported alongside it.
Table 4 presents the average performance metrics obtained across all folds and repetitions.
As shown in
Table 4, the Random Forest model yielded a mean AUC of 0.8247 across the repeated 10-fold cross-validation procedure, with relatively low variation across resamples (SD = 0.0226). The sensitivity and specificity estimates were also broadly comparable to those obtained from the hold-out test-set evaluation. Overall, the consistency of the performance estimates across the cross-validation and hold-out test-set evaluations provides additional evidence of the stability of the Random Forest model.
4.2. Random Forest Variable Importance Analysis
The contribution of each variable to ROI classification in the Random Forest model was assessed using the MeanDecreaseAccuracy and MeanDecreaseGini criteria. Results are presented in
Table 5.
Table 5 shows that the vote_count_log variable is the most important predictor according to both variable importance criteria. This finding indicates that the total number of votes, representing audience engagement, plays a more important predictive role in ROI classification compared to production budget, popularity, and content attributes.
The budget_log variable is among the most important predictors according to both variable importance measures. However, the difference in its relative ranking across the two variable importance measures suggests that the relationship between production budget and ROI may be more complex than that of the other variables. This nonlinear relationship is examined in more detail through SHAP analysis in
Section 4.3.
popularity_log ranks third in the variable importance analysis. This finding suggests that a film’s visibility and level of audience interest on online platforms are closely associated with its ROI classification. While vote_average also contributes to ROI classification, its predictive relevance is more limited than that of vote_count_log and popularity_log.
runtime_log exhibited moderate variable importance, whereas the genre variables generally showed the lowest importance values. Although the relative ranking of the genre variables varied slightly between the two importance measures, they consistently ranked below the numerical variables. This finding suggests that audience engagement and visibility indicators carry greater predictive relevance for ROI classification than content-related characteristics.
4.3. SHAP Analysis and Examination of Variable Effects
SHAP analysis was applied to examine the predictive mechanism of the Random Forest model at the variable level. The SHAP summary plot is presented in
Figure 2.
Figure 2 indicates that vote_count_log, budget_log, and popularity_log exhibit the widest SHAP value distributions in relation to model predictions, in that order. This ranking is consistent with the variable importance analysis results presented in
Section 4.2. SHAP values for the genre variables are largely concentrated around zero, while runtime_log shows low variation. This pattern suggests that numerical audience engagement metrics may carry substantially more predictive information for ROI classification than content-based variables.
The marginal effects of variables across different value ranges and potential threshold values were examined through SHAP dependence plots, the results of which are presented in
Figure 3. In these plots, the red curves represent LOESS (Locally Estimated Scatterplot Smoothing) non-parametric regression trend lines, illustrating the nonlinear relationship between feature values and their corresponding SHAP values.
budget_log: SHAP values remain limited at low budget levels and move into positive territory across mid-range budget values. Around a log value of 17.03 (corresponding to roughly 25 million USD), SHAP values begin to decline and eventually become negative. This pattern indicates that the fitted model assigns a less positive contribution to ROI predictions at higher budget levels. When the interaction with vote_count_log is examined, higher audience engagement appears to be associated with more positive SHAP values for budget.
vote_count_log: The SHAP dependence plot indicates a marked change in the contribution of vote count to the model predictions. Beyond an approximate threshold of 6.2 on the log scale- corresponding to roughly 493 votes- SHAP values rise sharply into positive territory. This positive effect is more pronounced in films with high popularity scores, and this interaction is effectively captured by the fitted model.
popularity_log: This variable exhibits a positive and non-linear association with ROI. Around a log value of 3.09 (corresponding to approximately 22 popularity points), SHAP values begin to shift into positive territory, indicating that the fitted model assigns increasingly positive contributions to ROI predictions as popularity increases.
vote_average: SHAP values tend to be negative at lower rating levels. Above an approximate threshold of 6.4 out of 10, SHAP values become predominantly positive, though the rate of increase appears to slow at higher rating values. This pattern indicates that higher audience ratings receive increasingly positive SHAP contributions in the fitted model.
runtime_log: The concentration of SHAP values largely around zero suggests that film runtime contributes relatively little to the fitted model’s predictions. Mild negative effects were observed in the mid-length range, though the overall SHAP variation for this variable remains low.
4.4. Robustness Analysis Using Continuous ROI
To assess the robustness of the findings, an additional Random Forest regression model was constructed using continuous ROI as the dependent variable. Since the continuous ROI variable exhibited substantial right skewness, a natural logarithmic transformation, ln(1 + ROI), was applied prior to the analysis to reduce the influence of extreme values and improve the distributional properties of the outcome variable. The results regarding model performance are presented in
Table 6.
As shown in
Table 6, the performance of the Random Forest regression model was evaluated using the coefficient of determination (R
2), Root Mean Squared Error (RMSE), and Mean Absolute Error (MAE), with RMSE and MAE providing complementary measures of prediction error. The R
2 value of 0.403 indicates that the model explains a moderate proportion of the variation in the log-transformed ROI. To evaluate whether using ROI as a binary or continuous outcome variable affects the variable importance ranking, the variable importance rankings obtained from the classification and regression models were compared. Since the classification and regression models use different permutation-based variable importance measures (Mean Decrease Accuracy and %IncMSE, respectively), the comparison focuses on the relative ranking of the predictors rather than numerical importance values. The rankings obtained for each method for the top five variables are given in
Table 7.
Table 7 shows that the most important predictors in both models consist of the same variables. Only the positions of the budget_log and vote_count_log variables in the top two positions have changed, indicating that the relative contribution of these two variables to the model can show small differences depending on the modelling approach. However, the fact that the most important predictors generally remain consistent supports the robustness of the study’s findings against treating ROI as a binary or continuous outcome variable.
5. Discussion
The model performance findings of this study broadly support the effectiveness of ensemble learning methods in film investment classification and are consistent with patterns documented in the existing literature. Random Forest outperformed both XGBoost and CatBoost on the classification task, a result in line with prior findings on the relative advantages of decision tree ensembles on structured data problems of this kind (
Lee et al., 2020;
e Souza et al., 2023). The economic interpretation of this performance matters more than its magnitude. In an industry characterised by structurally limited financial predictability (
De Vany & Walls, 1999), the fact that a small set of observable variables separates above- and below-median ROI films with meaningful accuracy indicates that uncertainty in cultural markets is not entirely random and that investment outcomes retain at least some systematic structure. Equally informative is where the models fail: XGBoost identified high-ROI films almost without exception but classified low-ROI films poorly at the default threshold, meaning that the observable characteristics examined here distinguish commercial success more sharply than they distinguish commercial failure. This asymmetry is itself substantively meaningful for investors, since it is exposure to loss rather than participation in success that governs portfolio construction in hit-driven industries, and it suggests that threshold optimisation or class weighting would be a necessary calibration step before any applied use of such a model.
The most striking finding of the variable importance and SHAP analyses is that audience engagement volume represented by vote_count_log ranks first among all variables in the classification framework, with SHAP values rising sharply into positive territory beyond an approximate threshold of 493 votes. From a cultural markets perspective, this suggests that audience participation may function not only as an indicator of demand but also as a leading signal of investment success.
Liu (
2006) found that word-of-mouth volume, independent of whether its content was positive or negative in tone, served as the primary audience-driven signal for box office performance, while
Duan et al. (
2008) identified a self-reinforcing dynamic between engagement volume and sales performance.
Moul’s (
2007) econometric estimate that roughly 10% of consumer expectation variation traces back to information diffusion lends further support to this pattern. The present findings suggest that this dynamic may extend beyond the revenue level to the level of investment returns. Viewed through the lens of information asymmetry and signalling theory in cultural economics, this pattern suggests that traditional quality signals circulating among producers and distributors- including genre, cast, and critical ratings- may carry comparatively weaker predictive relevance than audience visibility.
Fan et al. (
2021) similarly found that external signals tended to outweigh internal ones in post-opening periods, while
Castillo et al. (
2021) confirmed the positive association between audience engagement and box office revenues across different market contexts.
Dellarocas et al. (
2007) reached a similar conclusion using diffusion models that incorporated online review metrics into revenue forecasting. Viewed through
Rosen’s (
1981) superstar economics framework, the top ranking of vote_count_log is perhaps unsurprising: the concentration of audience attention on a small number of productions, and the disproportionate shaping of financial returns by that concentration, is a defining feature of cultural markets.
Elberse and Eliashberg’s (
2003) finding on exhibitors’ screen allocation decisions points to a similar two-way relationship between performance and audience accessibility.
An important qualification concerns the temporal availability of the variables that emerge as the strongest signals in this analysis. Both vote count and popularity are recorded after theatrical release, and ROI is likewise realised ex post. The framework developed here should therefore be understood as an explanatory decomposition of realised investment outcomes rather than as a pre-release forecasting instrument. Of the variables examined, only production budget, runtime, and genre are known at the greenlighting stage, whereas audience engagement metrics become observable only as demand realises itself in the market. This distinction carries direct implications for how the findings should be used. The practical relevance of the engagement finding lies less in ex ante project selection than in decisions taken during and after the theatrical window, including screen retention and release-window extension, the reallocation of marketing expenditure, the valuation of ancillary and streaming rights, and sequel or franchise commitments. Read in this way, the result that engagement volume dominates production budget does not imply that investors are able to anticipate returns before production begins; it indicates that once a film has entered the market, the volume of audience attention it accumulates carries more information about its eventual investment efficiency than the scale of expenditure committed to it. This does not imply that pre-release prediction is infeasible. Studies restricting themselves to information available before a film reaches audiences have in fact achieved strong results, but they do so by constructing features considerably richer than standard production metadata: cast collaboration networks and plot topic distributions (
Lash & Zhao, 2016), screen counts at release alongside budget and seasonality (
Ghiassi et al., 2015), or sentiment and purchase intention extracted from pre-release trailer reviews (
Ahmad et al., 2020a).
Lash and Zhao (
2016) report that removing such constructed features reduces the AUC of their pre-production model from 0.921 to 0.707, which indicates that the predictive content of ex ante designs rests precisely on the kind of variables the present dataset lacks. The explanatory orientation adopted here is therefore a property of the available data rather than a claim about the limits of pre-release prediction, and it reinforces the case for treating the construction of pre-release engagement proxies as the natural extension of this work.
The findings on the effect of production budget on investment returns may represent the most distinctive contribution of this study from a cultural markets economics perspective. In the fitted model, the SHAP contribution of budget declines across the upper part of the budget range and becomes negative in the region corresponding to roughly 25 million USD. This describes how the model distributes contributions across the observed range rather than identifying a point at which the underlying relationship changes sign, and the particular value should be read as approximate and sample-specific. The cultural economics literature has proposed that budget functions as a signalling mechanism, with high production expenditure interpreted as a quality assurance (
Ravid, 1999). The findings of the present study indicate, however, that within this sample the association between expenditure and investment efficiency weakens at higher budget levels. This is consistent with the possibility that budget growth translates into cost pressure rather than proportionate return growth, although the design does not permit that mechanism to be identified directly. This pattern is consistent with
Yoong Hon and Yen’s (
2023) finding of a high-budget, low-ROI relationship in ancillary markets. Given the extreme right-skewed income distribution documented by
De Vany and Walls (
1999), the interpretation that high budgets may sometimes amplify rather than reduce uncertainty in this environment gains some support;
Filson (
2026) similarly found that the majority of project-level ROI variation in film stems from structural uncertainties tied to the development process rather than post-release demand discovery.
Einav (
2007) showed that the largest productions tend to concentrate strategically around periods of peak demand, yet found that this strategic timing does not automatically translate into improved investment efficiency. From a cultural markets’ perspective, this threshold finding provides some empirical basis for questioning whether the budget-centric investment decision-making prevalent in film financing may carry systematic inefficiencies. The assumption that increasing production scale secures investment success appears, at least within the context of this dataset, to lack consistent empirical support.
The third-place ranking of popularity_log and the shift in SHAP values into positive territory beyond an approximate threshold of 22 popularity points suggest that platform-specific visibility may be meaningfully associated with investment success in cultural markets. This is broadly consistent with the literature on the growing role of digital platforms in shaping the financial performance of cultural products (
Castillo et al., 2021;
Ahmad et al., 2020a). From a cultural economics perspective, the emergence of popularity metrics as investment signals may point to a dynamic that extends beyond conventional signalling theory: a film’s visibility in the cultural marketplace could function both as a reflection of audience demand and as an independent factor that reinforces that demand.
The limited effect of vote_average- which shifts into positive territory only above an approximate threshold of 6.4 out of 10 and remains weaker than vote_count_log and popularity_log- suggests that the translation of perceived content quality into financial success may follow an indirect rather than direct path in cultural markets.
Liu’s (
2006) finding that engagement volume carries stronger signal value than valence finds broad support in the present results: how widely a film is discussed appears to be more closely tied to investment returns than how favourably it is assessed.
Eliashberg and Shugan’s (
1997) view of critics as signals that trail quality rather than drive audience behaviour fits this pattern: the failure of vote_average to match vote_count_log in predictive relevance- the quality signal yielding to the volume signal- is consistent with both
Liu (
2006) and
Eliashberg and Shugan (
1997).
Basuroy et al.’s (
2003) budget-as-shield finding limited to films facing unfavourable reviews points in the same direction. This suggests that investment in content quality may translate into financial returns only under specific conditions. From a cultural economics standpoint, the limited effectiveness of content quality investment across cast, script, and production value in securing financial returns may serve as an empirical indication that value creation mechanisms in cultural markets operate differently from standard production quality models.
The emergence of genre variables as the lowest-ranked group raises questions about the role commonly attributed to content characteristics in the cultural economics literature.
Ahmad et al. (
2020b) noted that while genre is widely used in film performance research, its explanatory power has varied considerably across studies. The present findings suggest that genre carries relatively limited predictive relevance for investment returns once audience visibility and financial scale are accounted for, pointing to a more circumscribed role for film category in shaping financial performance than is typically assumed among producers and investors.
This study contributes to the cultural economics literature in several respects. The first concerns the shift from box office revenue as the dominant outcome measure in film success research towards ROI as the dependent variable. It has been argued that an investment efficiency framework is necessary for understanding producer and investor behaviour in cultural markets, given the limited informational content of absolute revenue measures (
Lauria & Phillips, 2021;
Kim et al., 2020). The median-based binary classification approach adopted to address the extreme distributional skewness documented by
De Vany and Walls (
1999) directly engages with this methodological gap and extends
Ravid’s (
1999) foundational ROI analysis using more recent analytical tools.
Caves’s (
2003) characterisation of demand uncertainty as a structural feature of cultural industries, and
Hadida’s (
2009) comprehensive synthesis of the multi-layered nature of film success, provide theoretical grounding for this study and support the case for ROI-oriented approaches in film research.
The second contribution concerns the non-linear form of the association between production scale and investment efficiency. The existence and limits of scale economies in cultural markets remain a contested issue. The present findings describe a pattern in which the estimated contribution of budget weakens at higher expenditure levels rather than increasing monotonically, which is difficult to reconcile with the assumption that budget growth translates automatically into investment success. We would not characterise this as evidence that scale advantages give way to scale disadvantages at an identifiable point, since the analysis describes the behaviour of a model fitted to a single observational sample and cannot isolate the mechanism responsible. Read at this level of generality, the pattern indicates that the budget range over which additional expenditure improves investment efficiency may be narrower than budget-centric decision heuristics assume. Any policy reading should be correspondingly cautious: the result raises a question about incentive structures calibrated to budget size rather than providing a basis for redesigning them.
The third contribution concerns the standing of audience engagement volume among the determinants of investment efficiency. In the classification specification, engagement ranks first on both importance criteria, ahead of production scale and content characteristics; in the continuous specification, engagement and budget exchange the top two positions. Their relative order is therefore specification-dependent, while engagement remains among the leading predictors in both. The theoretical significance of this result lies in the outcome variable against which it is established rather than in the ordering itself. The existing word-of-mouth literature has demonstrated that engagement volume is associated with box office revenue (
Liu, 2006;
Duan et al., 2008;
Dellarocas et al., 2007;
Castillo et al., 2021), and this relationship is by now well documented. Revenue and investment efficiency, however, are not interchangeable outcomes, and a finding established on one does not transfer automatically to the other.
Lash and Zhao (
2016) report a correlation of only 0.077 between box office revenue and ROI in their sample, indicating that the two measures capture substantially different dimensions of financial performance. Audience attention is itself partly a product of expenditure: larger budgets purchase wider releases, heavier advertising, and more visible casts, all of which generate attention. On this reasoning, engagement and budget should co-move, and once returns are normalised by production cost the apparent advantage of engagement over budget should attenuate or disappear. The present findings indicate that it does not. Engagement retains its explanatory standing alongside budget on a cost-normalised outcome measure, which suggests that audience attention carries informational content that is not fully embodied in production expenditure.
Interpreted through
Rosen’s (
1981) superstar framework, this points to a mechanism that the revenue-based literature is not positioned to identify. Superstar economics explains why attention, and therefore revenue, concentrates among a small number of productions. It does not by itself establish that returns concentrate in the same way: if attention were anticipated at the financing stage, it would be capitalised into input prices (talent fees, acquisition costs, marketing commitments) and the resulting rents would accrue to the factors of production rather than to the investor. That engagement retains its explanatory strength after costs are netted out suggests that this capitalisation is incomplete, and that the attention a film ultimately attracts is imperfectly anticipated by the expenditure committed to it. This offers one empirically grounded reading of why demand uncertainty in cultural markets, as characterised by
Caves (
2003), persists as a structural rather than incidental feature: the uncertainty resides not only in whether audiences will attend, but in the weak correspondence between what producers can commit ex ante and the attention that materialises ex post.
Read together with the budget threshold, the two findings describe a single pattern rather than two separate ones. Expenditure exhibits diminishing and eventually negative marginal contributions to investment efficiency beyond an approximate threshold, while audience engagement contributes sharply above a comparatively modest level of accumulated attention. The two mechanisms therefore point in opposite directions across their respective ranges, which is consistent with
Ravid’s (
1999) finding that budget and star power failed to predict returns and with
Yoong Hon and Yen’s (
2023) evidence of high-budget underperformance in ancillary markets, while extending both by identifying where the reversal occurs and what displaces expenditure as the dominant signal. The practical implication for cultural economics is that the value creation mechanism in film appears to run through attention rather than through scale, and that these are separable: attention is not simply purchased. This provides empirical purchase on the questions of demand uncertainty and investor behaviour that
Eliashberg et al. (
2006) identified as unresolved, and indicates that investment decision frameworks in cultural markets require a more systematic treatment of audience visibility than budget-centric heuristics currently allow.
The robustness check reported in
Section 4.4 indicates that the ranking of the leading predictors is largely invariant to whether ROI is treated as a binary or a continuous outcome, which suggests that the substantive conclusions drawn here are not artefacts of the median-based classification. The one difference worth noting is that budget and vote count exchange the first two positions between the specifications, so that engagement leads in the classification framework and budget in the continuous one. We therefore treat the two as joint leading predictors rather than claiming a stable ordering between them; what is invariant is their standing relative to popularity, audience rating, runtime, and genre, which indicates that the mechanisms identified operate across the ROI distribution rather than only at the median boundary. The binary specification is nonetheless retained as the primary framework, since the objective of the study is to distinguish films with high and low realised investment efficiency rather than to estimate ROI values precisely, and classification remains the appropriate design for that purpose.
Taken together, these findings carry practical implications that differ from prevailing industry heuristics in three respects. For producers and financiers, the weakening of the budget contribution at higher expenditure levels suggests that additional expenditure in this range is better understood as an increase in exposure than as an improvement in expected efficiency; where the objective is return rather than absolute revenue, a portfolio of moderately budgeted productions is indicated over a smaller number of large commitments. For distributors and rights holders, the strength of the engagement signal points to early accumulated audience attention as a more informative input for post-release decisions—screen retention, window extension, marketing reallocation, and ancillary and streaming valuation—than either budget or genre. For public agencies designing film support schemes, the weakening of the budget contribution at higher expenditure levels raises the question of whether incentive structures calibrated to budget size direct resources towards the range in which investment efficiency is weakest. We note this as a question the present design can raise but not settle.
Several limitations should be kept in mind when interpreting the findings of this study. The budget and revenue data used were drawn from TMDB’s publicly available dataset, which does not cover marketing expenditure and reflects only theatrical box office revenue. This means that ROI calculations may not fully capture actual investment costs or total revenues. Similar data constraints apply to the majority of existing studies in this area and are widely acknowledged in the literature (
Mammadov et al., 2025;
Leukel et al., 2026). The filtering criteria applied in this study, retaining only observations with budget and revenue values above zero and above 1000 USD, should be regarded as a methodological step taken to improve data quality. First, the dataset covers the period 1916–2017 and relies on a single platform-specific source. The exclusion of the post-2017 period, during which streaming platforms have become dominant, and revenue models have undergone substantial transformation, limits the generalisability of the findings to current cultural market conditions. As
McKenzie (
2023) has noted, the effects of digital transformation on the economic structure of the film industry remain incompletely documented, and examining how this transformation affects variable importance rankings represents a productive direction for future research. Second, variables that could have a strong bearing on financial performance in cultural markets including marketing expenditure, distribution width, and advertising investment are absent from the dataset. The absence of such variables is recognised in the cultural economics literature as one of the primary factors limiting model explanatory power (
Leukel et al., 2026;
Eliashberg et al., 2006). Their inclusion would likely enrich the interpretation of the relationship between audience engagement and budget. Third, the median-based binary classification approach partially obscures the nuances within the ROI distribution. A multi-class categorisation or a graduated approach could allow for a more detailed assessment of investment performance in cultural markets. This study is based on observational data and the associations identified are correlational in nature. The patterns among variables have been interpreted within a relational rather than causal framework, and causal claims have been systematically avoided. Testing the causal dimensions of these relationships through instrumental variable approaches, difference-in-differences designs, or natural experiment methods is suggested as a priority for future research. Finally, both the potential endogeneity and the ex post availability of the vote_count and popularity variables require that interpretations of these variables be treated with particular care. Films that perform strongly in the market attract more votes and greater platform visibility, so the association identified here is likely to operate in both directions; more importantly, because both variables are recorded only after release, they cannot serve as inputs to pre-release investment decisions. A genuinely ex ante framework would require pre-release proxies for anticipated audience engagement, such as trailer view counts, search volume trends, pre-sale ticket data, or social media activity in the weeks preceding release, none of which are available in the TMDB dataset. Constructing and testing such a design, alongside investigating the causal dimension of this relationship through instrumental variable approaches or natural experiment designs, would considerably strengthen the analytical basis for conclusions regarding the role of audience engagement in investment success in cultural markets.