1. Introduction
Hedonic models are employed to capture the heterogeneous effects of intrinsic and extrinsic factors of residences and their location, respectively, on real estate prices. Using regression techniques, these models quantify the impact of each factor on the price of a house. The identification of relevant factors, the selection of the regression formulation, and the application of the model to real-world data constitute the three general steps in developing such models. In this paper, we consider generalized additive models (GAMs) for the valuation of completed sale transactions of homes based on intrinsic and extrinsic factors of the residences. We develop a methodology for identifying the most significant contributory factors for home prices by evaluating the changes in adjusted values resulting from removing selected factors. Combining this with information from correlation matrices, we identify significant factors that provide little added predictive value. We apply this methodology to home prices in three U.S. cities.
For this study, three cities were chosen to represent variations in geography, primary economic activity, and population density and size: Denver (CO), Jacksonville (FL), and Phoenix (AZ). Phoenix represents the American southwest, Denver the interior, and Jacksonville the southeast. Phoenix is a nationwide leader in semiconductor and advanced manufacturing, Denver is a leading professional and technology services and energy center, and Jacksonville is a major cargo and distribution hub for the southeastern US. The population densities of Jacksonville, Phoenix, and Denver are 490, 1198, and 1805 people per square kilometer, respectively. Data was restricted to home sales within each city boundary. This choice reflects the fact that city boundaries are well established, whereas the boundaries of the suburban sprawl comprising these metropolitan areas are less precise.
The number of bedrooms and bathrooms, indoor and outdoor areas, and the categorization of the dwelling type (single-family, condominium, etc.) are commonly used and accepted intrinsic factors for real estate price models. More interesting are the variety of extrinsic factors considered. A well-known extrinsic factor is location (neighborhood desirability). Example location-related measures include postal codes and GPS coordinates, the latter providing more precise location granularity. Many publicly available geocoding websites provide the latitudinal and longitudinal coordinates of an estate, and such refinements have allowed for more extensive analyses.
Eiling et al. (
2019) used monthly housing returns for 9831 zip codes across 178 U.S. Metropolitan Statistical Areas (MSAs) to quantify the idiosyncratic zip-code-specific risk and systematic market risk within each MSA.
Hill and Scholz (
2018) established the superiority of a nonparametric spline surface based on GPS data over postal code proxy information.
Helbich et al. (
2013) examined the explanatory power of exposure to solar radiation on the pricing of owner-occupied flats in Vienna by employing airborne LIDAR maps.
Olszewski et al. (
2017) verified the significance of such factors as the distances to the nearest metro station, green space, and the city center.
Cohen and Coughlin (
2008) studied the effects of home proximity to airports.
Extrinsic macroeconomic factors have an effect on real estate prices. In their study,
Olszewski et al. (
2017) also analyzed the effects of housing policy on prices.
Belke and Keil (
2017) investigated several macroeconomic factors, including the per capita number of newly constructed apartments, the per capita number of real estate market transactions, the unemployment rate, the purchasing power index of the area, and the number of hospitals.
Environmental, social, and governance (ESG) components represent the sustainability factors of a property. The risk of a natural disaster, the installation of renewable energy systems, and resiliency to global warming are examples of environmental factors. Construction worker labor standards, homeowner satisfaction, and noise pollution are examples of social factors. Regulatory compliance with standards set at all governmental levels, overall transparency, and legal issues related to property owner practices are examples of governance factors.
Lauper et al. (
2013) analyzed the green home acquisition and installation process from the point of view of a homebuilder. Social norms and policies have been shown to not only heighten consumer spending and interest in environmentally friendly appliances but also significantly impact the energy-relevant decisions made during homebuilding (
Reposa, 2009;
Palm, 2017;
Rakha et al., 2018).
Ma et al. (
2019) analyzed the impact of governmental policymaking processes on the adoption of residential green energy additions and construction. Specifically, they noted that stringent governmental policies on residential green energy subsidies can have an adverse effect on household installations.
Under global climate change, environmental factors (flood risk, wildfires, etc.) can be expected to play a role in homebuyer decisions and, as a consequence, real estate pricing.
Lavaine (
2019) found that the closure of a toxic site leading to a decrease in atmospheric SO
2 levels was associated with an increase in the average house price but a decrease in the average price of flats. Quantitative environmental indices have been developed to provide guidance to consumers in assessing house prices.
Mahanama et al. (
2021) formulated an index to measure the level of future systemic risk caused by natural disasters. The study by
Contat et al. (
2023) confirmed that the risks of wildfire and flooding correlated inversely with home prices, as the higher risks resulted in discounts on said prices. Thus, the inclusion of ESG factors in home price models is desirable.
Beyond direct contact with a real estate agent, online real estate databases such as those provided by Zillow, Realtor.com, and Redfin are the most common resources used by potential home buyers. Accordingly, our study focuses on information these websites provide. Unfortunately, information on ESG factors in such databases is limited (see, e.g.,
Brown, 2025); using the Redfin database, we were able to obtain information only on the three ESG-related factors discussed in
Section 2.
As hedonic models aim to estimate the contributory value of each external or internal factor, the decomposition allows for the appropriate use of generalized additive, logarithmic, or linear models to identify the contributive power of each factor.
Pace (
1998) was one of the earliest to use a GAM in the context of real estate pricing and demonstrated that GAMs could outperform more unsophisticated polynomial and parametric models.
Owusu-Ansah (
2011) presented a review of semi-parametric, parametric, and non-parametric models.
Silver (
2016) proposed a hedonic regression pricing methodology.
Colonnello et al. (
2021) considered a linear hedonic model for housing yield (rent-to-price ratio) and incorporated a relatively large number of demographic, extrinsic, and local economic factors.
Brunauer et al. (
2013) used a four-level hierarchical additive regression model to quantify the contribution of each level of geographic detail to housing prices.
Bárcena et al. (
2013) employed a geographically weighted and semi-parametric hedonic model to create an index of housing prices in Bilbao, Spain, over the time period before and after the Great Recession.
Bax and Chasomeris (
2019) used a generalized linear model (GLM) to measure apartment rent prices from a set of statistically significant factors.
Doszyń and Gnat (
2017) used predictive and studentized residuals of a properly specified linear model to regress the price per square meter of plots of land to six factors. However, many models are not properly specified and correctly applied, which can result in models of poor quality, an issue especially prevalent with linear regression models. A series of simulations involving varying levels of price disturbances in multiple linear regression models found that even minor disturbances meaningfully reduced
values (
Kokot & Gnat, 2019). Linear models can only be expected to perform adequately in well-developed and well-functioning real estate markets whose influencing factors on real estate valuations exhibited strong and linear relationships.
Bailey et al. (
2022) used intrinsic and extrinsic factors (including some ESG factors) in GAMs and GLMs to analyze the variance in the logarithm of the expected sales price of homes. When ESG factors easily accessible from real estate vendor websites were included, minor improvements were observed in the adjusted
values of the model. A further analysis using these ESG factors and new home constructions to estimate the average annual home prices of eight U.S. cities over two decades found that the ESG factors had city-dependent significance in predictive power (
Bailey et al., 2024). In both studies, GAMs were found to significantly outperform GLMs.
We note that intrinsic and extrinsic factors have been shown to have varying impacts on home valuations within the distribution of the local housing prices. Analyses of home sales found that segmentation by price quantile was vital in the assessment of the impact of several input factors (
Zeitz et al., 2007,
2008).
Our first results (
Section 3.1) evaluate the factor contributions using a P-spline GAM containing no interaction terms. We identified the relative significance of each factor by evaluating the change in adjusted
value resulting from its removal from the model. We combined this with information from correlation matrices to identify the added predictive value of each factor. Based on data for the three U.S. cities, the GAM consistently achieved higher adjusted
values compared to a benchmark generalized linear model and identified all factors as statistically significant at the 0.5% level. The tests showed that living area and location (latitude, longitude) were the three most significant factors, and each added an independent predictive value. The number of bedrooms and bathrooms was strongly correlated with living area; consequently, these two variables, which are commonly included in hedonic housing models, contribute little additional predictive value. The year of construction completion ranked fourth or fifth in significance across cities and consistently showed a moderate correlation with the number of bathrooms. Of the ESG factors, elderly/disabled accessibility ranked fifth in significance for one city. For each city, this accessibility factor had a moderate, negative correlation with the number of bathrooms.
In
Section 3.2 we consider the concurvity of the GAM used in
Section 3.1. Using the correlation and prediction results of
Section 3.1 as guidance, we propose a GAM with improved concurvity via the introduction of interaction terms and orthogonalized smoothed terms. The new model was then tested on the data from the three cities. Both concurvity and explained deviance values improved under the new GAM.
3. Results
3.1. Models with No Interaction Terms
The models were run using the R
lm function and the
mgcv package (version 1.9-4) (
Wood, 2025). Specifically, using the notation of the
mgcv package, the P-spline GAM computed was
where “ps” denotes P-spline functions. The default value of ten knots was used. The restricted maximum likelihood method (REML) was used for smooth parameter selection. The
-values associated with the factors for each of the three cities as fit by the models are shown in
Table 2.
The number of factors with -values below 0.01 was larger for the GAMs than the GLMs in all three cities. Additionally, every factor was found to be significant at a -value threshold of 0.01 for all three GAMs, whereas each GLM had at least one factor -value in excess of 0.05. Six to ten percentage point increases occurred in the adjusted for all three cities under the GAM.
In order to assess the contribution of each factor to the predictive power of the models, each GAM was rerun with one of the ten factors excluded. We computed
defined as the GAM baseline adjusted
from
Table 2 minus the adjusted
from rerunning the model with one factor dropped. Thus, a positive value of
corresponds to a decrease in the adjusted
when the factor is removed. The resulting changes are shown in
Table 3. For each city, the factors have been listed in order of the magnitude of
. The ESG factors are in green font.
We define the significance of a factor by its impact. The home square footage and its latitude and longitude are the three most significant pricing factors for all three cities. The latter two factors reflect the well-known importance of location in housing price, while square footage reflects the costs of construction, maintenance, and luxury. What begins to differentiate the cities are the fourth and fifth most significant factors. While the year of construction (reflecting age as well as the change in society’s housing “tastes” and preferences with time) constitutes one of these two factors for each city, the other factor varies by city. The number of bathrooms has stronger significance than year of construction for Denver. For Phoenix, lot size is just as significant a factor as year of construction. Interestingly, the accessibility ESG factor is the fifth factor of significance for Jacksonville.
The five least significant factors are populated (with the noted exception of accessibility for Jacksonville) by the ESG factors, as well as number of bedrooms, number of bathrooms (with the noted exception of Denver), and lot size (with the noted exception of Phoenix). While the number of bedrooms and bathrooms always appear in real estate listings (reflecting the importance of family size to a potential buyer), their impact on house pricing is not all that significant for the selected cities. Jacksonville was the only city with a measurable for all three ESG factors.
With
used as a measure of factor significance, we considered factor–factor correlations to measure added predictive value. For each city, we computed the correlation matrix
,
,
, where
is the Pearson correlation coefficient between factor observations
and
with
denoting the sample covariance; and
and
the sample standard deviations. The correlation matrices for each city are displayed in
Table A2,
Table A3 and
Table A4 in
Appendix B. In assessing the Pearson correlation values, we adopt the qualitative descriptions of the strength of association (of a pair of variables) summarized in
Table 4.
If the strength of association between factors and is strong (), then used together the two factors add little predictive value, so one could be dropped from the model. If their association is weak (), both factors are needed. In the case of moderate association (), either both can be kept, or a different combination of the two factors should be considered.
Table A2,
Table A3 and
Table A4 provide the correlation values
,
, and
for the three most significant factors. All values are weak, with the exception of
for Denver, which is moderate. The relevant correlation values
,
,
and
for the
th and 5th significant factors are given in data rows 4 and 5 of
Table A2,
Table A3 and
Table A4. It is especially notable that all correlation values excluding Beds and Baths are negligible for Jacksonville (
Table A3). For Phoenix (
Table A4), both
and
have moderate values, perhaps indicative of a general construction increase in house size over time combined with movement out of a denser city center.
Finally, we consider the correlation values (last five data rows of
Table A2,
Table A3 and
Table A4) corresponding to the least significant factors. The data indicate a very strong value for
and strong values of
and
for all three cities. In particular, the very strong
values indicate that the number of bathrooms does not add a great deal of predictive power to the GAM. This is reinforced by the observation that
has a moderate correlation. The strong correlation values
and
are an obvious reflection of construction. The strong correlation values of
reflect family dynamics. Since the number of bedrooms and bathrooms is not generating significant
values (with the exception of Baths for Denver), their inclusion in the GAM does not add significant predictive value.
For Denver, all other correlations are negligible or weak except and . The magnitudes are moderate; interestingly, the signs of both correlations are negative. For Jacksonville, all other correlations are negligible or weak except for and , which are moderate in magnitude. For Phoenix, all other correlations are negligible or weak except for , and . For all three cities, is negative. Except for Phoenix, lot size had low significance but did add predictive value. For Phoenix, lot size was competitive in significance with year of construction and was moderately correlated with living area and number of bathrooms.
3.2. Concurvity and a GAM with Interaction Terms
Concurvity for non-linear functions is the analogy to collinearity for linear functions (
Wood, 2017). In the
mgcv package, each smooth function
in (1) is regressed against the other smooth functions, and three measures of concurvity are presented: estimated, observed and worst case (
Wood, 2017). These results are presented as pairwise regressions,
on
,
, or as full regressions
. We concentrate on presenting results from the full regressions. The estimated values measure concurvity after penalization and smoothing parameter selection and reflect the extent of (nonlinear) redundancy remaining in the fitted model. Observed concurvity values capture the redundancy in the data as sampled, before smoothing. It helps to explain why concurvity exists and whether it is structural in the data. Worst case concurvity is the theoretical upper bound on redundancy. It can flag problems that never materialize in practice, it ignores penalties and fitted coefficients, and it can produce high values even in well-behaved models. It reflects potential risk and is not often evidence of an actual problem. We consider estimated concurvity to be the primary diagnostic, with observed concurvity as secondary. Typical (though not strict) heuristic threshold values for concurvity are: negligible (<0.3), moderate (0.3–0.7), and high (>0.7).
Table 5 presents the full regression concurvity values for the P-spline GAM used in
Section 3.1. The table also gives the deviance explained (Dev Exp) for each city. While the percent deviance explained (68–77%) is relatively high, it is apparent that moderate estimated and observed concurvity redundancy is predominant in this model.
The redundancy in concurvity is due to the correlations observed in
Section 3.1. There are various tactics that can be adopted to reduce these dependencies. We adopt two strategies. The first is to replace smoothing functions of several interdependent variables with a single smooth of a composite (tensor product) variable. The second, used in combination with the first, is to linearly orthogonalize interdependent factors.
Guided by the correlations noted in
Section 3.1, we ran the sequential linear orthogonalizations:
Note that, in (9) and (10), we have not included any orthogonalization of Baths and Beds against Lot size, but we have orthogonalized Beds against the residuals .
We considered the following GAM:
This GAM specifically accounts for the correlations between continuous variables noted in
Section 3.1. It does not, however, address the binary variables and their correlations with continuous variables. The implementation of the GAM (11) via the
mgcv package used the formula
In (12), res_SqFt refers to the values of , with similar notation for the other three residuals used in the fit. The notation “tp” denotes thin plate spline functions, which were employed as they are translation and rotation-invariant as required for the tensor product forms in (12). The REML method was used for smooth parameter selection.
Table 6 presents the deviation explained as well as the concurvity measures for the composite terms. Compared to the model of
Table 5, the explained deviation has increased by 2.5% (JAX) to 5.1% (DEN). Estimated and observed values of concurvity are negligible or weakly moderate. Worst-case values have improved considerably for PHX and JAX, although DEV still has some strong, worst case concurvity values.
4. Discussion
The significance and correlation tests revealed that living area and latitude–longitude location exerted the strongest impact on the explained variance of the GAM, and each independently adds predictive value. The number of bedrooms and bathrooms correlated strongly with living area, indicating that these two factors, commonly included in hedonic housing models, contribute little predictive value. The year of construction completion ranked fourth or fifth in significance across cities, consistently having moderate correlation with the number of bathrooms (and additional moderate correlation with square footage and longitude in the case of Denver).
Of the ESG factors, the waterfront and green energy binary factors had low significance but did add predictive value independent of the other factors. The ESG factors had greater significance for Jacksonville than for the other two cities, with the accessibility factor being the fifth most significant. Arizona and Florida (home to Phoenix and Jacksonville, respectively) are both well-known retirement states. While 14.6% of the population of Jacksonville and 11.9% of Phoenix are older adults (65+), that difference alone does not explain why accessibility is more significant in Jacksonville (
than in Phoenix
. In fact, Denver has an older-adult population percentage of 12.3%, slightly larger than Phoenix, but its accessibility factor has no significant effect on the explained variance in the GAM fit. The accessibility difference may be because the average sold home in Jacksonville is not as new as those in Denver or Phoenix, which makes accessibility differences more pronounced. Furthermore, Jacksonville’s relative significance of the waterfront factor may reflect Jacksonville’s acreage bordering the Atlantic Ocean and the St. John’s River (which meanders through the city). Moreover, the negative correlation between Access and Baths may correspond to the fact that older-adult homes often correspond to retired, “empty nesters” who have downsized their living quarters. Our concurvity results (
Table 5) indicate that these data correlations are nonlinear and require careful development of a GAM to control correlations in order to develop a model that is both predictive and enables term-wise inference. Our model (12) is a relatively good development in that direction. However, it is clear (to us) that a single model, employing the same factor set, will not perform equally well for all cities, as local factors will produce different correlations (or correlations of differing importance). It was not our intention here to develop in-depth, specific models for the three cities studied but rather to investigate the general correlation and predictive properties of real estate factors commonly represented in online real estate databases.
We acknowledge that the conclusions of this study are limited by the small sample of cities. As noted in the introduction, we have taken some care with our choice to represent variations in geography, primary economic activity, and population density and size. Nonetheless, city-by-city variations will exist in general and will necessitate some level of city-by-city model modification. As the contribution rankings are sample-dependent, a natural next step would be to continue such an analysis with more cities of further geographic, demographic, and economic variance.
One limitation of this study is that all three qualitative ESG factors do not contain any further specificity than was available in the dataset. For example, a home was coded as either not having green energy (0) or having green energy (1). However, different green energy units can have varying efficiency ratings. We would expect efficiency levels (such as solar panel quality) to factor into a home’s valuation. Suppose that a solar panel factor was stratified based on effectiveness; for example, none (0), low (1), medium (2), and high (3). While this would represent an improvement in data refinement, it presupposes a stratification into “equidistant” intervals. However, a homebuyer may view the difference between low and medium quality panels as more significant than that between medium and high-quality panels. We speculate that more precise information, such as age and/or wattage of the panels, would further strengthen the model’s predictive power, and the compression of green energy information down into a single binary for each house may explain the lack of predictive power associated with the green factor for each city.
Our prior work has considered the presence of central air conditioning as an ESG factor. Although the option existed to filter by central air conditioning, no homes were specifically identified as such in our Redfin data set for these three cities. Given that the prior literature (see the Introduction) has identified the significance of central air conditioning on home pricing, we were unfortunately unable to estimate its contribution to the adjusted values.