Next Article in Journal
The Effects of Green Innovation on Stock Liquidity: Evidence from US Companies
Next Article in Special Issue
Reassessing Residential REITs: Performance and Resilience After COVID-19
Previous Article in Journal
Board Gender Diversity and Innovation Strategies: Sectoral Effects on ESG Performance in Financial and Non-Financial Firms
Previous Article in Special Issue
Property Tax, Local Sales Tax and Business Activity in Nevada: A Spatial Analysis
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Evaluating Factor Contributions for Sold Homes

by
Jason R. Bailey
*,
W. Brent Lindquist
and
Svetlozar T. Rachev
Department of Mathematics & Statistics, Texas Tech University, Lubbock, TX 79409-1042, USA
*
Author to whom correspondence should be addressed.
J. Risk Financ. Manag. 2026, 19(2), 146; https://doi.org/10.3390/jrfm19020146
Submission received: 3 December 2025 / Revised: 10 February 2026 / Accepted: 11 February 2026 / Published: 13 February 2026
(This article belongs to the Special Issue Real Estate Finance and Risk Management)

Abstract

We evaluated the contributions of ten intrinsic and extrinsic factors readily available from website data to individual home sale prices for three major U.S. cities using a P-spline generalized additive model (GAM). We identified the relative significance of each factor by evaluating the change in the adjusted R 2 value resulting from its removal from the model. We combined this with information from correlation matrices to identify the added predictive value of a factor. For these three cities, the tests revealed that living area and location (latitude, longitude) had the strongest impact on explained variance, and each factor independently added predictive value. Relative impacts of the other factors were city-dependent. We utilized this information to develop an improved GAM with superior concurvity values. The improved GAM required the use of linear orthogonalization of factors combined with smoothing functions based on tensor products of correlated factors.

1. Introduction

Hedonic models are employed to capture the heterogeneous effects of intrinsic and extrinsic factors of residences and their location, respectively, on real estate prices. Using regression techniques, these models quantify the impact of each factor on the price of a house. The identification of relevant factors, the selection of the regression formulation, and the application of the model to real-world data constitute the three general steps in developing such models. In this paper, we consider generalized additive models (GAMs) for the valuation of completed sale transactions of homes based on intrinsic and extrinsic factors of the residences. We develop a methodology for identifying the most significant contributory factors for home prices by evaluating the changes in adjusted R 2 values resulting from removing selected factors. Combining this with information from correlation matrices, we identify significant factors that provide little added predictive value. We apply this methodology to home prices in three U.S. cities.
For this study, three cities were chosen to represent variations in geography, primary economic activity, and population density and size: Denver (CO), Jacksonville (FL), and Phoenix (AZ). Phoenix represents the American southwest, Denver the interior, and Jacksonville the southeast. Phoenix is a nationwide leader in semiconductor and advanced manufacturing, Denver is a leading professional and technology services and energy center, and Jacksonville is a major cargo and distribution hub for the southeastern US. The population densities of Jacksonville, Phoenix, and Denver are 490, 1198, and 1805 people per square kilometer, respectively. Data was restricted to home sales within each city boundary. This choice reflects the fact that city boundaries are well established, whereas the boundaries of the suburban sprawl comprising these metropolitan areas are less precise.
The number of bedrooms and bathrooms, indoor and outdoor areas, and the categorization of the dwelling type (single-family, condominium, etc.) are commonly used and accepted intrinsic factors for real estate price models. More interesting are the variety of extrinsic factors considered. A well-known extrinsic factor is location (neighborhood desirability). Example location-related measures include postal codes and GPS coordinates, the latter providing more precise location granularity. Many publicly available geocoding websites provide the latitudinal and longitudinal coordinates of an estate, and such refinements have allowed for more extensive analyses. Eiling et al. (2019) used monthly housing returns for 9831 zip codes across 178 U.S. Metropolitan Statistical Areas (MSAs) to quantify the idiosyncratic zip-code-specific risk and systematic market risk within each MSA. Hill and Scholz (2018) established the superiority of a nonparametric spline surface based on GPS data over postal code proxy information. Helbich et al. (2013) examined the explanatory power of exposure to solar radiation on the pricing of owner-occupied flats in Vienna by employing airborne LIDAR maps. Olszewski et al. (2017) verified the significance of such factors as the distances to the nearest metro station, green space, and the city center. Cohen and Coughlin (2008) studied the effects of home proximity to airports.
Extrinsic macroeconomic factors have an effect on real estate prices. In their study, Olszewski et al. (2017) also analyzed the effects of housing policy on prices. Belke and Keil (2017) investigated several macroeconomic factors, including the per capita number of newly constructed apartments, the per capita number of real estate market transactions, the unemployment rate, the purchasing power index of the area, and the number of hospitals.
Environmental, social, and governance (ESG) components represent the sustainability factors of a property. The risk of a natural disaster, the installation of renewable energy systems, and resiliency to global warming are examples of environmental factors. Construction worker labor standards, homeowner satisfaction, and noise pollution are examples of social factors. Regulatory compliance with standards set at all governmental levels, overall transparency, and legal issues related to property owner practices are examples of governance factors.
Lauper et al. (2013) analyzed the green home acquisition and installation process from the point of view of a homebuilder. Social norms and policies have been shown to not only heighten consumer spending and interest in environmentally friendly appliances but also significantly impact the energy-relevant decisions made during homebuilding (Reposa, 2009; Palm, 2017; Rakha et al., 2018). Ma et al. (2019) analyzed the impact of governmental policymaking processes on the adoption of residential green energy additions and construction. Specifically, they noted that stringent governmental policies on residential green energy subsidies can have an adverse effect on household installations.
Under global climate change, environmental factors (flood risk, wildfires, etc.) can be expected to play a role in homebuyer decisions and, as a consequence, real estate pricing. Lavaine (2019) found that the closure of a toxic site leading to a decrease in atmospheric SO2 levels was associated with an increase in the average house price but a decrease in the average price of flats. Quantitative environmental indices have been developed to provide guidance to consumers in assessing house prices. Mahanama et al. (2021) formulated an index to measure the level of future systemic risk caused by natural disasters. The study by Contat et al. (2023) confirmed that the risks of wildfire and flooding correlated inversely with home prices, as the higher risks resulted in discounts on said prices. Thus, the inclusion of ESG factors in home price models is desirable.
Beyond direct contact with a real estate agent, online real estate databases such as those provided by Zillow, Realtor.com, and Redfin are the most common resources used by potential home buyers. Accordingly, our study focuses on information these websites provide. Unfortunately, information on ESG factors in such databases is limited (see, e.g., Brown, 2025); using the Redfin database, we were able to obtain information only on the three ESG-related factors discussed in Section 2.
As hedonic models aim to estimate the contributory value of each external or internal factor, the decomposition allows for the appropriate use of generalized additive, logarithmic, or linear models to identify the contributive power of each factor. Pace (1998) was one of the earliest to use a GAM in the context of real estate pricing and demonstrated that GAMs could outperform more unsophisticated polynomial and parametric models. Owusu-Ansah (2011) presented a review of semi-parametric, parametric, and non-parametric models. Silver (2016) proposed a hedonic regression pricing methodology. Colonnello et al. (2021) considered a linear hedonic model for housing yield (rent-to-price ratio) and incorporated a relatively large number of demographic, extrinsic, and local economic factors. Brunauer et al. (2013) used a four-level hierarchical additive regression model to quantify the contribution of each level of geographic detail to housing prices. Bárcena et al. (2013) employed a geographically weighted and semi-parametric hedonic model to create an index of housing prices in Bilbao, Spain, over the time period before and after the Great Recession. Bax and Chasomeris (2019) used a generalized linear model (GLM) to measure apartment rent prices from a set of statistically significant factors.
Doszyń and Gnat (2017) used predictive and studentized residuals of a properly specified linear model to regress the price per square meter of plots of land to six factors. However, many models are not properly specified and correctly applied, which can result in models of poor quality, an issue especially prevalent with linear regression models. A series of simulations involving varying levels of price disturbances in multiple linear regression models found that even minor disturbances meaningfully reduced R 2 values (Kokot & Gnat, 2019). Linear models can only be expected to perform adequately in well-developed and well-functioning real estate markets whose influencing factors on real estate valuations exhibited strong and linear relationships.
Bailey et al. (2022) used intrinsic and extrinsic factors (including some ESG factors) in GAMs and GLMs to analyze the variance in the logarithm of the expected sales price of homes. When ESG factors easily accessible from real estate vendor websites were included, minor improvements were observed in the adjusted R 2 values of the model. A further analysis using these ESG factors and new home constructions to estimate the average annual home prices of eight U.S. cities over two decades found that the ESG factors had city-dependent significance in predictive power (Bailey et al., 2024). In both studies, GAMs were found to significantly outperform GLMs.
We note that intrinsic and extrinsic factors have been shown to have varying impacts on home valuations within the distribution of the local housing prices. Analyses of home sales found that segmentation by price quantile was vital in the assessment of the impact of several input factors (Zeitz et al., 2007, 2008).
Our first results (Section 3.1) evaluate the factor contributions using a P-spline GAM containing no interaction terms. We identified the relative significance of each factor by evaluating the change in adjusted R 2 value resulting from its removal from the model. We combined this with information from correlation matrices to identify the added predictive value of each factor. Based on data for the three U.S. cities, the GAM consistently achieved higher adjusted R 2 values compared to a benchmark generalized linear model and identified all factors as statistically significant at the 0.5% level. The tests showed that living area and location (latitude, longitude) were the three most significant factors, and each added an independent predictive value. The number of bedrooms and bathrooms was strongly correlated with living area; consequently, these two variables, which are commonly included in hedonic housing models, contribute little additional predictive value. The year of construction completion ranked fourth or fifth in significance across cities and consistently showed a moderate correlation with the number of bathrooms. Of the ESG factors, elderly/disabled accessibility ranked fifth in significance for one city. For each city, this accessibility factor had a moderate, negative correlation with the number of bathrooms.
In Section 3.2 we consider the concurvity of the GAM used in Section 3.1. Using the correlation and prediction results of Section 3.1 as guidance, we propose a GAM with improved concurvity via the introduction of interaction terms and orthogonalized smoothed terms. The new model was then tested on the data from the three cities. Both concurvity and explained deviance values improved under the new GAM.

2. Materials and Methods

2.1. Price and Factor Data

Data were assembled for the cities of Denver (DEN), Jacksonville (JAX), and Phoenix (PHX). Our data set1 was based on completed sale transactions of homes within the 36-month period preceding the end of 2024 (see Appendix A for details on the collection process). Table 1 shows the number of homes sold (within the city limits for each city) for each of the three years. Sales per year were essentially constant in 2022 and 2023 and rose in 2024. There was no significant city-to-city difference in percent of total sales for each year.
The data set consisted of dwelling prices and ten factors. Seven are intrinsic factors: living area (SqFt), lot size (Lot), the number of bedrooms (Beds), the number of bathrooms (Baths), the year during which the construction of the dwelling was completed (Year), whether the home was green-rated (Green), and whether the home was considered accessible to the elderly and disabled (Access). Three are extrinsic location factors: latitude (Lat), longitude (Long), and whether the home was along a waterfront (Water). We note that the data are restricted to home sales within city limits. Due to the heavy-tailed nature of dwelling prices, we used log10(Price) to express dwelling price (log-price). The seven non-binary factors (SqFt, Lot, Beds, Baths, Lat, Long, and Year) are normalized for this study by computing the sample mean and standard deviation for each factor and converting each data value to a z -score. The three ESG factors (Green, Access, and Water) are only available from Redfin.com as binary (yes/no) factors.

2.2. Generalized Additive and Linear Models

A GAM relates a univariate response variable Y to a set of predictive (intrinsic and extrinsic) factors x j ,   j = 1 , ,   m (Hastie & Tibshirani, 1990). It relates the expected value μ = E [ Y ] to the factors through functional dependencies
g μ = β 0 + f 1 x 1 + f 2 x 2 + + f m x m .
It is assumed that Y ~ E F ( μ , θ ) , where E F ( μ , θ ) denotes the exponential family of distributions with mean μ and scale parameter θ . The link function g ( · ) relates conditional expectations of Y to the factors through
μ = g 1 β 0 + f 1 x 1 + f 2 x 2 + + f m x m .
We used the identity function for g ( · ) and P-splines (Eilers & Marx, 1996) for the functions f j ( · ) , which minimize the penalized sum of squares
i = 1 N Y i j = 1 m f j x j i 2 + j = 1 m λ j f j z 2 d z ,
where Y i , i = 1 ,   ,   N , and x j i , j = 1 ,   ,   m , i = 1 ,   ,   N , are N data observations. The weight given to the smoothness of function f j ( · ) is determined by the tuning parameter λ j > 0 . The values x j i , i = 1 ,   ,   N , are referred to as the knots of the function f j ( · ) . In Section 3.2, we utilize thin plate splines (Duchon, 1977), which minimize a penalized sum of squares analogous to that specified by (3) but with the important distinction that higher derivative orders can be employed in the penalty term. Additionally, thin plate splines are useful for tensor product terms of the form f x i , x j , i j , where mixed derivative terms also appear in the penalty term.
As a benchmark, we compared the results obtained from the P-spline GAM to those obtained from a standard GLM of the form
g ( E Y Y X ) = β 0 + β 1 x 1 + + β m x m X β .
In (4), X is an N × ( m + 1 ) matrix, Y =   Y 1 ,   ,   Y N T is the column vector of values of the response variable, x j =   x j 1 ,   ,   x j N T is the column vector of values for factor x j , and β =   β 0 , β 1   ,   β m T is the column vector of unknown parameters. The first column of X is a vector of ones, and the remaining columns correspond to the factor column vectors. As the identity function was used for g     , (4) becomes a pure linear model.

3. Results

3.1. Models with No Interaction Terms

The models were run using the R lm function and the mgcv package (version 1.9-4) (Wood, 2025). Specifically, using the notation of the mgcv package, the P-spline GAM computed was
log10(Price) ~ s(SqFt, bs = “ps”) + s(Lot, bs = “ps”) + s(Beds, bs =  
     “ps”) + s(Baths, bs = “ps”) + s(Lat, bs = “ps”) + s(Long, bs = “ps”)
+ s(Year, bs = “ps”) + Water + Green + Access,      
where “ps” denotes P-spline functions. The default value of ten knots was used. The restricted maximum likelihood method (REML) was used for smooth parameter selection. The p -values associated with the factors for each of the three cities as fit by the models are shown in Table 2.
The number of factors with p -values below 0.01 was larger for the GAMs than the GLMs in all three cities. Additionally, every factor was found to be significant at a p -value threshold of 0.01 for all three GAMs, whereas each GLM had at least one factor p -value in excess of 0.05. Six to ten percentage point increases occurred in the adjusted R 2 for all three cities under the GAM.
In order to assess the contribution of each factor to the predictive power of the models, each GAM was rerun with one of the ten factors excluded. We computed R 2 defined as the GAM baseline adjusted R 2 from Table 2 minus the adjusted R 2 from rerunning the model with one factor dropped. Thus, a positive value of R 2 corresponds to a decrease in the adjusted R 2 when the factor is removed. The resulting changes are shown in Table 3. For each city, the factors have been listed in order of the magnitude of R 2 . The ESG factors are in green font.
We define the significance of a factor by its R 2 impact. The home square footage and its latitude and longitude are the three most significant pricing factors for all three cities. The latter two factors reflect the well-known importance of location in housing price, while square footage reflects the costs of construction, maintenance, and luxury. What begins to differentiate the cities are the fourth and fifth most significant factors. While the year of construction (reflecting age as well as the change in society’s housing “tastes” and preferences with time) constitutes one of these two factors for each city, the other factor varies by city. The number of bathrooms has stronger significance than year of construction for Denver. For Phoenix, lot size is just as significant a factor as year of construction. Interestingly, the accessibility ESG factor is the fifth factor of significance for Jacksonville.
The five least significant factors are populated (with the noted exception of accessibility for Jacksonville) by the ESG factors, as well as number of bedrooms, number of bathrooms (with the noted exception of Denver), and lot size (with the noted exception of Phoenix). While the number of bedrooms and bathrooms always appear in real estate listings (reflecting the importance of family size to a potential buyer), their impact on house pricing is not all that significant for the selected cities. Jacksonville was the only city with a measurable R 2 for all three ESG factors.
With R 2 used as a measure of factor significance, we considered factor–factor correlations to measure added predictive value. For each city, we computed the correlation matrix R = r i j , i = 1 , ,   10 , j = 1 , ,   10 , where
r i j = c o v x i , x j σ x i σ x j
is the Pearson correlation coefficient between factor observations x i and x j with c o v x i , x j denoting the sample covariance; and σ x i and σ x j the sample standard deviations. The correlation matrices for each city are displayed in Table A2, Table A3 and Table A4 in Appendix B. In assessing the Pearson correlation values, we adopt the qualitative descriptions of the strength of association (of a pair of variables) summarized in Table 4.
If the strength of association between factors i and j is strong ( r i , j 0.50 ), then used together the two factors add little predictive value, so one could be dropped from the model. If their association is weak ( r i , j < 0.30 ), both factors are needed. In the case of moderate association ( r i , j 0.30 , 0.50 ), either both can be kept, or a different combination of the two factors should be considered.
Table A2, Table A3 and Table A4 provide the correlation values r S q F t , L a t , r S q F t , L o n g , and r L a t , L o n g for the three most significant factors. All values are weak, with the exception of r L a t , L o n g = 0.33 for Denver, which is moderate. The relevant correlation values r i , S q F t , r i , L o n g , r i , L a t and r 5,4 for the i = 4 th and 5th significant factors are given in data rows 4 and 5 of Table A2, Table A3 and Table A4. It is especially notable that all correlation values excluding Beds and Baths are negligible for Jacksonville (Table A3). For Phoenix (Table A4), both r Y e a r , S q F t and r L o t , S q F t have moderate values, perhaps indicative of a general construction increase in house size over time combined with movement out of a denser city center.
Finally, we consider the correlation values (last five data rows of Table A2, Table A3 and Table A4) corresponding to the least significant factors. The data indicate a very strong value for r B a t h s , S q F t and strong values of r B e d s , S q F t and r B a t h s , B e d s for all three cities. In particular, the very strong r B a t h s , S q F t values indicate that the number of bathrooms does not add a great deal of predictive power to the GAM. This is reinforced by the observation that r Y e a r , B a t h s has a moderate correlation. The strong correlation values r B a t h s , S q F t and r B e d s , S q F t are an obvious reflection of construction. The strong correlation values of r B a t h s , B e d s reflect family dynamics. Since the number of bedrooms and bathrooms is not generating significant R 2 values (with the exception of Baths for Denver), their inclusion in the GAM does not add significant predictive value.
For Denver, all other correlations are negligible or weak except r A c c e s s , S q F t and r A c c e s s , B a t h . The magnitudes are moderate; interestingly, the signs of both correlations are negative. For Jacksonville, all other correlations are negligible or weak except for r B a t h s , Y e a r and r B a t h s , A c c e s s , which are moderate in magnitude. For Phoenix, all other correlations are negligible or weak except for r B a t h s , Y e a r , r B a t h s , L o t and r A c c e s s , B a t h s . For all three cities, r A c c e s s , B a t h s is negative. Except for Phoenix, lot size had low significance but did add predictive value. For Phoenix, lot size was competitive in significance with year of construction and was moderately correlated with living area and number of bathrooms.

3.2. Concurvity and a GAM with Interaction Terms

Concurvity for non-linear functions is the analogy to collinearity for linear functions (Wood, 2017). In the mgcv package, each smooth function f i · in (1) is regressed against the other smooth functions, and three measures of concurvity are presented: estimated, observed and worst case (Wood, 2017). These results are presented as pairwise regressions, f i · on f j · , j i , or as full regressions f i · = j i f j · . We concentrate on presenting results from the full regressions. The estimated values measure concurvity after penalization and smoothing parameter selection and reflect the extent of (nonlinear) redundancy remaining in the fitted model. Observed concurvity values capture the redundancy in the data as sampled, before smoothing. It helps to explain why concurvity exists and whether it is structural in the data. Worst case concurvity is the theoretical upper bound on redundancy. It can flag problems that never materialize in practice, it ignores penalties and fitted coefficients, and it can produce high values even in well-behaved models. It reflects potential risk and is not often evidence of an actual problem. We consider estimated concurvity to be the primary diagnostic, with observed concurvity as secondary. Typical (though not strict) heuristic threshold values for concurvity are: negligible (<0.3), moderate (0.3–0.7), and high (>0.7).
Table 5 presents the full regression concurvity values for the P-spline GAM used in Section 3.1. The table also gives the deviance explained (Dev Exp) for each city. While the percent deviance explained (68–77%) is relatively high, it is apparent that moderate estimated and observed concurvity redundancy is predominant in this model.
The redundancy in concurvity is due to the correlations observed in Section 3.1. There are various tactics that can be adopted to reduce these dependencies. We adopt two strategies. The first is to replace smoothing functions of several interdependent variables with a single smooth of a composite (tensor product) variable. The second, used in combination with the first, is to linearly orthogonalize interdependent factors.
Guided by the correlations noted in Section 3.1, we ran the sequential linear orthogonalizations:
E S q F t = β 10 + β 11 x L a t + β 12 x L o n g + β 13 x Y e a r + β 14 x L a t x L o n g + β 15 x L a t x Y e a r + β 16 x L o n g x Y e a r + β 17 x L a t x L o n g x Y e a r + ε S q F t .
E L o t = β 20 + β 21 x L a t + β 22 x L o n g + β 23 x Y e a r + β 24 x L a t x L o n g + β 25 x L a t x Y e a r + β 26 x L o n g x Y e a r + β 27 x L a t x L o n g x Y e a r + β 28 ε S q F t   + ε L o t ,
E B a t h s = β 30 + β 31 x L a t + β 32 x L o n g + β 33 x Y e a r + β 34 x L a t x L o n g   + β 35 x L a t x Y e a r + β 36 x L o n g x Y e a r + β 37 x L a t x L o n g x Y e a r   + β 38 ε S q F t   + ε B a t h s ,
E B e d s = β 40 + β 41 x L a t + β 42 x L o n g + β 43 x Y e a r + β 44 x L a t x L o n g + β 45 x L a t x Y e a r + β 46 x L o n g x Y e a r + β 47 x L a t x L o n g x Y e a r + β 48 ε S q F t + β 49 ε B a t h s +   ε B e d s .
Note that, in (9) and (10), we have not included any orthogonalization of Baths and Beds against Lot size, but we have orthogonalized Beds against the residuals ε B a t h s .
We considered the following GAM:
E l o g 10 P r i c e = β 0 + f 1 x L a t + f 2 x L o n g + f 3 x Y e a r + f 4 x L a t x L o n g + f 5 x L a t x Y e a r + f 6 x L o n g x Y e a r + f 7 x L a t x L o n g x Y e a r + f 8 ε S q F t + f 9 ε L o t + f 10 ε S q F t ε L o t + f 11 ε B a t h s + f 12 ε B e d s + f 13 ε B a t h s ε B e d s .
This GAM specifically accounts for the correlations between continuous variables noted in Section 3.1. It does not, however, address the binary variables and their correlations with continuous variables. The implementation of the GAM (11) via the mgcv package used the formula
log10(Price) ~ te(Lat, Long, Year, bs = c(“tp”, “tp”, “tp”), k = c(5,  
  5, 5)) + te(res_SqFt, res_Lot, bs = c(“tp”, “tp”)) + te(res_Baths,
res_Beds, bs = c(“tp”, “tp”)) + Water + Green + Access. 
In (12), res_SqFt refers to the values of ε S q F t , with similar notation for the other three residuals used in the fit. The notation “tp” denotes thin plate spline functions, which were employed as they are translation and rotation-invariant as required for the tensor product forms in (12). The REML method was used for smooth parameter selection.
Table 6 presents the deviation explained as well as the concurvity measures for the composite terms. Compared to the model of Table 5, the explained deviation has increased by 2.5% (JAX) to 5.1% (DEN). Estimated and observed values of concurvity are negligible or weakly moderate. Worst-case values have improved considerably for PHX and JAX, although DEV still has some strong, worst case concurvity values.

4. Discussion

The significance and correlation tests revealed that living area and latitude–longitude location exerted the strongest impact on the explained variance of the GAM, and each independently adds predictive value. The number of bedrooms and bathrooms correlated strongly with living area, indicating that these two factors, commonly included in hedonic housing models, contribute little predictive value. The year of construction completion ranked fourth or fifth in significance across cities, consistently having moderate correlation with the number of bathrooms (and additional moderate correlation with square footage and longitude in the case of Denver).
Of the ESG factors, the waterfront and green energy binary factors had low significance but did add predictive value independent of the other factors. The ESG factors had greater significance for Jacksonville than for the other two cities, with the accessibility factor being the fifth most significant. Arizona and Florida (home to Phoenix and Jacksonville, respectively) are both well-known retirement states. While 14.6% of the population of Jacksonville and 11.9% of Phoenix are older adults (65+), that difference alone does not explain why accessibility is more significant in Jacksonville ( R 2 = 0.014 ) than in Phoenix ( R 2 = 0.003 ) . In fact, Denver has an older-adult population percentage of 12.3%, slightly larger than Phoenix, but its accessibility factor has no significant effect on the explained variance in the GAM fit. The accessibility difference may be because the average sold home in Jacksonville is not as new as those in Denver or Phoenix, which makes accessibility differences more pronounced. Furthermore, Jacksonville’s relative significance of the waterfront factor may reflect Jacksonville’s acreage bordering the Atlantic Ocean and the St. John’s River (which meanders through the city). Moreover, the negative correlation between Access and Baths may correspond to the fact that older-adult homes often correspond to retired, “empty nesters” who have downsized their living quarters. Our concurvity results (Table 5) indicate that these data correlations are nonlinear and require careful development of a GAM to control correlations in order to develop a model that is both predictive and enables term-wise inference. Our model (12) is a relatively good development in that direction. However, it is clear (to us) that a single model, employing the same factor set, will not perform equally well for all cities, as local factors will produce different correlations (or correlations of differing importance). It was not our intention here to develop in-depth, specific models for the three cities studied but rather to investigate the general correlation and predictive properties of real estate factors commonly represented in online real estate databases.
We acknowledge that the conclusions of this study are limited by the small sample of cities. As noted in the introduction, we have taken some care with our choice to represent variations in geography, primary economic activity, and population density and size. Nonetheless, city-by-city variations will exist in general and will necessitate some level of city-by-city model modification. As the contribution rankings are sample-dependent, a natural next step would be to continue such an analysis with more cities of further geographic, demographic, and economic variance.
One limitation of this study is that all three qualitative ESG factors do not contain any further specificity than was available in the dataset. For example, a home was coded as either not having green energy (0) or having green energy (1). However, different green energy units can have varying efficiency ratings. We would expect efficiency levels (such as solar panel quality) to factor into a home’s valuation. Suppose that a solar panel factor was stratified based on effectiveness; for example, none (0), low (1), medium (2), and high (3). While this would represent an improvement in data refinement, it presupposes a stratification into “equidistant” intervals. However, a homebuyer may view the difference between low and medium quality panels as more significant than that between medium and high-quality panels. We speculate that more precise information, such as age and/or wattage of the panels, would further strengthen the model’s predictive power, and the compression of green energy information down into a single binary for each house may explain the lack of predictive power associated with the green factor for each city.
Our prior work has considered the presence of central air conditioning as an ESG factor. Although the option existed to filter by central air conditioning, no homes were specifically identified as such in our Redfin data set for these three cities. Given that the prior literature (see the Introduction) has identified the significance of central air conditioning on home pricing, we were unfortunately unable to estimate its contribution to the adjusted R 2 values.

Author Contributions

Conceptualization, J.R.B. and S.T.R.; methodology, all authors; software, J.R.B.; validation, all authors; formal analysis, all authors; investigation, J.R.B. and W.B.L.; data curation, J.R.B.; writing—original draft preparation, J.R.B.; writing—review and editing, W.B.L.; visualization, J.R.B. and W.B.L.; supervision, S.T.R. and W.B.L. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

This study’s data and source code are available upon request to the corresponding author.

Conflicts of Interest

The authors declare no conflicts of interest.

Appendix A. Data Collection

Data were downloaded via a Python 3 script accessing RedFin’s application programming interface (API). The data-filter values specified in the script are provided in Table A1. As RedFin limits each API request to ten thousand homes, the script collected the data via repeated requests with each restricted to an ascending price range until data for all homes for all three cities was obtained. As per Table A1, this required at least 13 API requests.
Table A1. Filter values used for the RedFin data.
Table A1. Filter values used for the RedFin data.
FilterInputFilterInput
StatusSold Last
36 Months
Price RangeMIN: $100 k,
MAX: $10 M
Number
Bedrooms
1+Number
Bathrooms
1+
Home TypeHouse
Square FeetMIN: 750,
MAX: NS 1
StoriesMIN: NS,
MAX: NS
Lot SizeMIN: 1000,
MAX: NS
Year BuiltMIN: NS,
MAX: NS
Garage SpotsNSPool TypeNS
Exclude 55+
Communities
NSBasementNS
Air ConditioningNSWasher/Dryer
Hookup
NS
FireplaceNSElevatorNS
Primary Bedroom
on Main Floor
NSPets AllowedNS
Guest HouseNSHas a ViewNS
WaterfrontESG 2Fixer-UpperNS
Green HomeESGAccessible HomeESG
1 NS = Not Specified; 2 “Yes” when filtering for those houses and “NS” otherwise.

Appendix B. Correlation Coefficient Tables

Table A2. Pearson correlation coefficients, r r o w   f a c t o r , c o l u m n   f a c t o r , for Denver.
Table A2. Pearson correlation coefficients, r r o w   f a c t o r , c o l u m n   f a c t o r , for Denver.
SqFtLongLatBathsYearLotBedsWaterGreenAccess
SqFt1.00
Long0.111.00
Lat−0.150.331.00
Baths0.830.18−0.091.00
Year0.370.420.090.441.00
Lot0.26−0.13−0.250.140.041.00
Beds0.630.05−0.110.630.260.201.00
Water0.08−0.04−0.050.050.050.050.021.00
Green0.050.000.000.050.030.010.040.001.00
Access−0.37−0.25−0.08−0.46−0.280.10−0.210.000.001.00
Colors indicate strength of association: white—negligible, weak; green—moderate; light red—strong; dark red—very strong.
Table A3. Pearson correlation coefficients, r r o w   f a c t o r , c o l u m n   f a c t o r , for Jacksonville.
Table A3. Pearson correlation coefficients, r r o w   f a c t o r , c o l u m n   f a c t o r , for Jacksonville.
SqFtLongLatYearAccessBathsWaterLotBedsGreen
SqFt1.00
Long0.181.00
Lat−0.110.001.00
Year0.250.11−0.041.00
Access−0.27−0.050.01−0.071.00
Baths0.810.17−0.110.34−0.311.00
Water0.280.12−0.010.11−0.020.221.00
Lot0.030.000.01−0.05−0.040.190.231.00
Beds0.610.06−0.010.25−0.170.600.130.131.00
Green0.030.010.000.02−0.010.030.040.020.031.00
Colors indicate strength of association: white—negligible, weak; green—moderate; light red—strong; dark red—very strong.
Table A4. Pearson correlation coefficients, r r o w   f a c t o r , c o l u m n   f a c t o r , for Phoenix.
Table A4. Pearson correlation coefficients, r r o w   f a c t o r , c o l u m n   f a c t o r , for Phoenix.
LongSqFtLatYearLotBathsBedsAccessWaterGreen
Long1.00
SqFt0.201.00
Lat0.130.131.00
Year−0.140.320.111.00
Lot0.170.460.16−0.021.00
Baths0.150.810.080.330.351.00
Beds0.020.640.030.210.230.611.00
Access−0.22−0.37−0.14−0.26−0.07−0.33−0.261.00
Water0.020.02−0.030.010.010.020.00−0.021.00
Green0.040.070.020.020.030.070.06−0.040.031.00
Colors indicate strength of association: white—negligible, weak; green—moderate; light red—strong; dark red—very strong.

Note

1
Price and factor data were obtained from Redfin.com. Data was obtained by specification of the city and the entries for “All filters” in Appendix A.

References

  1. Bailey, J. R., Lauria, D., Lindquist, W. B., Mittnik, S., & Rachev, S. T. (2022). Hedonic models of real estate prices: GAM models; environmental and sex-offender-proximity factors. Journal of Risk and Financial Management, 15, 601. [Google Scholar] [CrossRef] [Scilit]
  2. Bailey, J. R., Lindquist, W. B., & Rachev, S. T. (2024). Hedonic models incorporating environmental, social, and governance factors for time series of average annual home prices. Journal of Risk and Financial Management, 17, 375. [Google Scholar] [CrossRef] [Scilit]
  3. Bax, D., & Chasomeris, M. (2019). Listing price estimation of apartments: A generalized linear model. Journal of Economic and Financial Sciences, 12, a204. [Google Scholar] [CrossRef] [Scilit]
  4. Bárcena, M. J., Mendez, P., Palacios, M. B., & Tusell-Palmer, F. (2013). Measuring the effect of the real estate bubble: A house price index for Bilbao. In H. Augustyniak, J. Lascek, K. Olszewski, & J. Waszczuk (Eds.), Recent trends in the real estate market and its analysis (Vol. 2, pp. 127–156). Narodowy Bank Polski. [Google Scholar]
  5. Belke, A., & Keil, J. (2017). Fundamental determinants of real estate prices: A panel study of german regions. Ruhr Economic Papers, No. 731. Ruhr-University. Available online: https://link.springer.com/article/10.1007/s11294-018-9671-2 (accessed on 18 September 2024).
  6. Brown, C. (2025, November 30). Zillow removes climate risk scores from home listings. The New York Times.
  7. Brunauer, W., Lang, S., & Umlauf, N. (2013). Modeling house prices using multilevel structured additive regression. Statistical Modelling, 13, 95–123. [Google Scholar] [CrossRef] [Scilit]
  8. Cohen, J. P., & Coughlin, C. C. (2008). Spatial hedonic models of airport noise, proximity, and housing prices. Journal of Regional Science, 48, 859–878. [Google Scholar] [CrossRef] [Scilit]
  9. Colonnello, S., Marfè, R., & Xiong, Q. (2021). Housing yields. Department of Economics Working Papers, No. 21/WP/2021. University of Venice “Ca’ Foscari”. [Google Scholar] [CrossRef] [Scilit]
  10. Contat, J., Hopkins, C., Mejia, L., & Suandi, M. (2023). When climate meets real estate: A survey of the literature. Working Paper 23-05. Federal Housing Finance Agency. [Google Scholar]
  11. Doszyń, M., & Gnat, S. (2017). Econometric identification of the impact of real estate characteristics based on predictive and studentized residuals. Real Estate Management and Valuation, 25(1), 84–92. [Google Scholar] [CrossRef] [Scilit]
  12. Duchon, J. (1977). Splines minimizing rotation-invariant semi-norms in Solobev spaces. In W. Schemp, & K. Zeller (Eds.), Construction theory of functions of several variables (pp. 85–100). Springer. [Google Scholar]
  13. Eilers, P. H. C., & Marx, B. D. (1996). Flexible smoothing with B-spines and penalties. Statistical Science, 11, 89–121. [Google Scholar] [CrossRef] [Scilit]
  14. Eiling, E., Giambona, E., Aliouchkin, R. L., & Tuijp, P. (2019). Homeowners’ risk premia: Evidence from zip code housing returns. SSRN. Available online: https://ssrn.com/abstract=3312391 (accessed on 8 February 2026). [CrossRef] [Scilit]
  15. Hastie, T. J., & Tibshirani, R. J. (1990). Generalized additive models. Chapman and Hall. [Google Scholar]
  16. Helbich, M., Joachem, A., Mücke, W., & Höfle, B. (2013). Boosting the predictive accuracy of urban hedonic house price models through airborne laser scanning. Computers, Environment, and Urban Systems, 39, 81–92. [Google Scholar] [CrossRef] [Scilit]
  17. Hill, R. J., & Scholz, M. (2018). Can geospatial data improve house price indexes? A hedonic imputation approach with splines. The Review of Income and Wealth, 64, 737–756. [Google Scholar] [CrossRef] [Scilit]
  18. Kokot, S., & Gnat, S. (2019). Simulative verification of the possibility of using multiple regression models for real estate appraisal. Real Estate Management and Valuation, 27(3), 109–123. [Google Scholar] [CrossRef] [Scilit]
  19. Lauper, E., Bruppacher, S., & Kaufmann-Hayoz, R. (2013). Energy-relevant decisions of home buyers in new home construction. Umweltpsychologie, 17, 109–123. [Google Scholar]
  20. Lavaine, E. (2019). Environmental risk and differentiated housing values: Evidence from the north of France. Journal of Housing Economics, 44, 74–87. [Google Scholar] [CrossRef] [Scilit]
  21. Ma, J., Hou, A., & Tian, Y. (2019). Research on the complexity of green innovative enterprise in dynamic game model and governmental policy making. Chaos, Solitons & Fractals: X, 2, 1000008. [Google Scholar] [CrossRef] [Scilit]
  22. Mahanama, T., Shirvani, A., & Rachev, S. T. (2021). A natural disasters index. Environmental Economics and Policy Studies, 24, 263–284. [Google Scholar] [CrossRef] [Scilit]
  23. Olszewski, K., Waszczuk, J., & Widlak, M. (2017). Spatial and hedonic analysis of house price dynamics in Warsaw, Poland. Journal of Urban Planning and Development, 143, 04017009. [Google Scholar] [CrossRef] [Scilit]
  24. Owusu-Ansah, A. (2011). A review of hedonic pricing models in housing research. Journal of International Real Estate and Construction Studies, 1, 19–38. [Google Scholar]
  25. Pace, K. (1998). Appraisal using generalized additive models. Journal of Real Estate Research, 15, 77–99. [Google Scholar] [CrossRef] [Scilit]
  26. Palm, A. (2017). Peer effects in residential solar photovoltaics adoption—A mixed methods study of Swedish users. Energy Research & Social Science, 26(4), 1–10. [Google Scholar] [CrossRef] [Scilit]
  27. Rakha, T., Moss, T. W., & Shin, D. (2018). A decade analysis of residential LEED buildings market share in the United States: Trends for transitioning sustainable societies. Sustainable Cities and Society, 39(5), 568–577. [Google Scholar] [CrossRef] [Scilit]
  28. Reposa, J. H., Jr. (2009). Comparison of USGBC LEED for homes and the NAHB national green building program. International Journal of Construction Education and Research, 5(2), 108–120. [Google Scholar] [CrossRef] [Scilit]
  29. Silver, M. (2016). How to better measure hedonic residential property price indexes. IMF Working Papers, WP/16/213. International Monetary Fund. Available online: https://www.imf.org/external/pubs/ft/wp/2016/wp16213.pdf (accessed on 18 September 2024).
  30. Wood, S. N. (2017). Generalized additive models: An introduction with R (2nd ed.). Chapman and Hall/CRC. [Google Scholar]
  31. Wood, S. N. (2025). Package ‘mgcv’ (Version 1.9-4). R Foundation for Statistical Computing. Available online: https://cran.r-project.org/web/packages/mgcv/mgcv.pdf (accessed on 5 January 2026).
  32. Zeitz, J., Sirmans, G. S., & Smersh, G. T. (2008). The impact of inflation on home prices and the valuation of housing characteristics across the price distribution. Journal of Housing Research, 17, 119–138. [Google Scholar] [CrossRef] [Scilit]
  33. Zeitz, J., Zietz, E. N., & Sirmans, G. S. (2007). Determinants of house prices: A quantile regression approach. The Journal of Real Estate Finance and Economics, 37, 317–333. [Google Scholar] [CrossRef] [Scilit]
Table 1. Recorded home sales comprise the data set.
Table 1. Recorded home sales comprise the data set.
PhoenixDenverJacksonville
YearSales% TotalSales% TotalSales% Total
202222,46530.110,28430.5633831.2
202321,16228.4978329.0596929.3
202430,97941.513,65040.5802739.5
total74,606 33,717 20,334
Table 2. Significance (p-value) of the factors in the GLM and GAM fits.
Table 2. Significance (p-value) of the factors in the GLM and GAM fits.
FactorDENJAXPHXDENJAXPHX
GLMGAM
SqFt******************
Lot***0.104************
Beds******************
Baths******************
Lat******************
Long******************
Year******************
Water0.086***************
Green***0.0440.282***0.003***
Access0.0100.005************
Adj. R 2 0.6660.6230.6940.7610.6870.769
*** Indicates p-value < 0.001.
Table 3. Value of ∆R2 for the GAM fits with the stated factor removed.
Table 3. Value of ∆R2 for the GAM fits with the stated factor removed.
DENJAXPHX
SqFt0.043SqFt0.082Long0.078
Long0.040Long0.051SqFt0.036
Lat0.023Lat0.027Lat0.027
Baths0.023Year0.022Year0.023
Year0.010Access0.014Lot0.022
Lot0.002Baths0.006Baths0.005
Beds0.001Water0.005Beds0.004
Water***Lot0.004Access0.003
Green***Beds0.003Water***
Access***Green0.001Green***
*** Indicates change in adjusted R2 of <0.001.
Table 4. Strength of association descriptions.
Table 4. Strength of association descriptions.
r i , j Description
0.00 ,   0.10 Negligible
0.10 ,   0.30 Weak
0.30 ,   0.50 Moderate
0.50 ,   0.70 Strong
0.70 ,   1.00 Very Strong
Table 5. Concurvity measures 1 for the GAM (5).
Table 5. Concurvity measures 1 for the GAM (5).
MeasureDev ExpSqFtLotBedsBathsLatLongYear
DEN
Est0.7610.7030.4610.3470.5420.3870.3360.506
Obs0.7960.3810.3300.7580.4690.6510.427
Worst0.7970.5960.5470.7630.5760.6750.736
JAX
Est0.6870.5870.2190.4000.5270.2130.3010.429
Obs0.7550.3600.3250.6750.1570.2570.331
Worst0.7710.4030.4870.7480.4100.4800.607
PHX
Est0.7690.7570.3470.4110.6250.3740.2280.519
Obs0.8160.5280.4660.6680.4110.3280.503
Worst0.8420.5440.7020.8410.6510.4850.716
1 Concurvity regions—green: negligible, yellow: moderate, red: strong.
Table 6. Concurvity measures 1 for the GAM (12).
Table 6. Concurvity measures 1 for the GAM (12).
MeasureDev Expte(Lat,Long,Year) t e ε S q F t , ε L o t t e ε B a t h s , ε B e d s
DEN
Est0.8120.0800.3520.118
Obs0.2820.4480.121
Worst0.7070.8510.851
JAX
Est0.7120.0410.1730.087
Obs0.0890.3490.116
Worst0.3620.3670.243
PHX
Est0.7970.0690.2870.087
Obs0.3430.3840.117
Worst0.5410.5120.325
1 Concurvity regions—green: negligible, yellow: moderate, red: strong.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Bailey, J.R.; Lindquist, W.B.; Rachev, S.T. Evaluating Factor Contributions for Sold Homes. J. Risk Financ. Manag. 2026, 19, 146. https://doi.org/10.3390/jrfm19020146

AMA Style

Bailey JR, Lindquist WB, Rachev ST. Evaluating Factor Contributions for Sold Homes. Journal of Risk and Financial Management. 2026; 19(2):146. https://doi.org/10.3390/jrfm19020146

Chicago/Turabian Style

Bailey, Jason R., W. Brent Lindquist, and Svetlozar T. Rachev. 2026. "Evaluating Factor Contributions for Sold Homes" Journal of Risk and Financial Management 19, no. 2: 146. https://doi.org/10.3390/jrfm19020146

APA Style

Bailey, J. R., Lindquist, W. B., & Rachev, S. T. (2026). Evaluating Factor Contributions for Sold Homes. Journal of Risk and Financial Management, 19(2), 146. https://doi.org/10.3390/jrfm19020146

Article Metrics

Back to TopTop