4.1. Instrument Validity and Internal Consistency Measurements
Following the specification of the model, the next step involves establishing the psychometric properties of the scales, primarily their reliability, followed by their factorial structure to ensure a clear substantive and empirical definition of the constructs, as well as the convergent and discriminant validity of the scales [
174]. Only when the measurement instrument is confirmed to provide reliable and valid results is it possible to proceed with considering further theoretical issues.
Construct validity was examined through factor analysis. The collected data were first tested using the Kaiser–Meyer–Olkin (KMO) measure of sampling adequacy to verify the suitability of the dataset for factor analysis. The KMO coefficient indicates whether the variables in the applied scale are sufficiently correlated to justify the application of factor analysis. Values of the KMO coefficient range from 0 to 1, where values above 0.70 are considered acceptable, and those above 0.90 are considered excellent [
175,
180]. The analysis was conducted using the Statistical Package for the Social Sciences (SPSS version 20.0). Since two scales were applied in the research, the adequacy of both for factor analysis was tested. The KMO coefficient obtained for the indicators of the EXQ scale amounted to 0.969, indicating excellent sampling adequacy for factor analysis [
57] confirming that the applied measurement instrument, the EXQ scale, is both valid and reliable.
The absolute contribution of an indicator and its significance are assessed through the analysis of indicator loadings. Outer loadings refer to the individual regression coefficients between an indicator (or measurement variable) and the latent variable (construct) as estimated in the model, indicating the strength of the relationship between the observed variables (indicators) and the unobserved constructs (latent variables) they represent [
167].
The accepted rule is that a latent variable should explain a substantial portion of the variance of each indicator, typically at least 50%. This also implies that the shared variance between a construct and its indicator must exceed the variance attributable to measurement error. Outer loadings should exceed 0.708, since the square of this value (0.708
2) equals 0.50 [
169,
176]. If an indicator’s impact is not statistically significant, its corresponding loading will fall below 0.50. Statistical significance is demonstrated by loadings greater than 0.50. All indicators that do not significantly contribute to the latent variable they are intended to form or define were excluded from further analysis (
Appendix A,
Table A1).
Testing internal consistency is necessary to assess whether the questionnaire truly measures what it is intended to investigate. Reliability testing was conducted by calculating the coefficient of internal consistency, commonly known as Cronbach’s Alpha (α), as well as the values of Composite Reliability (CR) and Average Variance Extracted (AVE) for each construct (
Appendix A,
Table A1).
Cronbach’s Alpha is the most widely used coefficient for measuring the internal consistency of a questionnaire, as it provides a single estimate of reliability and indicates the extent to which the items forming a given construct are closely interrelated. The commonly accepted lower threshold for Cronbach’s Alpha is 0.70. Composite Reliability is an additional measure of reliability that, unlike Cronbach’s Alpha, does not assume equal factor loadings across items but accounts for the varying factor loadings for each indicator. The recommended threshold for Composite Reliability is 0.70 [
178].
Convergent validity reflects the degree to which multiple indicators of the same construct converge or share a high proportion of variance. The standard metric for establishing convergent validity at the construct level is Average Variance Extracted (AVE), defined as the mean of the squared loadings of indicators associated with the construct [
169]. An AVE value of 0.50 or higher indicates that, on average, the construct explains more than half of the variance of its indicators [
167].
As shown in
Appendix A,
Table A1, the values of Cronbach’s Alpha for all constructs indicate a high level of reliability of the applied scales. The reliability of the scale is satisfactory, particularly considering its length and the characteristics of the population to which it was applied. The values of CR also meet the recommended threshold, while the AVE confirms that the criterion of convergent validity has been achieved, i.e., the majority of the variance in the latent variables can be explained by their respective indicators.
Following the analysis of individual construct indicators generated from the applied measures, the overall Cronbach’s Alpha coefficient was also calculated for the MOCE scale. The obtained value of 0.966 confirms that the measurement instrument is internally consistent and valid, making it applicable in the Serbian context. As shown in
Appendix A,
Table A2, the values of Cronbach’s Alpha for all constructs indicate a high level of reliability of the applied scales. The reliability of the scale is satisfactory, particularly considering its length and the characteristics of the population to which it was applied. The CR values also meet the recommended threshold. At the same time, the AVE confirms that the criterion of convergent validity has been met, i.e., the majority of the variance in the latent variables is explained by their respective indicators.
Based on the conducted tests (
Appendix A,
Table A2), internal consistency was established, thereby confirming that the applied questionnaire, encompassing the EXQ and the MOCE measurement scales, can be reliably employed as an instrument for identifying the explanatory model of the relationship between customer experience in banking services and the marketing outcomes of customer experience in Serbia. In this way, the general aim of the research has been achieved, and the testing of the specific model explaining the relationships among constructs can be undertaken.
4.2. Structural Equation Modelling of the Inner Model
Hypotheses were tested and conclusions drawn at a 95% significance level (
p < 0.05). This level of significance represents the probability that the observed results are due to chance; the lower the
p-value, the lower the likelihood that the findings are random, thus suggesting a genuine effect or the consequence of the influence of the selected independent variables in the model [
167].
The problem of multicollinearity arises when two or more independent variables are highly correlated, meaning that the variation in one explanatory variable can be explained mainly by the variation in another explanatory variable. A fundamental assumption in classical multiple linear regression models is that no explanatory variable is a perfect linear function of another explanatory variable.
The presence of multicollinearity was assessed using the Variance Inflation Factor (VIF) for all model indicators. The acceptable lower threshold of VIF varies across authors. Kock [
181] argues that if VIF < 3.3, there is no multicollinearity among formative constructs, whereas other scholars suggest a more liberal threshold of VIF < 5, indicating that even when values fall between 3.3 and 5, multicollinearity is not present [
167,
182]. In the results obtained for the variable Customer Experience in relation to its dimensions, none of the VIF values approached the lower limit of acceptability (
Appendix B,
Table A3).
After confirming VIF values for all three independent variables and establishing the absence of multicollinearity, the analysis proceeded to assess the partial least squares analysis of the first-order model [
167] presented in
Appendix C,
Figure A1. Additionally, the coefficient of determination (R
2), which indicates the explanatory power of independent variables in a statistical model to account for the variance of the dependent variable, shows a high value R
2 = 0.998, which leads to the conclusion that only 0.2% of the variance in the latent variable Customer Experience is attributable to influences outside the model [
167]. To eliminate issues with an overly high R
2 value, a stepwise approach was conducted using latent variable scores, dividing the estimation of the model into separate steps.
After applying the Partial Least Squares (PLS) method, the bootstrapping procedure was conducted. Bootstrapping can be defined as a resampling technique that, based on the available data from an initial sample, generates a large number of new samples of the same size as the original sample by randomly drawing observations with replacement [
183]. In this way, each unit has an equal probability of being included in a sample, and once included, it is returned to the population from which it was drawn, allowing it to be selected multiple times, as its probability of selection remains constant throughout the resampling process.
The bootstrapping results presented in
Appendix C,
Figure A2, determined the path coefficients, along with an assessment of the statistical significance of the retained indicators in the model and the significance of the path coefficients themselves. The evaluation of path coefficients reflects the relevance of formative constructs in the model. Since all significance parameters equal 0 (
p < 0.001), it can be concluded that all path coefficients are highly statistically significant [
167,
184].
All three independent variables form a positive (defined by the values of β in
Table 2) and highly significant (represented by
p values in
Table 2) relationship with Customer Experience. Thus, variable Customer Experience is formatively determined by independent variables. After establishing the direction, strength, and statistical significance of the relationships in the model, it is necessary to assess the model’s predictive power. The results of the predictive analysis for the retained parameters in the model are presented in
Appendix D,
Table A5. For each indicator, the minimum, mean, and maximum values are shown, along with the standard deviation value. The most important parameter is the predictive relevance value (Q
2 predict). Predictive relevance is defined as the ratio of the sum of squared prediction errors of the PLS path model to the sum of squared prediction errors of the mean benchmark model, expressed as one minus this ratio. A positive value of this parameter provides evidence that the prediction error of the PLS path model is smaller than the error produced by the most naïve benchmark criterion [
185]. To evaluate the strength of predictive power, the predictive relevance values are interpreted such that any positive values above 0.35 may be considered to indicate a high level of predictive relevance of the observed effect [
167]. For the indicators presented in
Appendix D,
Table A5, it can be concluded that they exhibit an exceptionally strong level of predictive relevance for the latent variable they form, since the predictive relevance values range from 0.409 and above.
Based on the research results, it can be concluded that Hypothesis 1, Brand Experience, Service (Provider) Experience, and Post-Purchase/Consumption Experience have a positive influence on Customer Experience, is supported.
4.3. Structural Equation Modeling of the Overall Hierarchical Model
The analysis for the overall hierarchical model was assessed using the same techniques. First, the presence of multicollinearity was assessed using the Variance Inflation Factor (VIF) for all model indicators (see
Appendix B,
Table A4). The analysis proceeded to evaluate the outer loadings of indicators and their significance [
167], as presented in
Appendix E,
Figure A3.
Discriminant validity is demonstrated when each measurement item correlates weakly with constructs other than the one it is theoretically associated with. Discriminant validity was tested by examining the cross-loadings of the indicators [
169,
186] presented in
Appendix F,
Table A7. The highlighted items represent the factor loadings for each construct, while the cross-loadings are reported in the unshaded cells. It can be observed that the cross-loadings for each construct are low, which indicates good discriminant validity. For the Customer Experience construct, the factor scores of the latent variables were considered as indicators, and these loadings correlate more strongly with the construct they define than with the other constructs in the model.
Discriminant validity was also assessed using the Heterotrait–Monotrait (HTMT) ratio of correlations. This measure assesses how well the constructs in the model differ from one another, ensuring that each captures a distinct aspect of the phenomenon under investigation. An HTMT value below 0.90 generally indicates good discriminant validity, meaning that the constructs are sufficiently distinct to be treated as separate [
167,
186]. The results of the HTMT analysis are presented in
Appendix G,
Table A8. The results presented in
Appendix F and G, based on two different analyses, confirm that the criterion of discriminant validity of the model has been satisfied.
After establishing discriminant validity, a metric invariance test was conducted to examine the potential moderating influence of categorical variables. The goal is to statistically confirm that the instrument measures the same underlying concept consistently across groups, ensuring that any observed differences in scores reflect true differences in the construct rather than measurement bias, thereby making comparisons valid.
Configural invariance (Model 1) was tested for gender, with the factor structure free across groups. Model 1 exhibited a good fit (Χ2 (df) = 677.188(258), CFI = 0.959, RMSEA = 0.051), confirming that the same pattern of latent factors existed for both men and women. Next, a metric invariance (Model 2) was tested by constraining the factor loadings to be equal across groups. Compared to the configural model, the metric model (Model 2) did not show a significant decrease in model fit: (Delta CFI = −0.001) with p = 0.124. Therefore, metric invariance was established, indicating that the items contributed to the constructs to a similar extent for both men and women.
For customer segments, Model 1 exhibited good fit (Χ2 (df) = 719.161(258), CFI = 0.956, RMSEA = 0.054), confirming that the same pattern of latent factors existed for both individual and corporate segments of customers. The metric model (Model 2) did not show a significant decrease in model fit: (Delta CFI = 0) with p = 0.364. Therefore, metric invariance was established, indicating that the items contributed to the constructs to a similar extent for both individual and corporate customer segments.
For regions Model 1 showed good fit (Χ2 (df) = 672.333(258), CFI = 0.955, RMSEA = 0.052), confirming that the same pattern of latent factors existed for both Belgrade and Vojvodina region respondents. The metric model (Model 2) did not exhibit a significant decrease in model fit: (Delta CFI = 0) with p = 0.167. Therefore, metric invariance was established, indicating that the items contributed to the constructs to a similar extent for both Belgrade and Vojvodina region respondents.
The results of bootstrapping analysis for the hierarchical model are presented in
Table 3 and
Figure 3.
After the analysis presented in
Table 3, it is confirmed that there is a positive and statistically significant relationship between the variable Customer Experience and the variable Customer Satisfaction as an aspect of the marketing outcomes of customer experience. The path coefficient of this relationship (β = 0.851) indicates a positive relationship, while the corresponding
p-value (
p < 0.001) demonstrates a high level of statistical significance. The interpretation of the results confirms the theoretical assumption that customers who perceive better CX (in terms of brand, service, and post-purchase activities) also demonstrate higher levels of customer satisfaction [
56,
58].
Based on the research results, it can be concluded that Hypothesis 2, Customer Experience has a positive influence on Customer Satisfaction, is supported.
Table 3 presents results that confirm a positive and statistically significant relationship between the variable Customer Experience and the variable Behavior Loyalty Intentions as another aspect of the marketing outcomes of customer experience. The path coefficient (β = 0.766) indicates a positive relationship, while the corresponding
p-value (
p < 0.001) demonstrates a high level of statistical significance. The interpretation of the results supports the theoretical proposition that customers who perceive better customer experience (in terms of brand, service, and post-purchase activities) are more likely to be loyal [
134,
187].
Based on the research results, it can be concluded that Hypothesis 3, Customer Experience has a positive influence on Behavior Loyalty Intentions, is supported.
Finally, the results presented in
Table 3 show a positive and statistically significant relationship between Customer Experience and Word-of-Mouth as a marketing outcome of customer experience. The path coefficient (β = 0.623) is positive, and the corresponding
p-value indicates high statistical significance (
p < 0.001). This supports the theoretical proposition that customers who perceive a better customer experience (in terms of brand, service, and post-purchase activities) are more likely to engage in positive WOM, thereby recommending the service or company to others [
145,
149].
Based on the research results, it can be concluded that Hypothesis 4, Customer Experience has a positive influence on Word-of-Mouth, is supported.
After establishing the direction, strength, and statistical significance of the relationships in the model, it is necessary to evaluate the predictive power of the model. The results of the predictive analysis for the retained parameters in the hierarchical model are presented in
Appendix D,
Table A6. For each indicator, the minimum, mean, and maximum values are reported, along with the standard deviation value. The most important parameter is the predictive relevance value (Q
2 predict), where a positive value provides evidence that the prediction error of the PLS path model is smaller than the error generated by the most naïve benchmark criterion [
185].
To assess the strength of predictive power, the values of predictive relevance are interpreted such that any positive values above 0.35 can be considered to represent a strong degree of predictive significance for the observed effect [
167]. For the indicators analyzed, it can be concluded that they demonstrate an exceptionally high level of predictive significance for the latent variable they form. The lowest predictive power was observed for the indicators of the variable WOM, whose predictive parameters fall within the interval of 0.15 to 0.35, thereby reflecting moderate predictive power (
Appendix D,
Table A6).
Within the scope of this study, the potential moderating influence of variables on the relationship between customer experience and the aspects of marketing outcomes of customer experience was examined [
167,
188,
189].
The results of testing the structural relationships of the model provide insight into the existence of a directional effect of independent variables, such as customer segment, respondents’ gender, or regional affiliation of respondents, on the relationship between customer experience and the aspects of marketing outcomes related to customer experience.
The assessment of the directional influence of independent variables was conducted stepwise, based on potential moderating effects on the observed relationships in the model, and is presented in
Table 3. The results suggest that the existence of a positive directional impact of the variable Customer Segment on the relationship between Customer Experience and Customer Satisfaction is debatable. The path coefficient (β = 0.068) indicates a positive directional effect, which can be interpreted as follows: the relationship between Customer Experience and Customer Satisfaction is strengthened by 0.068 units with each standard deviation of the variable Customer Segment. Thus, the variable Customer Segment may be considered to have a strengthening effect.
For the conclusion to be meaningful, it is also necessary to consider the
p-value (
p = 0.81), indicating exceptionally weak statistical significance of the parameter (
Table 3). Although this may not be sufficient evidence of statistical significance, it does indicate a trend (as
p < 0.1). To gain a better understanding of the potentially moderating effect of the variable Customer Segment, which divides the sample into sub-samples of individual customers and commercial customers,
Figure 4 may be helpful.
Figure 4 illustrates the interaction of variables within the moderating relationship. On the abscissa, the variable Customer Experience is presented using centered values, which also highlight negative values, thereby distinguishing low, medium, and high levels of Customer Experience. On the ordinate, the variable Customer Satisfaction is displayed, likewise using centered values, allowing for a clear identification of low, medium, and high levels of satisfaction.
The lines, presented in two colors, depict the relationship of different segments to the connection between the specified variables, for which a significant directional effect of the two subsamples can be established. Respondents from the commercial segment (represented by the green line) are portrayed with a steeper slope line compared to respondents from the individual customer segment (represented by the red line).
In summary, the conclusion is that a high level of satisfaction may be expected not only as a result of a high level of customer experience, but also as influenced by the customer segment to which they belong. Put differently, when customer experience is at a low level, and respondents are individual customers, higher customer satisfaction is expected, as a noted trend.
Based on the research results, it can be concluded that Hypothesis 5, Customer segment has a moderating effect on the relationship between customer experience and customer satisfaction, is not supported.
The analysis of path coefficients presented in
Table 3 enables us to conclude the moderating effect of the variable Gender on the observed relationships within the model. The value of the path coefficient (β = 0.071) confirms a positive directional impact of the variable Gender on the relationship between Customer Experience and Customer Satisfaction, which can be interpreted as follows: the relationship between Customer Experience and Customer Satisfaction strengthens by 0.071 units with each standard deviation in the variable Gender. Consequently, the variable Gender may be regarded as an enhancing factor.
To ensure the conclusion is meaningful, it is also necessary to consider the
p-value, which indicates the statistical significance of the parameter. For this effect (
Table 3), the relationship is statistically significant at the
p = 0.05 level, which is considered sufficient evidence of statistical significance. To better understand the moderating effect of the variable Gender, which divides the sample into subsamples of female and male respondents, a graphical representation may be helpful, as shown in
Figure 5.
Figure 5 illustrates the interaction of variables within the moderating relationship. The lines depict the relationship between different genders with respect to the specified variables, for which a statistically significant directional effect can be identified in the two subsamples. Respondents of male gender (represented by the green line) are portrayed with a steeper slope line compared to respondents of female gender (represented by the red line).
The conclusion is that a high level of satisfaction can be expected not only from a high level of customer experience, but also because of a particular gender. When customer experience is at a low level, female respondents tend to contribute to higher customer satisfaction, whereas male respondents may report lower customer satisfaction.
Based on the research results, it can be concluded that Hypothesis 6, Gender has a moderating effect on the relationship between customer experience and customer satisfaction, is supported.
Based on the results of the analysis in
Table 3, it is possible to conclude that the variable Region has a moderating effect on the observed relationships in the model. The value of the path coefficient (β = 0.002) barely establishes a positive and extremely weak relationship. The
p-value = 0.933, indicating that the regional affiliation of respondents does not significantly influence the relationship between Customer Experience and Customer Satisfaction. Therefore, there is no need to perform a slope analysis graphically. Although the variable regional affiliation of respondents divides the sample into subsamples for Belgrade, Vojvodina, and Other, no influence could be determined in the specified model.
Based on the research results, regional affiliation does not moderate the relationship between customer experience and customer satisfaction. Thus, Hypothesis 7 is not supported.
Potential moderating effects of the chosen variables were not found to indicate even a trend in the remaining structural relationships between customer experience and behavioral loyalty intentions, nor between customer experience and word-of-mouth.
Following the statistical confirmation of the direction, strength, and statistically significant effects of the variables, the discriminant validity of the final model, which includes all retained relationships, was tested. The results of the HTMT analysis are presented in
Appendix G,
Table A8. All values support the previously drawn conclusion based on the results measuring correlations presented in
Appendix F,
Table A7. The conducted analysis indicates satisfactory discriminant validity, meaning that the constructs are sufficiently distinct and can be regarded as separate [
167].
Empirical evidence shows that gender influences how trust develops in mobile banking, with factors like perceived security, ease of use, and service quality affecting men and women differently. This suggests that males and females rely on different cognitive and emotional processes when building trust in digital financial services [
190]. Customer experience alone doesn’t predict trust outcomes uniformly; instead, gender shapes both the strength and direction of these relationships, indicating its role as a crucial moderator, not just a control variable. These insights are especially relevant for banking systems powered by Industry 4.0 and Industry 5.0, where human–technology interactions and perceived risks are key to shaping user experience.
Additional evidence supports the idea that consumer demographics, such as gender, influence how customer experience affects satisfaction, word-of-mouth intentions, and loyalty [
191]. The effect of customer experience on marketing outcomes differs notably between men and women, especially regarding relationships like loyalty and WOM, indicating that gender not only influences how customers evaluate their experience but also shapes their behavioral intentions and how they communicate after purchase.
Prior empirical evidence in the extant literature, along with the findings reported in this study, emphasize that the inclusion of demographic moderating variables significantly enhances the accuracy and practical usefulness of customer experience (CX) models. Demographic moderators, such as gender, customer segments, or regions, enable researchers to capture systematic differences in how distinct customer groups perceive service encounters and translate experiential evaluations into attitudinal and behavioral outcomes.
From both methodological and managerial perspectives, incorporating demographic moderating effects enhances model validity and relevance. By acknowledging demographic differences, CX models offer more precise diagnostic value and support the development of targeted, human-centric, and context-sensitive customer experience strategies that are consistent with a systems-oriented view of service design and the principles of Industry 5.0.