Next Article in Journal
Algebraic Reduction and Periodic Solvability in a Coupled Ternary Rational System
Previous Article in Journal
An Extended BEM Model for 2-D Elasticity Problems
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Heterogeneous Transfer Learning for Linear Regression Model with Heteroscedasticity

1
School of Statistics and Data Science, Nanjing Audit University, Nanjing 211815, China
2
Joint Laboratory for Statistics and Finance, Nanjing Audit University, Nanjing 211815, China
*
Author to whom correspondence should be addressed.
Mathematics 2026, 14(8), 1395; https://doi.org/10.3390/math14081395
Submission received: 28 February 2026 / Revised: 7 April 2026 / Accepted: 16 April 2026 / Published: 21 April 2026
(This article belongs to the Section D1: Probability and Statistics)

Abstract

We consider the transfer learning problem in the linear regression model, where the source domain and target domain have different features and the model exhibits heteroscedasticity. The existing homogeneous transfer learning methods cannot yet handle this type of problem. In this work, a transfer learning algorithm is proposed, which integrates data of varying dimensions and accounts for heteroscedasticity, thereby yielding a data-pooling estimator. The algorithm is simple to implement and easy to operate. The theoretical properties of the proposed estimator are established, including an upper bound on the generalization error and robustness against negative transfer. Simulation studies indicate that the proposed method performs well in terms of parameter estimation. The effectiveness of the proposed method is also validated using real-life datasets in UCI public repository, demonstrating favorable performance compared to the existing methods.

1. Introduction

Transfer learning is an important research technique in machine learning and also a new learning paradigm. It focuses on using existing knowledge (namely, source domain) to solve different but related tasks (namely, target domain), addressing the problem of having only a small amount of labeled sample data. It also overcomes the tediousness and inefficiency of learning each task from scratch and building different learning models for each problem. In many practical situations, collecting a large amount of labeled data for each target domain is very expensive and difficult. Therefore, it is necessary for transferring knowledge from related tasks to improve problems in the target domain. Moreover, transfer learning has achieved significant results in fields such as natural language processing, computer vision, and healthcare. As evidenced by these successes, transfer learning represents a powerful and meaningful strategy for improving data efficiency and model performance across diverse applications [1].
Recently, several authors have explored the problem of transfer learning in many statistical models. Tian and Feng [2] discussed the transfer learning under high-dimensional generalized linear models; they proposed a transfer learning algorithm in generalized linear models and achieved a theoretical analysis. In the logistic regression framework, Hou and Song [3] developed a two-step transfer learning algorithm and a non-algorithmic transferable source detection method to offer consistent and robust privacy protection. Transfer learning for nonparametric regression is considered in [4], which studies the non-asymptotic minimax risk of transfer learning in nonparametric regression and develop a confidence thresholding estimator that achieves the minimax optimal risk up to a logarithmic factor. Lin and Reimherr [5] studied the transfer learning for the functional linear regression under the reproducing kernel Hilbert space framework. Jin et al. [6] extended transfer learning methodology to quantile regression settings, developing a framework for detecting informative sources and enhancing estimation accuracy in applications such as flight safety analysis. He et al. [7] developed a semiparametric transfer learning framework that leverages shared representations while modeling domain-specific heterogeneity. Transfer learning has also been applied in other models, such as in [8,9,10].
While some of the aforementioned methods offer statistical frameworks in various statistical models, they are limited by a crucial assumption. It is assumed that the target and source populations have identical covariates or features at their disposal, which is also known as homogeneous transfer learning. Nevertheless, this assumption is impractical in many real-world applications. For example, in biomedical studies, the target domain on a specific patient population has a limited amount of data, and only a few key variable features can be collected, while the source data or related electronic health record can be easily obtained, and more features can be measured, which is referred to as heterogeneous transfer learning. At this point, homogeneous transfer learning limits their practical applicability. The challenge of learning from source datasets with diverse feature spaces has been investigated in research, such as [11,12]. Peng and Wang [13] assumed that auxiliary samples were collected from different sub-populations with non-negligible heterogeneity and developed a strategy to leverage possible shared information from relevant source datasets. Karbalayghareh et al. [14] proposed a novel Bayesian method to allow a robust partial information transfer by imposing a spike-and-slab prior on the joint distribution of parameters in the proxy and target domain. Chang et al. [15] introduced a heterogeneous transfer learning for high dimensional regression with different features. Silvestrin et al. [16] developed a transfer learning technique in the linear regression versions with different input dimensions. Recently, considerable progress has been made in heterogeneous transfer learning under high-dimensional settings. Most existing high-dimensional heterogeneous frameworks focus on feature heterogeneity or distribution shift while rarely integrating systematic heteroscedasticity modeling into linear regression transfer inference. In this work, we consider this issue and develop a transfer learning strategy to heteroscedastic linear regression under cross-domain heterogeneity, thereby extending the scope of the existing heterogeneous transfer learning methods to accommodate heteroscedastic error structures.
As is well known, heteroscedasticity is a critical phenomenon that significantly impacts the validity and reliability of regression models. Its presence can lead to inefficient and biased parameter estimates, thereby undermining the accuracy of predictions and inferences. White [17] demonstrates that heteroscedasticity biases standard error estimation and proposes a consistent covariance matrix estimator to restore valid inference. Wooldridge [18] further stresses that, in finite samples and complex data settings, uncorrected heteroscedasticity substantially distorts test statistics and confidence intervals. Berenguer-Rico and Wilms [19] reveal that even preliminary data-cleaning steps, such as outlier removal, can interact with heteroscedasticity to induce size distortion in hypothesis testing. Recognizing and addressing heteroscedasticity is essential for enhancing the robustness of statistical analysis and ensuring that the conclusions drawn from data are sound and trustworthy. Based on the above analysis, it is of great significance to conduct research on heterogeneous transfer learning in linear regression with heteroscedasticity. Informed by prior research, a novel estimator termed heteroscedastic data-pooling (HDP) is developed herein for linear regression with heteroscedasticity, which effectively integrates datasets from source and target domains that contain different features. To address this issue, an efficient and easily implementable heterogeneous transfer learning algorithm is developed, for which an upper bound on the generalization error is established through a theoretical analysis. The robustness of the estimator is confirmed by numerous simulation experiments and real data analysis.
Although the existing heterogeneous regression-based transfer methods (e.g., [13,15,16]) have made progress in handling feature heterogeneity across domains, they generally fail to explicitly model domain-specific heteroscedastic noise or construct dedicated estimators that can adapt to both heterogeneous feature spaces and heteroscedasticity. Different from these prior approaches, the HDP estimator proposed in this paper explicitly characterizes the heteroscedastic structure in both source and target data, and achieves effective integration of cross-domain data with different features, which distinguishes our method from existing heterogeneous regression-based transfer approaches.
The main contributions and innovations of this paper can be summarized as follows:
  • We address the problem of transfer learning when the source and target domains have inconsistent feature dimensions and heteroscedastic noise, which is rarely considered in existing studies.
  • We propose a novel Heteroscedastic Data-fusion Predictor (HDP), which constructs a robust transfer learning objective by weighted fusion of source and target domain data.
  • We derive the estimation error upper bound for the proposed HDP estimator, providing theoretical guarantees for its performance under heteroscedasticity and feature misalignment.
  • Extensive simulations and experiments demonstrate that HDP effectively leverages source domain information while maintaining robustness against feature mismatch and heteroscedastic noise in the target domain.
The differences between the contributions of this paper and those of existing works are summarized as follows. Chang et al. [15] considers a homoscedastic linear regression model and assumes that some features present in the source data are missing in the target domain. In contrast, our work accounts for heteroscedasticity in the model and assumes that some features in the target data are missing in the source domain. Compared with Silvestrin et al. [16], our work explicitly incorporates heteroscedasticity into the modeling framework. Zhao et al. [20] studies heterogeneous transfer learning for high-dimensional generalized linear models. Although it allows the model parameters to differ between the source and target domains, it assumes that the feature spaces of all datasets are identical. By contrast, our work considers the more general setting where the feature spaces of the source and target domains are allowed to be different.
The rest of the paper is organized as follows: Section 2 introduces the model setup and formulates the heterogeneous transfer learning problem under heteroscedastic linear regression. Section 3 presents the heteroscedastic data-pooling estimator. Section 4 provides the theoretical analysis, including the transfer gain and its properties. Section 5 reports both simulation experiments and real-data case studies to evaluate the performance of the proposed method. Finally, Section 6 concludes the paper and discusses future research directions.

2. Model Setup

This article focuses on the problem of learning models based on two datasets: historical data and new data containing additional features. The first one will be referred to as the source dataset, the second as the target dataset. Linear regression is adopted as the modeling framework. The labels in two datasets represent the same linear regression task, but the variance of the error terms in both the source and target domains is heteroscedastic. In this section, we will define the dataset and its distribution in the context of linear regression.
Let the target dataset be defined as ( X T , y T ) , where X T R n T × d T denotes a full-rank matrix consisting of n T independent observations with d T input features, and y T R n T is a random vector of the corresponding labels. These are associated with the heteroscedastic linear model y T = X T β + w T , where β is a d T -dimensional parameter vector capturing the linear relationship between input features and labels, and w T is noise that follows a Gaussian distribution N 0 n T , diag ( σ 1 2 , σ 2 2 , , σ n T 2 ) , where σ i 2 > 0 , for i = 1 , , n T . The objective of this paper is to learn the parameter vector β .
Additionally, define the source dataset as ( X S , y S ) , where X S R n S × d S is a full-rank matrix containing n S observations with d S input features ( d S < d T ). And y S R n S stands for the random vector of corresponding labels. The setting d S < d T naturally arises from the incremental input transfer learning problem that motivates our work. The opposite case d S > d T would essentially become a feature selection or dimensionality reduction problem, which is a different research direction that can be investigated in the future. The relationship between labels and input features in the source dataset is given by y S = X S β + X β + w S , where β and β correspond to the first d S and last d T d S components of β , respectively. Here, X represents a n S × ( d T d S ) random matrix, whose entries X i j are mutually independent, and each follows a standard normal distribution, i.e., X i j i . i . d . N ( 0 , 1 ) . The assumption that unobserved covariates follow a Gaussian distribution is justified on three grounds. First, it greatly simplifies the theoretical analysis, enabling closed-form derivations of the estimators and their associated properties. Second, although real-world data may not be strictly normal, many features become approximately normally distributed as the sample size increases. Third, the real-data study by Silvestrin et al. [16] demonstrates that, even when the normality condition is not met, the proposed method based on this assumption still yields satisfactory performance. It is further assumed that w S follows a Gaussian distribution, N ( 0 n S , diag ( σ ˜ 1 2 , , σ ˜ n S 2 ) ) , where σ ˜ j 2 > 0 , for j = 1 , , n S , and is independent of w T . Evidently, the source data follows a heteroscedastic linear model.
Thus, this paper focuses on a linear regression model with a heteroscedastic structure. Furthermore, in the source domain, only a subset of the features available in the target dataset is observable; that is, the newly added features are unknown, while their impact on the labels of both datasets remains consistent (captured by β ). The unobserved new features are substituted with random values to account for their influence on the labels. Consequently, the distributions of the source and target domain labels y S and y T can be derived as follows:
y S N ( X S β , diag ( σ ˜ 1 2 + β 2 , σ ˜ 2 2 + β 2 , , σ ˜ n S 2 + β 2 ) ) ,
y T N ( X T β , diag ( σ 1 2 , σ 2 2 , , σ n T 2 ) ) ,
where β 2 = k = 1 d T d S ( β k ) 2 , and β k is the k-th element of the vector β . All · used in this paper denote the 2 norm; i.e., for a matrix, R m × n , the 2 norm is defined as = i = 1 m j = 1 n ( i j ) 2 . Proof details are presented in Appendix A.
In this work, the parameters of the heteroscedastic model are estimated using the generalized least squares method (GLS). For illustration, we focus on the estimation in the target domain. Let Σ T = diag ( σ 1 2 , σ 2 2 , , σ n T 2 ) . Based on the distributional assumption of w T above, it follows that Σ T is a positive definite matrix. According to the Cholesky decomposition theorem, there exists a lower triangular matrix, C T , such that Σ T = C T C T T . Multiplying both sides of the model y T = X T β + w T by the matrix C T 1 , we obtain the transformed model y ˜ T = X ˜ T β + ω ˜ T , where y ˜ T = C T 1 y T , X ˜ T = C T 1 X T , and ω ˜ T = C T 1 ω T . Since Var ( ω ˜ T ) = C T 1 Σ T ( C T 1 ) T , substituting Σ T = C T C T T yields Var ( ω ˜ T ) = C T 1 ( C T C T T ) ( C T 1 ) T = I . This transformation converts the original heteroscedastic models into homoscedastic forms, allowing the application of ordinary least squares (OLS) to the transformed model and yielding a generalized least squares estimator that inherently accounts for heteroscedasticity. The resulting estimator is given by:
β ^ T * = ( X T T ( C T C T T ) 1 X T ) 1 X T T ( C T C T T ) 1 y T .
Similarly, for the source domain model, an analogous transformation can be applied. It is evident from the assumption on w S that Σ S = diag ( σ ˜ 1 2 , , σ ˜ n S 2 ) is positive-definite. Specifically, there exists a matrix C S satisfying Σ S = C S C S T , such that, after multiplying both sides of the source model by C S 1 , the transformed model has homoscedastic errors.
At this stage, the source and target datasets, together with their underlying distributions, have been formally specified. Within the heteroscedastic linear regression framework, the GLS estimation procedure for both domains has been introduced. By applying the Cholesky decomposition to the covariance structure of the noise terms, the original heteroscedastic models are transformed into equivalent homoscedastic representations, which enables the direct application of OLS techniques. These results establish the statistical foundation and estimation tools required for the subsequent construction of transfer learning estimators under incremental input settings.
Basic Estimator: To highlight the performance improvement brought by the source dataset, this paper selects a model that uses only the target data as the baseline and chooses the OLS estimator, which is widely used to solve linear regression problems to calculate the loss by minimizing the sum of squared residuals of the target data:
R T ( β ) = y T X T β 2 .
Therefore, the basic estimator is defined as:
β ^ T = X T T X T 1 X T T y T .
The OLS is chosen as the baseline because it is the maximum likelihood estimator using ( X T , y T ) and is widely used to solve linear regression problems. Furthermore, ordinary least squares also possesses two favorable statistical properties: unbiasedness and efficiency.

3. Heteroscedastic Data-Pooling Estimator

Building upon the heteroscedastic linear regression model and the estimation framework developed in the previous section, this section proposes a transfer learning approach for incremental input settings. The method leverages information from both the source and target domains while explicitly accounting for heteroscedasticity and discrepancies in the observed feature spaces.
Heteroscedastic data-pooling loss: To integrate information from the source and target domains under heteroscedastic noise, we construct a unified empirical risk that jointly incorporates domain-specific loss functions. The proposed data-pooling loss combines the source and target risks through a weighted scheme, where each component is adjusted according to its corresponding noise covariance structure.
This formulation provides a principled mechanism for balancing the contributions of the two domains while accounting for discrepancies in noise levels and observed feature spaces. A common transfer learning approach is to learn by minimizing the convex sum of errors from two datasets. In the case of linear regression, this method is known as data-pooling. This paper defines the data-pooling loss R α ( β ) as the weighted sum of the losses from the source and target datasets (denoted as R S and R T , respectively):
R α ( β ) = α S R S ( β ) + α T R T ( β ) .
It is clear that, when α T α S , the error of the target dataset is magnified compared to the error of the source dataset, and the solution obtained by optimizing R T ( β ) will be closer to the baseline estimator. On the other hand, when α T α S , R S ( β ) will dominate the loss, and the optimal solution will deviate from the baseline estimator.
In the incremental input setting, a key consideration is to isolate the influence of parameters corresponding to newly introduced features on the source-domain error. To this end, this paper defines R S ( β ) = y S x S I β 2 , where I represents a d T × d S matrix, such that I i i = 1 and I i j = 0 when i j . The operation X S I effectively performs dimension expansion on the source-domain feature matrix, where absent dimensions are filled with zero values. Through the algebraic substitution of both R S ( β ) and R T ( β ) into the unified optimization framework, we derive the following analytical solution. Proof details are presented in Appendix A.
R T ( β ) = y T X T β 2 = C T ( y ˜ T X ˜ T β ) 2 R S ( β ) = y S X S I T β 2 = C S ( y ˜ S X ˜ S I T β ) 2 R α ( β ) = α S C S ( y ˜ S X ˜ S I T β ) 2 + α T C T ( y ˜ T X ˜ T β ) 2 .
 Proposition 1.
Given the source and target datasets defined in Section 2, the data-pooling loss R α ( β ) is convex for any chosen α s , α t R + . Therefore, β ^ α has a unique minimum solution, which is defined as follows:
β ^ α = A 1 b ,
where A = α S I ( C S 1 X S ) T ( C S 1 X S ) I T + α T ( C T 1 X T ) T ( C T 1 X T ) , b = α S I ( C S 1 X S ) T ( C S 1 y S ) + α T ( C T 1 X T ) T ( C T 1 y T ) .
Algorithm 1 presents the computational procedure for estimating β ^ α . Based on this procedure and with the leveraging of the probability distributions of both y S and y T , the first and second moments of the data-pooling estimator can be derived analytically. The closed-form expressions for its expectation and variance are formally established in Proposition 2.
Algorithm 1 Heteroscedastic data-pooling estimator
Input:     source dataset ( X S , y S ) , target dataset ( X T , y T )
1.             Initialization:  X ¯ T = 0 d T .
2.             For j > d S : X ¯ j T 1 n T i = 1 n T X i j T .
3.             For i = 1 , , n T : X i T X i T X ¯ T .
4.              β ^ T X T ( C T C T ) 1 X T 1 X T ( C T C T ) 1 y T .
5.              β ^ S X S ( C S C S ) 1 X S 1 X S ( C S C S ) 1 y S .
6.              α S ( n S d S ) / C S 1 ( y S X S β ^ S ) 2 .
7.              α T ( n T d T ) / C T 1 ( y T X T β ^ T ) 2 .
8.             Compute β ^ α with Equation (8).
 Proposition 2.
The data-pooling estimator β ^ α is an unbiased estimator. Its expectation and variance are:
E [ β ^ α ] = β , Var ( β ^ α ) = M 1 α S 2 I ( C S 1 X S ) T ( C S 1 X S ) I T + α T 2 ( C T 1 X T ) T ( C T 1 X T ) M 1 ,
where M = α S I ( C S 1 X S ) T ( C S 1 X S ) I T + α T ( C T 1 X T ) T ( C T 1 X T ) . Proof details are presented in Appendix A.
From Proposition 2, it is evident that the data-pooling estimator β ^ α remains unbiased for the true parameter β across all admissible values of the weighting coefficients α S and α T . This property ensures consistent convergence, irrespective of the specific weight assignment. Clearly, Var ( β ^ α ) is a positive semidefinite matrix; i.e., for any l R d T , l T Var ( β ^ α ) l 0 . The variance, however, as explicitly shown in Formula (9), exhibits a direct dependency on the choice of these hyperparameters. Although selecting the optimal α = ( α S , α T ) to minimize Var ( β ^ α ) is of direct interest, the complex functional form of the variance with respect to α makes this a challenging problem.

4. Theoretical Analysis

To assess the effectiveness of the proposed heteroscedastic data-pooling framework, we introduce the concept of transfer gain as a principled evaluation metric. This metric quantifies the benefit of incorporating source-domain data by comparing the generalization performance of the proposed estimator against a baseline that uses only target-domain data.
Formally, consider predicting the label for a new target-domain instance represented by a feature vector, X R d T . In accordance with the heteroscedastic linear regression model, the corresponding label y follows y = X β + w , where w N ( 0 , Ω ) . Assume Ω = C C T ; then, the transformed model is given by y ˜ = X ˜ β + w ˜ , where y ˜ = C 1 y and X ˜ = C 1 X . For any estimator, β ^ , of β , the generalization error of the prediction for the transformed label y ˜ is given by E [ ( y ˜ X ˜ β ^ ) 2 ] .
The transfer gain G ( X ) is then defined as the reduction in this generalization error achieved by using the data-pooling estimator β ^ α compared to the baseline target-only estimator β ^ T :
G ( X ) = E y ˜ X ˜ β ^ T 2 E y ˜ X ˜ β ^ α 2 .
The transfer gain G ( X ) , as defined in Equation (10), serves as a quantitative measure to evaluate whether incorporating source-domain data improves predictive performance on the target task. A positive value of G ( X ) indicates beneficial knowledge transfer, while a negative value suggests a detrimental negative transfer.
Given the complexity of directly computing the expectations in Equation (10), under the assumption of linear regression model, i.e., E [ w T ] = 0 , E [ w S ] = 0 and E [ X ] = 0 , we then derive a simplified expression through the following proposition:
 Proposition 3.
For unbiased estimators β ^ α and β ^ T , the transfer gain can be characterized by the difference between their covariance matrices when evaluated on whitened new instances:
G ( X ) = X ˜ Var ( β ^ T ) Var ( β ^ α ) X ˜ T .
Proof details are presented in Appendix A. This reformulation reveals that the sign of the transfer gain is determined by the difference in estimation precision between the two estimators. By explicitly computing the variances, we obtain the operational expression:
G ( X ) = X ˜ Q T 1 X T T y T M 1 α S 2 I ( C S 1 X S ) T ( C S 1 X S ) I T + α T 2 ( C T 1 X T ) T ( C T 1 X T ) M 1 X ˜ T ,
where Q T = X T T X T .
Equation (12) provides a complete analytical characterization of the transfer gain G ( X ˜ ) through its dependence on the covariance matrices of the estimators. The key insight from this formulation is that both the sign and magnitude of the transfer gain are entirely determined by the difference in estimation precision between β ^ α and β ^ T . In particular, if Var ( β ^ α ) is smaller than Var ( β ^ T ) in the positive semi-definite ordering, then the transfer gain is guaranteed to be non-negative for all non-zero whitened input vectors, X ˜ .
This observation establishes a direct link between variance reduction and performance improvement, and it provides a principled criterion for selecting the transfer weights ( α S , α T ) , so as to enhance generalization performance. As a direct consequence, we can identify a special but practically relevant setting under which a non-negative transfer gain is always ensured, which is formally stated in the following theorem:
 Theorem 1.
Consider any datasets, ( X S , y S ) and ( X T , y T ) , and let t = α S / α T > 0 . Define J = I ( C S 1 X S ) T ( C S 1 X S ) I T , D = ( C T 1 X T ) T ( C T 1 X T ) , and let λ max ( J , D ) denote the largest generalized eigenvalue of the pair ( J , D ) . Then, the transfer gain G ( X ) is non-negative for all X R d T if and only if t 1 λ max ( J , D ) . Proof details are presented in Appendix A.
Theorem 1 characterizes the conditions under which the proposed heteroscedastic data-pooling estimator β ^ α provides a non-negative transfer gain relative to the target-only estimator β ^ T . In particular, if λ max ( J , D ) 1 , any positive weights ( α S , α T ) satisfy this condition; otherwise, t must not exceed 1 / λ max ( J , D ) to avoid directions with negative gain. This result generalizes the previous special case commonly assumed in practice and provides a principled guideline for selecting the relative weighting of source and target information.

5. Simulation Experiment and Case Study

5.1. Simulation Experiment

This section presents a series of simulation experiments designed to evaluate the performance of the proposed Heteroscedastic Data-Pooling (HDP) estimator in linear regression models with heteroscedastic errors and mismatched input dimensions across domains.
Throughout all simulation experiments, data are generated from linear regression models with an intercept term. Specifically, the response variable follows y = 2 + X T β + w , where β = ( 2 , 2 ) T . An intercept is incorporated by augmenting the design matrix with a constant column. Additional data-generating characteristics, including the covariance structure of covariates, noise variance patterns, and sample sizes in the source and target domains, are specified separately in each subsection.
The proposed heteroscedastic data-pooling estimator is compared with three baseline approaches: a data-pooling method that ignores heteroscedasticity (DP), ordinary least squares applied solely to the target domain (OLS), and pooled ordinary least squares applied to the combined source and target datasets (DP_OLS), which treats the aggregated data as homoscedastic. These baselines are selected to disentangle the respective contributions of data integration and heteroscedasticity correction. In particular, DP evaluates the effect of explicitly modeling heteroscedasticity, OLS serves as a target-only benchmark, and DP_OLS assesses whether naive pooling combined with standard least squares can achieve comparable performance.
To ensure reproducibility, the random seed is fixed across all simulations. Estimation accuracy is primarily evaluated using the mean squared error (MSE), with the root mean squared error (RMSE) reported when relevant, defined as MSE = 1 n i = 1 n ( y i y ^ i ) 2 and RMSE = MSE , respectively. All simulations are conducted with fixed, reproducible settings: the random seed is set to 10, and each experiment is repeated 200 times.

5.1.1. The Impact of n S on Estimation Performance

In this subsection, we investigate the effect of the source-domain sample size n S on estimation performance, while all other experimental settings remain fixed as specified at the beginning of this section. Specifically, the target-domain sample size is fixed at n T = 10 , and the source-domain sample size varies over a moderate range given by the closed interval n S [ 220 , 280 ] . The mean squared error (MSE) is used to evaluate the performance of the proposed HDP estimator in comparison with DP, OLS, and DP_OLS. The corresponding results are reported in Figure 1.
Figure 1 shows that the HDP estimator achieves the lowest MSE when the source-domain sample size is moderate, particularly for 220 n S 260 . In this regime, HDP consistently outperforms DP and DP_OLS, while the performance gap gradually diminishes as n S increases. When the source-domain sample size becomes sufficiently large, the performance of HDP converges to that of the two pooled estimators, suggesting that the benefit of carefully weighting source information is partially offset by the sheer volume of source data.
Across all values of n S , the OLS estimator exhibits a substantially higher MSE than the other methods, highlighting its inability to effectively leverage source-domain information. Overall, these results indicate that the proposed HDP estimator is particularly effective in regimes with moderate source sample sizes.

5.1.2. Correlation Experiment

This subsection investigates the impact of feature correlation on estimation performance. In particular, we examine how varying the correlation coefficient between covariates affects the effectiveness of information transfer from the source domain to the target domain.
The source-domain sample size is fixed at n S = 100 , and the target-domain sample size is fixed at n T = 8 . The correlation coefficient c between X 1 and X 2 is varied over the interval [ 0 , 0.9 ] , while all other data-generating settings remain unchanged. For each value of c, the results are averaged over 200 independent Monte Carlo replications. The mean squared error (MSE) is used as the evaluation metric. The corresponding results are reported in Figure 2.
As shown in Figure 2, the performance of all methods is influenced by the correlation structure of the covariates. The proposed HDP estimator consistently achieves lower MSE than DP, OLS, and DP_OLS across the full range of correlation values, demonstrating robustness to changes in feature dependence.
Notably, the MSE curves exhibit a non-monotonic pattern as the correlation coefficient increases. When the correlation is low ( c < 0.3 ), both DP and DP_OLS suffer from relatively large estimation errors, indicating limited transferability between weakly related features. As the correlation increases, their performance improves, yet they remain inferior to HDP.
For the HDP estimator, moderate levels of correlation may introduce additional estimation variability due to the interaction between heterogeneous noise variances and partially aligned feature structures. However, in high-correlation regimes ( c > 0.8 ), the MSE of HDP decreases, suggesting that strong feature similarity enables more effective utilization of source-domain information for variance reduction. Overall, these results highlight the advantage of the HDP estimator in accommodating heterogeneous noise structures across a wide range of correlation settings.

5.1.3. Residual-Based Weighting

This subsection investigates a practical weighting strategy based on residual information, motivated by scenarios in which observation-level noise variances are unknown and the HDP estimator is, therefore, infeasible. To this end, we consider a residual-based weighted data-pooling approach (WDP), in which observation-specific weights are constructed from squared residuals obtained via preliminary ordinary least squares fits on the source and target domains. This procedure yields a heteroscedasticity-aware pooled estimator without requiring explicit knowledge of the underlying noise variances and can be viewed as an implementable approximation of variance-based weighting. Specifically, domain-specific residual variances are estimated via OLS fits on the source and target samples, and the resulting squared residuals are used to construct diagonal weighting matrices in a pooled weighted least squares (WLS) estimator.
The experimental setting follows the previous configuration, where the correlation coefficient c between newly introduced and existing target-domain features serves as the independent variable, and MSE on the test set is used for evaluation. The performance of four estimators, DP, WDP, OLS, and DP_OLS, is summarized in Figure 3. We note that HDP estimator is not included here, as its behavior under varying correlation structures has already been examined in Section 5.1.2. The purpose of this experiment is to evaluate whether residual-based weighting can partially recover the benefits of variance-aware pooling in the absence of oracle variance information.
Overall, both DP and DP_OLS exhibit decreasing MSE as the correlation increases, reflecting the growing effectiveness of naive data pooling when the new and existing features become strongly aligned. In contrast, the OLS estimator remains relatively insensitive to feature correlation, showing limited variation across different settings. The residual-weighted estimator achieves the lowest MSE under low to moderate correlation levels, indicating that adaptive weighting based on residual information can effectively mitigate heteroscedastic effects in this regime.
When the correlation becomes sufficiently large, the performance gap between residual-weighted pooling and unweighted pooling narrows. In this case, the strong linear dependence between features enhances the utility of pooled data, reducing the relative benefit of residual-based reweighting. Nevertheless, across a broad range of correlation values, the residual-weighted approach remains competitive and demonstrates clear advantages in scenarios characterized by weak to moderate feature dependence and heteroscedastic noise.

5.1.4. The Impact of n T on Transfer Gain

Unlike the previous subsections, which focus on estimation accuracy measured by prediction error, this subsection directly investigates the effect of the target-domain sample size on the empirical transfer gain achieved by data pooling. The goal is to quantify when and to what extent incorporating source-domain information leads to performance improvement over target-only estimation.
To this end, the source sample size is fixed at n S = 100 , while the target sample size n T varies from 5 to 30. The data-generating process follows a linear regression model, y = 2 + 2 X 1 2 X 2 + w , where the target-domain model is deliberately misspecified by retaining only X 1 , while the source-domain model remains correctly specified. For each value of n T , the empirical transfer gain is evaluated using a fixed test set of size n test = 1000 , and the results are averaged over 200 independent repetitions.
The transfer gain for the i-th repetition is given by
G i = 1 n test ( X , y ) D test ( y X β ^ T ) 2 ( y X β ^ α ) 2 ,
and the reported empirical transfer gain corresponds to the average of G i across all repetitions. The specific execution process is listed in Algorithm 2.
Algorithm 2 Empirical Transfer Gain Algorithm
Input: Number of iterations L, test dataset D test
1:     for  i = 1 , , N  do
2:             Randomly sample D S and D T
3:             Calculate β ^ a and β ^ T
4:             Calculate empirical transfer gain using Formula (13)
5:     end
Output: Average of all G i
The empirical results in Figure 4 show that the transfer gain remains non-negative across all target sample sizes, which is consistent with the theoretical guarantee in Theorem 1. The gain attains its maximum when the target sample size is small and gradually decays toward zero as n T increases. This pattern reflects the diminishing error reduction effect of data pooling as the target-only estimator becomes more stable with an increasing sample size. Moreover, the transfer gain obtained using the estimated weighting parameter closely follows the oracle-based curve, indicating that the proposed data-driven estimation of α provides an accurate and effective approximation to the ideal variance-based weighting scheme.

5.1.5. The Impact of n S and n T Simultaneously

To examine the combined influence of source and target sample sizes on transfer performance, we implement a grid-based simulation that simultaneously adjusts n S and n T across a range of practical values. The source domain sample size is set to n S { 50 , 100 , 150 , 200 , 250 , 300 } , and the target domain sample size is set to n T { 5 , 10 , 15 , 20 , 25 , 30 } . Heteroscedastic noise is generated with variances uniformly sampled from U ( 0.1 , 2 ) for both domains, and feature mismatch is constructed by reducing the source feature dimension. The definition of transfer gain is consistent with that in Section 5.1.4.
As illustrated in Figure 5, the maximum positive transfer gain reaches 2.135 when n S = 50 and n T = 5 , followed by a gain of 0.719 for n S = 100 and n T = 5 ; all other combinations yield negative transfer gains, with the minimum value of −1.588 observed at n S = 250 and n T = 30 . The HDP estimator delivers a significant positive transfer gain only in the regime where the target domain is extremely data-scarce ( n T = 5 ) and the source domain sample size is small to moderate ( n S 100 ). As n T increases beyond 5, the transfer gain immediately turns negative and remains stable at approximately −1.0 to −1.6 across all n S values; when n S exceeds 100 (even for n T = 5 ), the positive gain diminishes rapidly, with n S = 150 already yielding a negative gain of −0.449. These patterns confirm that the relative scale of n S and n T is the dominant factor for transfer effectiveness: HDP excels in scenarios where target data is highly limited and source data is moderately available, while excessive source data or sufficient target data leads to a negative transfer due to an amplified feature mismatch bias and a reduced reliance on source information.

5.2. Case Study

This section extends the theoretical analysis to real-life datasets, aiming to examine the practical performance of the proposed method under heterogeneous noise and limited target-domain samples. By considering a diverse collection of benchmark regression tasks with varying dimensionality and sample sizes, the following experiments assess whether the empirical performance of different estimators aligns with the theoretical transfer gain and robustness properties established earlier.

5.2.1. Comparison of Four Methods Under Different Dataset Settings

The empirical evaluation is conducted on nine multivariate regression datasets from the UCI repository, with dataset statistics summarized in Table 1. For each dataset, samples are partitioned into source, target, and test sets according to predefined sizes. To simulate an incremental feature setting, a subset of covariates is removed from the source domain and treated as newly introduced features in the target domain. These features are selected based on their Pearson correlation with the response variable, ensuring that the OLS estimator operates on the most informative covariates. Model performance is assessed using RMSE computed on an independent test set.
To investigate the impact of data availability on transfer performance, two distinct sampling regimes are considered, following common practice in the transfer learning literature.
Small data regime: In the small data setting, the source dataset and the test dataset are fixed across runs, while multiple disjoint target datasets of limited size are repeatedly sampled. Each sampled target dataset is used to train the transfer learning estimators, as well as the OLS baseline. This regime is designed to mimic scenarios in which only a small amount of target-domain data is available, such that estimation performance is primarily dominated by variance. By repeatedly resampling the target data, we evaluate the stability and sensitivity of different methods under severe target-data scarcity.
Large data regime: In the large data setting, for each dataset, a fixed proportion (10%) of the samples is randomly drawn as the test set in each run. The target-domain sample size is fixed at a moderate level, as reported in Table 1, while the source-domain sample size is specified accordingly for each dataset. Depending on the dataset, this procedure can result in either relatively large source datasets or moderately sized target datasets. This regime allows us to examine how transfer performance evolves as the overall data availability increases and whether the benefits of data pooling persist in data-rich settings.
All experiments are repeated multiple times with independent random splits. The reported results correspond to the mean RMSE, together with one standard deviation, providing a quantitative summary of both predictive accuracy and performance stability across different datasets and sampling regimes.
Under the small data regime (Table 2), HDP yields lower RMSE than the OLS estimator in most datasets, with the exception of the energy dataset where all methods achieve comparable performance. Compared with DP_OLS, the improvement offered by HDP is particularly evident in datasets such as parkinsons, pumadyn32nm, and skillcraft, where accounting for heterogeneous noise leads to a noticeable reduction in the estimation error. When compared with naive DP, the performance advantage of HDP is dataset-dependent but generally manifests as comparable or slightly lower RMSE.
As the target sample size increases under the large data regime (Table 3), performance differences among the four methods become less pronounced. Nevertheless, HDP continues to outperform or closely match the OLS estimator in several datasets, including parkinsons, pol, pumadyn32nm, and skillcraft. Relative to DP and DP_OLS, HDP remains competitive and often achieves stable performance without exhibiting large degradation.
Overall, the empirical results suggest that the transfer gain, as defined in Section 4, is predominantly non-negative across most datasets and sampling regimes, in the sense that HDP rarely underperforms the OLS baseline in terms of RMSE. These observations are consistent with the theoretical generalization error analysis developed earlier and support the applicability of heteroscedasticity-aware data pooling in regression problems with heterogeneous noise structures.

5.2.2. Comparison of Four Methods on a Single Dataset

To provide a more fine-grained empirical understanding of the proposed transfer learning method, we conduct a detailed comparison of four estimators on individual real-world datasets. Unlike aggregated evaluations across multiple tasks, this subsection focuses on single-dataset analyses in order to clearly illustrate how transfer learning performance varies with the target-domain sample size under heterogeneous noise and partial feature overlap.
We select three representative regression datasets from the UCI repository: 3D Road Network, Airfoil, and Breastcancer. These datasets differ substantially in scale, noise characteristics, and feature–response relationships, making them suitable benchmarks for evaluating robustness and transfer gain in practical settings. Table 4 summarizes the variables used in the source and target domains for each dataset. In all cases, the feature X 1 is available only in the target domain, while X 2 is shared by both the source and target domains, representing a typical transfer learning scenario with incremental feature availability.
The strength of association between covariates and the response variable is reported in Table 5 through Pearson correlation coefficients. Across all three datasets, the target-only feature X 1 exhibits a stronger correlation with the response variable y than the shared feature X 2 . This implies that models trained solely on the source domain cannot exploit the most informative covariate and, therefore, are expected to suffer from higher variance or bias when the target sample size is small. Such a setting is particularly favorable for transfer learning methods that can effectively integrate source information while adapting to newly introduced target-specific features.
For each dataset, the source and test sets are fixed, while target datasets of varying sizes, n T , are sampled. Specifically, n T ranges from the 8th to the 30th data point, and for each value of n T , the experiment is repeated multiple times with independent random sampling. The test set remains unchanged throughout to ensure a fair comparison across different target sample sizes. Model performance is evaluated using MSE computed on the test set. Since the response variable y is measured on a fixed scale within each dataset, the MSE values are directly comparable across methods for the same task, and a lower MSE indicates better predictive accuracy.
The empirical results for the 3D Road Network, Airfoil, and Breastcancer datasets are illustrated in Figure 6, Figure 7, and Figure 8, respectively. For the 3D Road Network dataset, the OLS estimator exhibits substantial fluctuations when n T is small, with significantly higher MSE compared to the transfer learning methods. This behavior reflects the sensitivity of OLS to heteroscedastic noise and limited target-domain samples. In contrast, HDP and DP_OLS maintain relatively low and stable MSE values across different target sample sizes, demonstrating improved robustness.
A similar pattern is observed in the Airfoil dataset. Although the performance of OLS improves as n T increases, it consistently underperforms compared to transfer-based estimators, particularly in the small-sample regime. Both DP and DP_OLS achieve stable MSE values across the entire range of n T , indicating their effectiveness in handling heteroscedasticity and exploiting source-domain information.
For the Breastcancer dataset, the advantage of HDP is most pronounced when n T is small. In this regime, its MSE is substantially lower than that of OLS and DP, confirming its ability to achieve significant variance reduction under strong heteroscedastic effects. As n T increases, the performance gap narrows, yet the proposed estimator remains competitive. DP shows reasonable performance for a small n T but gradually deteriorates as the target sample size grows, eventually approaching the behavior of OLS, whereas DP_OLS maintains stable accuracy throughout.
Overall, these empirical findings are consistent with the theoretical analysis presented in Section 4. When the target-domain sample size is limited, the baseline estimator suffers from high variance, and the transfer gain achieved by incorporating source-domain information is maximized. As n T increases, the marginal benefit of transfer learning diminishes, but robust transfer-based methods continue to outperform or match classical estimators. This confirms that the proposed approach is particularly advantageous in small-sample, heteroscedastic settings commonly encountered in real-world applications.

5.2.3. Sensitivity Analysis on Shared Feature Proportion

To investigate the robustness of different transfer learning methods under feature misalignment, we conducted a sensitivity analysis using the UCI Breastcancer dataset. The dataset was randomly split into a training set and a test set, with 50 samples selected as the target domain and 200 samples as the source domain. Feature misalignment was simulated by reducing the source domain features to a certain proportion of the target domain dimension. We denote the feature misalignment level as γ [ 0 , 1 ] , where γ = 0 indicates full feature alignment, and larger values of γ correspond to more severe feature misalignment. The analysis was conducted at γ = 0 ,   0.2 ,   0.4 ,   0.6 ,   0.8 . Performance is evaluated using RMSE.
The results are shown in Figure 9. OLS exhibited strong sensitivity to feature misalignment, with RMSE increasing noticeably as γ increased, indicating that relying solely on target domain samples leads to rapid performance degradation when the source domain contains fewer features. In contrast, DP maintained relatively stable RMSE values across different levels of γ , demonstrating that distribution correction and weighting effectively mitigate the impact of missing source domain features. DP_OLS showed considerable performance fluctuation, suggesting that the simple combination of source and target domains via OLS is vulnerable to feature misalignment.
HDP demonstrated the best robustness among the four methods, with only a moderate increase in RMSE as γ increased, indicating that its weighted robust strategy effectively preserves predictive accuracy under severe feature misalignment. Notably, under high levels of feature misalignment, HDP consistently outperformed OLS and DP_OLS, highlighting its superiority in handling feature inconsistency between domains. These results suggest that HDP and DP are relatively insensitive to feature misalignment and can effectively leverage source domain information, thereby enhancing the interpretability of the experimental findings and supporting their applicability in practical transfer learning scenarios.

6. Conclusions

In terms of transfer learning, the existing homogeneous transfer learning methods across domains assume that the dimensions of the source and target datasets are the same and often assume homoscedasticity in linear regression. To address both of the above issues simultaneously, in this work, we have proposed a method for heterogeneous transfer learning that allows heteroscedasticity in linear regression and the different dimensions of the source and target datasets. The theoretical property of the proposed heteroscedasticity data-pooling estimator shows an upper bound for its generalization error and its robustness against adverse transfer. Extensive simulation experiments confirmed the effectiveness of the proposed algorithm. The empirical study in the UCI datasets demonstrates that the proposed methodology performs better than the baseline estimator. In summary, the methodology in this paper effectively addresses key challenges in homogeneous transfer learning with homoscedasticity in the linear regression, yielding satisfactory results in heterogeneous transfer learning for linear regression with heteroscedasticity. However, this paper relies on several assumptions, such as d S < d T and the incremental features X i j i . i . d . N ( 0 , 1 ) . When these conditions are not satisfied in practice, exploring more general settings beyond these assumptions constitutes an interesting direction for future research. In addition, we can also delve more deeply into the theoretical foundations, extend this algorithm to other models, consider differences between conditional distributions in source and target datasets, and further explore practical application scenarios.

Author Contributions

Conceptualization, H.H. (Hongxia Hao) and Y.C.; methodology, H.H. (Hongxia Hao); software, H.H. (Hongqian Hu); validation, H.H. (Hongqian Hu); formal analysis, H.H. (Hongqian Hu); investigation, H.H. (Hongqian Hu); resources, Y.C.; data curation, H.H. (Hongqian Hu); writing—original draft, H.H. (Hongqian Hu); writing—review and editing, H.H. (Hongxia Hao) and Y.C.; visualization, H.H. (Hongqian Hu); supervision, H.H. (Hongxia Hao) and Y.C.; project administration, H.H. (Hongqian Hu); funding acquisition, H.H. (Hongxia Hao). All authors have read and agreed to the published version of the manuscript.

Funding

This research was funded by the National Social Science Fund of China, grant number 23CTJ028; the National Natural Science Fund of China, grant number 12371267; the Priority Academic Program Development of Jiangsu Higher Education Institutions (Statistics); the Open Project of the Joint Lab for Statistics and Finance of NAU, grant number 2025JLSF314; and the Postgraduate Research & Practice Innovation Program of Jiangsu Province, grant number KYCX24_2387.

Data Availability Statement

The original data presented in the study are openly available in [the UC Irvine Machine Learning Repository] at [http://archive.ics.uci.edu/ml (accessed on 25 November 2024)].

Conflicts of Interest

The authors declare no conflicts of interest.

Appendix A

 Proof of the distributions (Formulas (1) and (2)).
y T = X T β + w T , w T N 0 n T , diag ( σ 1 2 , , σ n T 2 ) .
Since X T β is deterministic and w T is Gaussian, we have:
E [ y T ] = X T β , Cov ( y T ) = Cov ( w T ) = diag ( σ 1 2 , , σ n T 2 ) .
Thus,
y T N X T β , diag ( σ 1 2 , σ 2 2 , , σ n T 2 ) .
y S = X S β + X β + w S , X i j i . i . d . N ( 0 , 1 ) , w S N 0 n S , diag ( σ ˜ 1 2 , , σ ˜ n S 2 ) .
For the i-th component of X β , we have
( X β ) i = k = 1 d T d S X i k β k .
Since X i k N ( 0 , 1 ) are independent, it follows that
E [ ( X β ) i ] = 0 , Var ( ( X β ) i ) = k = 1 d T d S ( β k ) 2 = β 2 .
Hence,
X β N 0 n S , β 2 I n S .
Since X β and w S are independent Gaussian vectors, we have:
E [ y S ] = X S β , Cov ( y S ) = Cov ( X β ) + Cov ( w S ) = β 2 I n S + diag ( σ ˜ 1 2 , , σ ˜ n S 2 ) = diag ( σ ˜ 1 2 + β 2 , , σ ˜ n S 2 + β 2 ) .
Therefore,
y S N X S β , diag ( σ ˜ 1 2 + β 2 , , σ ˜ n S 2 + β 2 ) .
The distributions in Formulas (1) and (2) are proven. □
 Proof of the gradient (Formula (7)).
R T ( β ) = y T X T β 2 = C T 1 ( y T X T β ) 2 R S ( β ) = y S X S I T β 2 = C S 1 ( y S X S I T β ) 2 R T ( β ) = 2 C T 1 X T T C T 1 ( y T X T β ) R S ( β ) = 2 I C S 1 X S T C S 1 ( y S X S I T β ) R α ( β ) = R T ( β ) + R S ( β ) .
The objective function R α ( β ) is quadratic in β , and its Hessian matrix
2 R α ( β ) = 2 α T ( C T 1 X T ) T ( C T 1 X T ) + α S I ( C S 1 X S ) T ( C S 1 X S ) I T ,
is positive-definite. Therefore, R α ( β ) is strictly convex and admits a unique minimum.
We complete the derivation of the gradient in (7). □
 Proof of Proposition 2. 
Let M = α S I ( C S 1 X S ) T ( C S 1 X S ) I T + α T ( C T 1 X T ) T ( C T 1 X T ) .
E [ β ^ ] = E M 1 α S I ( C S 1 X S ) T ( C S 1 y S ) + α T ( C T 1 X T ) T ( C T 1 y T ) = M 1 α S I ( C S 1 X S ) T E [ C S 1 y S ] + M 1 α T ( C T 1 X T ) T E [ C T 1 y T ] = M 1 α S I ( C S 1 X S ) T ( C S 1 X S I T β ) + M 1 α T ( C T 1 X T ) T ( C T 1 X T β ) = M 1 α S I ( C S 1 X S ) T ( C S 1 X S ) I T + α T ( C T 1 X T ) T ( C T 1 X T ) β = M 1 M β = β ,
Var ( β ^ α ) = Var M 1 α S Π ( C S 1 X S ) T ( C S 1 y S ) + α T ( C T 1 X T ) T ( C T 1 y T ) = M 1 [ α S 2 Π ( C S 1 X S ) T Var ( C S 1 y S ) ( C S 1 X S ) Π T + α T 2 ( C T 1 X T ) T Var ( C T 1 y T ) ( C T 1 X T ) ] M 1 = M 1 α S 2 Π ( C S 1 X S ) T ( C S 1 X S ) Π T + α T 2 ( C T 1 X T ) T ( C T 1 X T ) M 1 .
where β = I T β . The solutions to the above two equations are obtained under the following conditions: E [ w T ] = 0 , E [ w S ] = 0 , E [ X ] = 0 , and the independence between the source noise term w S and the target noise term w T .
Proposition 2 is proven. □
 Proof of Proposition 3. 
G ( X ) = E ( y ˜ X ˜ β ^ T ) 2 E ( y ˜ X ˜ β ^ α ) 2 = Var ( y ˜ X ˜ β ^ T ) + E [ y ˜ X ˜ β ^ T ] 2 Var ( y ˜ X ˜ β ^ α ) + E [ y ˜ X ˜ β ^ α ] 2 = Var ( y ˜ X ˜ β ^ T ) Var ( y ˜ X ˜ β ^ α ) = Var ( X ˜ ( β β ^ T ) ) Var ( X ˜ ( β β ^ α ) ) = X ˜ Var ( β ^ T ) X ˜ T X ˜ Var ( β ^ α ) X ˜ T = X ˜ Var ( β ^ T ) Var ( β ^ α ) X ˜ T .
Proposition 3 is proven. □
 Proof of Theorem 1. 
The proof proceeds by analyzing the difference in the covariance matrices of the target-only estimator, β ^ T , and the transfer learning estimator, β ^ α , defined as
H = Var ( β ^ T ) Var ( β ^ α ) .
This matrix, H, quantifies the efficiency gain (or loss) from transfer. Substituting the expressions for the two variances yields:
H = Var ( β ^ T ) Var ( β ^ α ) = ( X T Σ T 1 X T ) 1 M 1 α S 2 I ( C S 1 X S ) T ( C S 1 X S ) I T + α T 2 ( C T 1 X T ) T ( C T 1 X T ) M 1 .
To simplify the analysis, we introduce the weight ratio t = α S / α T > 0 and define two key intermediate matrices, J and D, which capture the information structures of the source and target domains, respectively:
J = I ( C S 1 X S ) T ( C S 1 X S ) I T , D = ( C T 1 X T ) T ( C T 1 X T ) .
With the use of these definitions, the variance expressions and the composite matrix M in the definition of β ^ α can be rewritten concisely as:
Var ( β ^ T ) = D 1 , Var ( β ^ α ) = M 1 S M 1 ,
where
M = t J + D , S = t 2 J + D .
To bridge the expression of H and its diagonalized form, we first rewrite
H = D 1 / 2 I ( D 1 / 2 M D 1 / 2 ) 1 ( D 1 / 2 S D 1 / 2 ) ( D 1 / 2 M D 1 / 2 ) 1 D 1 / 2 .
By defining J ˜ = D 1 / 2 J D 1 / 2 , we obtain
D 1 / 2 M D 1 / 2 = t J ˜ + I , D 1 / 2 S D 1 / 2 = t 2 J ˜ + I .
Since J ˜ is symmetric positive semi-definite, it admits an eigen-decomposition, J ˜ = Q Λ Q T , where Q is an orthogonal matrix for standard eigen-decomposition. In the next steps, P will denote the generalized eigenvector matrix for the pair ( J , D ) , which ensures consistent notation for generalized eigenvalues, λ i , used in the diagonal representation of H.
The core of the proof involves a simultaneous diagonalization to decouple the interaction between J and D. We introduce a non-singular matrix, P, such that:
P T J P = Λ = diag ( λ 1 , , λ d T ) , P T D P = I ,
where the λ i 0 are the generalized eigenvalues of the pair ( J , D ) , solving ( J λ D ) v = 0 . Substituting this transformation into the expression for H allows us to analyze it in a diagonal basis:
H = P I ( t Λ + I ) 1 ( t 2 Λ + I ) ( t Λ + I ) 1 P T .
Let
H ˜ = I ( t Λ + I ) 1 ( t 2 Λ + I ) ( t Λ + I ) 1
denote the matrix H in this transformed coordinate system. The benefit of this representation is that H ˜ is diagonal, with its i-th diagonal entry given by:
H ˜ i = t λ i ( 1 t λ i ) ( t λ i + 1 ) 2 .
The condition for a non-negative efficiency gain (i.e., for H to be positive semi-definite) is that all diagonal entries, H ˜ i , are non-negative. From the above equation, H ˜ i 0 holds if and only if t λ i 1 . Since this must be true for all i, the most restrictive condition comes from the largest generalized eigenvalue λ max ( J , D ) , leading to the critical inequality:
t 1 λ max ( J , D ) .
The interpretation of this result completes the proof. If λ max ( J , D ) 1 , the condition t λ i 1 is satisfied for any t > 0 , meaning any positive weight combination yields a variance reduction. Conversely, if λ max ( J , D ) > 1 , a positive semi-definite H (and, thus, a guaranteed efficiency gain) is only assured when the weight ratio t does not exceed the reciprocal of this maximum eigenvalue. Exceeding this bound introduces at least one direction in the parameter space where the transfer learning estimator suffers an efficiency loss compared to using the target data alone.
Thus, Theorem 1 is proven. □

References

  1. Pan, S.J.; Yang, Q. A survey on transfer learning. IEEE Trans. Knowl. Data Eng. 2010, 22, 1345–1359. [Google Scholar] [CrossRef] [Scilit]
  2. Tian, Y.; Feng, Y. Transfer learning under high-dimensional generalized linear models. J. Am. Stat. Assoc. 2023, 118, 2684–2697. [Google Scholar] [CrossRef] [PubMed]
  3. Hou, Y.; Song, Y. Transfer learning for logistic regression with differential privacy. Axioms 2024, 13, 517. [Google Scholar] [CrossRef] [Scilit]
  4. Cai, T.T.; Pu, H. Transfer Learning for Nonparametric Regression: Non-Asymptotic Minimax Analysis and Adaptive Procedure. arXiv 2024, arXiv:2401.12272. [Google Scholar] [CrossRef] [Scilit]
  5. Lin, H.; Reimherr, M. On hypothesis transfer learning of functional linear models. In Proceedings of the 41st International Conference on Machine Learning, Vienna, Austria, 21–27 July 2024; Volume 235, pp. 30252–30285. [Google Scholar]
  6. Jin, J.; Yan, J.; Aseltine, R.H.; Chen, K. Transfer learning with large-scale quantile regression. Technometrics 2024, 66, 381–393. [Google Scholar] [CrossRef] [Scilit]
  7. He, B.; Liu, H.; Zhang, X.; Huang, J. Representation Transfer Learning for Semiparametric Regression. arXiv 2024, arXiv:2406.13197. [Google Scholar] [CrossRef] [Scilit]
  8. Li, S.; Cai, T.T.; Li, H. Transfer learning for large-scale Gaussian graphical models with false discovery rate control. J. Am. Stat. Assoc. 2023, 118, 1–15. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  9. Yan, H.; Chen, S.X. Transfer Learning with General Estimating Equations with Diverging Number of Covariates. arXiv 2024, arXiv:2410.04398. [Google Scholar]
  10. Paul, A.; Rottensteiner, F.; Heipke, C. Transfer learning based on logistic regression. Int. Arch. Photogramm. Remote Sens. Spatial Inf. Sci. 2015, 40, 145–152. [Google Scholar] [CrossRef] [Scilit]
  11. Wiens, J.; Guttag, J.; Horvitz, E. A study in transfer learning: Leveraging data from multiple hospitals to enhance hospital-specific predictions. J. Am. Med. Inform. Assoc. 2014, 21, 699–706. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  12. Yoon, J.; Jordon, J.; van der Schaar, M. INVASE: Instance-Wise Variable Selection Using Neural Networks. In Proceedings of the International Conference on Learning Representations (ICLR), New Orleans, LA, USA, 6–9 May 2019. [Google Scholar]
  13. Peng, Y.; Wang, L. Heterogeneity-Aware Transfer Learning for High-Dimensional Linear Regression Models. Comput. Stat. Data Anal. 2025, 206, 108129. [Google Scholar] [CrossRef] [Scilit]
  14. Karbalayghareh, A.; Qian, X.; Dougherty, E.R. Optimal Bayesian transfer learning. IEEE Trans. Signal Process. 2018, 66, 3724–3739. [Google Scholar] [CrossRef] [Scilit]
  15. Chang, J.H.; Russo, M.; Paul, S. Heterogeneous Transfer Learning for High-Dimensional Regression with Feature Mismatch. arXiv 2024, arXiv:2412.18081. [Google Scholar]
  16. Silvestrin, L.P.; van Zanten, H.; Hoogendoorn, M.; Koole, G. Transfer learning across datasets with different input dimensions: An algorithm and analysis for the linear regression case. J. Comput. Math. Data Sci. 2023, 9, 100086. [Google Scholar] [CrossRef] [Scilit]
  17. White, H. A heteroskedasticity-consistent covariance matrix estimator and a direct test for heteroskedasticity. Econometrica 1980, 48, 817–838. [Google Scholar] [CrossRef] [Scilit]
  18. Wooldridge, J.M. Econometric Analysis of Cross Section and Panel Data, 2nd ed.; MIT Press: Cambridge, MA, USA, 2010. [Google Scholar]
  19. Berenguer-Rico, V.; Wilms, I. Heteroscedasticity testing after outlier removal. Econom. Rev. 2021, 40, 51–85. [Google Scholar] [CrossRef] [Scilit]
  20. Zhao, R.; Kundu, P.; Saha, A.; Chatterjee, N. Heterogeneous Transfer Learning for Building High-Dimensional Generalized Linear Models with Disparate Datasets. arXiv 2024, arXiv:2312.12786. [Google Scholar]
Figure 1. Comparison of MSE values among four methods under different n S .
Figure 1. Comparison of MSE values among four methods under different n S .
Mathematics 14 01395 g001
Figure 2. MSE values among four methods under different correlations. The shaded regions represent the confidence interval of the mean squared error (MSE) over 200 independent Monte Carlo replications, reflecting the statistical variability of the results.
Figure 2. MSE values among four methods under different correlations. The shaded regions represent the confidence interval of the mean squared error (MSE) over 200 independent Monte Carlo replications, reflecting the statistical variability of the results.
Mathematics 14 01395 g002
Figure 3. The impact of residuals-based weighting on the MSE values of four methods. The shaded regions represent the confidence interval of the mean squared error (MSE) over 200 independent Monte Carlo replications, reflecting the statistical variability of the results.
Figure 3. The impact of residuals-based weighting on the MSE values of four methods. The shaded regions represent the confidence interval of the mean squared error (MSE) over 200 independent Monte Carlo replications, reflecting the statistical variability of the results.
Mathematics 14 01395 g003
Figure 4. Transfer gain comparison. The shaded regions represent the confidence interval of the transfer gain across 200 independent Monte Carlo replications.
Figure 4. Transfer gain comparison. The shaded regions represent the confidence interval of the transfer gain across 200 independent Monte Carlo replications.
Mathematics 14 01395 g004
Figure 5. Transfer gain heatmap under joint variations of n S and n T .
Figure 5. Transfer gain heatmap under joint variations of n S and n T .
Mathematics 14 01395 g005
Figure 6. Comparison of MSE values among four methods in the 3D Road dataset.
Figure 6. Comparison of MSE values among four methods in the 3D Road dataset.
Mathematics 14 01395 g006
Figure 7. Comparison of MSE values among four methods in the Airfoil dataset.
Figure 7. Comparison of MSE values among four methods in the Airfoil dataset.
Mathematics 14 01395 g007
Figure 8. Comparison of MSE values of four methods under the Breastcancer dataset.
Figure 8. Comparison of MSE values of four methods under the Breastcancer dataset.
Mathematics 14 01395 g008
Figure 9. RMSE sensitivity of four methods under different levels of shared features between source and target domains.
Figure 9. RMSE sensitivity of four methods under different levels of shared features between source and target domains.
Mathematics 14 01395 g009
Table 1. Dataset statistics.
Table 1. Dataset statistics.
Dataset n total n S n T n test d T #Runs
concrete103011438103821
energy7681143876815
kin40k40,000114384000850
parkinsons5875150505872050
pol15,0001685615002650
protein45,730117394573950
pumadyn32nm8192186628193250
skillcraft3338147493331950
sml4137156524132250
Table 2. RMSE of four methods under a small data regime.
Table 2. RMSE of four methods under a small data regime.
DatasetOLSDPDP_OLSHDP
concrete14.7 ± 2.514.4 ± 3.0214.7 ± 2.514.5 ± 2.4
energy3.4 ± 0.3953.63 ± 0.363.4 ± 0.3953.62 ± 0.177
kin40k1.14 ± 0.08461.06 ± 0.03751.14 ± 0.08461.09 ± 0.0656
parkinsons18.3 ± 5.9510.8 ± 0.53118.3 ± 5.9511 ± 0.44
protein0.851 ± 0.1410.713 ± 0.03530.851 ± 0.1410.733 ± 0.0487
pumadyn32nm1.47 ± 0.1711.09 ± 0.03021.47 ± 0.1711.12 ± 0.0489
skillcraft0.344 ± 0.05010.258 ± 0.00910.344 ± 0.05010.265 ± 0.0124
sml3.4 ± 0.7522.72 ± 0.4823.4 ± 0.7522.84 ± 0.62
Table 3. RMSE of four methods under large data regime.
Table 3. RMSE of four methods under large data regime.
DatasetOLSDPDP_OLSHDP
concrete11 ± 111.1 ± 0.5311 ± 111.3 ± 0.54
energy2.95 ± 0.3583.12 ± 0.352.95 ± 0.3583.16 ± 0.34
parkinsons11.1 ± 1.169.31 ± 0.22511.1 ± 1.169.33 ± 0.235
pol54.2 ± 20.330.7 ± 0.44954.2 ± 20.330.8 ± 0.527
pumadyn32nm1.21 ± 0.07321 ± 0.03011.21 ± 0.07321.01 ± 0.0298
skillcraft0.317 ± 0.1360.267 ± 0.03970.317 ± 0.1360.265 ± 0.0299
sml2.87 ± 0.9192.2 ± 0.2532.87 ± 0.9192.2 ± 0.286
Table 4. Variables of source domain data and target domain data in each dataset.
Table 4. Variables of source domain data and target domain data in each dataset.
Dataset X 1 X 2 y
3D Road NetworkOSM_IDLongitudeAltitude
AirfoilFrequencyChord-lengthScaled sound pressure
BreastcancerTimeFractal_dimension3Outcome
Table 5. Pearson correlation coefficients.
Table 5. Pearson correlation coefficients.
Dataset cor ( X 1 , y ) cor ( X 2 , y ) n test n S #Runs
3D Road Network0.1060.04243,000100115
Airfoil−0.390−0.23615010070
Breastcancer−0.3460.28720100117
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Hao, H.; Hu, H.; Cheng, Y. Heterogeneous Transfer Learning for Linear Regression Model with Heteroscedasticity. Mathematics 2026, 14, 1395. https://doi.org/10.3390/math14081395

AMA Style

Hao H, Hu H, Cheng Y. Heterogeneous Transfer Learning for Linear Regression Model with Heteroscedasticity. Mathematics. 2026; 14(8):1395. https://doi.org/10.3390/math14081395

Chicago/Turabian Style

Hao, Hongxia, Hongqian Hu, and Yao Cheng. 2026. "Heterogeneous Transfer Learning for Linear Regression Model with Heteroscedasticity" Mathematics 14, no. 8: 1395. https://doi.org/10.3390/math14081395

APA Style

Hao, H., Hu, H., & Cheng, Y. (2026). Heterogeneous Transfer Learning for Linear Regression Model with Heteroscedasticity. Mathematics, 14(8), 1395. https://doi.org/10.3390/math14081395

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop