1. Introduction
Clouds and precipitation systems are central components of the atmospheric energy and water cycles, and their representation remains a major source of uncertainty in weather and climate prediction [
1]. Among cloud microphysical variables, liquid water content (LWC) and ice water content (IWC) are particularly important because they describe not only the amount of condensed water, but also its phase and vertical distribution. These vertical profiles are closely linked to cloud radiative effects, precipitation development, and cloud lifetime [
2]. Because liquid droplets and ice particles interact differently with shortwave and longwave radiation, uncertainty in their vertical partitioning can affect estimates of cloud radiative effects and the atmospheric energy balance [
3]. In mixed-phase clouds, the relative amount and spatial distribution of liquid water and ice further influence cloud lifetime, precipitation processes, climate feedbacks, and radiative properties [
4,
5,
6]. Reliable profiling of LWC and IWC is therefore essential for understanding cloud processes and evaluating their representation in models.
Mixed-phase clouds are especially difficult to observe and model. They contain supercooled liquid droplets, ice particles, and water vapour under subfreezing conditions, and their evolution involves ice nucleation, vapour deposition, aggregation, riming, and other coupled microphysical pathways [
7]. Their phase partitioning is also shaped by thermodynamic conditions, aerosol abundance, and dynamical processes, which can modulate ice formation in stratiform cloud layers [
8]. These uncertainties have broader climate implications. A realistic representation of the supercooled liquid fraction has been linked to constraints on equilibrium climate sensitivity [
9], and recent observational constraints on supercooled cloud feedbacks have been shown to reduce uncertainty in projected warming [
10]. When mixed-phase processes are not well represented, models can develop biases in cloud state and thermodynamic structure, including Arctic wintertime temperature inversions [
11]. These issues highlight the need for observational methods that can resolve both liquid and ice water profiles in the vertical.
Accurate LWC and IWC profiles are commonly retrieved by combining observations that describe cloud vertical structure with information on the surrounding thermodynamic environment. Cloud radar, lidar, and microwave radiometer measurements provide complementary information for this task. Cloud radar reflectivity is sensitive to particle size and can penetrate optically thicker cloud layers, while lidar is sensitive to small droplets, aerosols, and cloud boundaries [
12]. Cloud phase information is often a prerequisite for retrieving cloud properties, because phase uncertainty can propagate into LWC and particle-size retrievals [
13]. Ground-based multisensor phase classification has therefore become an important basis for cloud-property retrieval [
14]. Cloudnet provides a broader framework by integrating radar, lidar, microwave radiometer, and model data into continuous time–height products of cloud classification and hydrometeor properties [
15]. Cloud-property retrievals in the Atmospheric Radiation Measurement (ARM) programme have also demonstrated the value of combining active remote sensing, microwave radiometer, and thermodynamic information for cloud microphysical studies [
16,
17]. The Cloudnet processing chain has since been implemented in CloudnetPy, supporting standardized processing of cloud remote-sensing datasets [
18]. Recent Cloudnet-based campaign products further show that such frameworks can provide LWC, IWC, cloud classification, and cloud macro- and microphysical properties at high temporal and vertical resolution [
19].
Despite this observational foundation, converting multisensor measurements into reliable LWC and IWC profiles remains challenging. Physically based synergistic retrievals, such as variational radar–lidar–radiometer methods, have provided an important framework for retrieving ice-cloud properties from combined observations [
20]. More recent EarthCARE retrieval developments extend this idea toward unified active–passive retrievals of clouds, aerosols, and precipitation [
21]. However, mixed-phase retrieval remains more difficult than single-phase retrieval because liquid, ice, and mixed-phase regions produce different observational responses and cannot always be processed using one set of assumptions [
12]. Lidar–radar synergistic retrievals are also limited when lidar signals become fully attenuated, preventing full-column characterization of cloud liquid in optically thick or multilayer mixed-phase clouds [
22,
23]. Cloud radar Doppler spectra can provide information beyond the lidar attenuation height, but liquid and ice signatures may be difficult to separate when turbulence broadens the spectra or when the liquid signal is weak [
24,
25]. Rapid changes in phase partitioning near mixed-phase transition regions add further vertical and temporal variability [
26]. Updates to DARDAR-CLOUD, a combined radar–lidar cloud classification and ice-cloud property product, also show that retrieved microphysical properties can be sensitive to parameter choices and assumed particle relationships [
27]. These limitations suggest that cloud-water retrieval depends not only on instrument availability, but also on how multisensor information is collocated, corrected, and represented vertically.
Machine-learning methods offer a complementary way to extract liquid- and phase-related information from complex remote-sensing observations. Interest in machine learning and deep learning has increased rapidly in Earth system science, partly because these methods can learn nonlinear relationships from large observational datasets [
28]. In cloud-radar applications, artificial neural networks have been used to infer riming from Doppler radar variables such as reflectivity, spectrum width, and skewness [
29]. Supervised learning has also been used for radar Doppler spectra peak finding, and structure-preserving spectral analysis methods have been developed to describe the morphology of Doppler spectra [
30,
31]. More directly related to liquid detection, deep convolutional neural networks have been used to infer the probability of supercooled cloud droplets from vertically pointing Doppler cloud radar observations, using Cloudnet target classification as supervision [
32]. Related work has evaluated Artificial Neural Network-based liquid classification against Cloudnet target classification and independent measurements such as microwave-radiometer liquid water path, ceilometer cloud-base height, radiosonde humidity, and satellite lidar observations [
22]. Recent multisensor machine-learning studies have also shown that random forests, multilayer perceptrons, and convolutional neural networks can use radar, lidar, microwave radiometer, and radiosonde inputs for thermodynamic cloud-phase classification [
33]. Earlier neural-network applications in atmospheric remote sensing include retrievals of precipitable water vapour and liquid water path from ground-based radiometer observations, cirrus cloud retrieval from satellite thermal observations, and analysis of marine liquid-water cloud occurrence from global observations [
34,
35,
36].
These studies show that data-driven methods can extract useful information from radar spectra, multisensor cloud observations, and reference products such as Cloudnet. However, most existing machine-learning applications in this area focus on phase classification, liquid-layer detection, cloud occurrence, or column-integrated quantities. They do not directly address the joint retrieval of vertically coherent LWC and IWC profiles. This distinction is important because cloud water content is not simply a collection of independent point values. LWC and IWC evolve with height through thermodynamic stratification, hydrometeor growth, cloud-boundary transitions, and phase changes. A pointwise regression framework may therefore lose part of the vertical coherence required for physically consistent retrieval, especially near cloud boundaries and mixed-phase transition regions.
This study addresses this gap by reformulating LWC and IWC retrieval as a vertically structured learning problem. Instead of treating each height level as an isolated sample, the proposed approach treats the cloud profile as the basic unit of representation. Year-round Cloudnet observations from the Lindenberg site in 2025 are used as the primary dataset, with Cloudnet-derived LWC and IWC profiles serving as reference targets. Reflectivity from the METEK MIRA-35 Ka-band Doppler cloud radar (METEK Meteorologische Messtechnik GmbH, Elmshorn, Germany), a 35 GHz cloud radar system, was used as the main vertically resolved observational predictor, while microwave-radiometer-derived temperature and relative humidity profiles describe the thermodynamic environment. Radiosonde observations are first used to evaluate and reduce height-dependent thermodynamic biases, improving the consistency of the predictor space before model training. The proposed framework then combines vertical-structure-enhanced features, profile-aware stacking of tree-based learners, and profile-level refinement based on cloud-geometry information. Through this design, the study evaluates whether explicit representation of vertical structure can improve both pointwise accuracy and profile-level coherence in data-driven LWC and IWC retrieval from collocated ground-based radar and thermodynamic observations.
3. Results
The performance of the proposed framework for retrieving cloud liquid water content (LWC) and ice water content (IWC) is evaluated using year-round observations from the Cloudnet Lindenberg site in 2025. The target variables are obtained from the Cloudnet LWC and IWC profile products, while the input features include METEK MIRA-35 radar reflectivity (Ze) and bias-corrected thermodynamic profiles of temperature and relative humidity derived from the Cloudnet microwave radiometer.
After applying unified time–height matching and quality control procedures, multisource paired datasets are constructed. The resulting dataset contains 457,940 valid samples for IWC and 244,950 samples for LWC. The valid atmospheric profiles were divided into training and testing subsets at a ratio of 4:1 using the profile index as the grouping unit, ensuring that all height levels from the same profile were assigned to the same subset.
Model performance is assessed using root mean square error (RMSE), mean absolute error (MAE), correlation coefficient (Corr), and coefficient of determination (R2). The following sections present comparisons across different models, examine height-dependent retrieval behavior, and analyze the contributions of individual variables and model components through ablation experiments.
3.1. Overall Performance Evaluation on the Year-Round Dataset
The predictive performance of different machine-learning models is summarized in
Table 3 (IWC) and
Table 4 (LWC).
For the IWC retrieval task, the multiple linear regression (MLR) model performs poorly (R2 = −0.105), indicating that a linear mapping cannot represent the nonlinear relationships between thermodynamic conditions, radar reflectivity, and ice microphysical variability. Ice microphysics involves nonlinear processes such as the exponential dependence of saturation vapor pressure on temperature and the power-law relationship between radar reflectivity and hydrometeor mass.
The multilayer perceptron (MLP) increases R2 to 0.060, but the improvement remains limited because it treats each height level independently and ignores vertical structural dependence. The LSTM model achieves R2 = 0.079, but the gain remains modest because the retrieval problem is formulated as pointwise regression rather than sequential prediction, preventing the recurrent architecture from exploiting sequence-learning capability.
Tree-based ensemble models perform substantially better. The Random Forest baseline increases R2 to 0.412 and reduces RMSE to 0.0152, indicating that nonlinear feature partitioning effectively captures interactions among thermodynamic variables and radar observations.
The proposed framework further improves performance, reducing RMSE to 0.0092 and increasing R2 to 0.784. These gains demonstrate that encoding vertical structure through VSE, together with profile-level refinement via PAS and VCR, enables the model to represent vertically coherent cloud processes rather than treating each level independently.
For LWC retrieval, overall performance is lower than for IWC. The Random Forest baseline achieves R2 = 0.303, whereas the proposed method increases R2 to 0.606 and reduces RMSE from 0.0786 to 0.0591 g m−3.
This difference reflects the physical nature of liquid clouds. LWC is strongly influenced by boundary-layer turbulence, aerosol activation, and rapid phase transitions, which introduce substantial small-scale variability. By incorporating vertical-structure features, the proposed framework distinguishes different cloud regimes, enabling more accurate retrieval of liquid-water variability.
While the above analysis focuses on overall performance aggregated across all height levels, cloud microphysical properties and retrieval difficulty exhibit strong vertical heterogeneity. To further characterize retrieval behavior, the following section examines how model performance varies with altitude.
3.2. Height-Dependent Performance Evaluation
The height-dependent performance metrics for IWC retrieval are summarized in
Table 5. In the 0–3 km layer, the model exhibits small absolute errors but explains limited variance due to the sparse occurrence of ice water content and mixed-phase processes.
Retrieval accuracy improves substantially in the 3–6 km and 6–9 km layers, where upper-tropospheric ice clouds exhibit more stable particle-size distributions, producing clearer relationships between reflectivity, thermodynamic structure, and IWC.
The vertical evolution of these metrics is illustrated in
Figure 6, where error metrics decrease with height while correlation metrics increase, indicating that the framework performs best in regimes with coherent vertical cloud structure.
For LWC retrieval, altitude dependence is weaker than for IWC. The best performance occurs in the 0–2 km layer, where the model achieves the highest coefficient of determination and the lowest error levels. Accuracy decreases in the 2–4 km layer despite the largest number of samples, likely due to frequent mixed-phase processes that introduce stronger microphysical variability and reduce the stability of the statistical relationship between radar observations and liquid water content.
Table 6 summarizes the height-dependent performance metrics for LWC retrieval. The results indicate moderate variability across altitude ranges.
Performance improves slightly in the 4–6 km layer
, suggesting more organized cloud structures above the mixed-phase transition region. The vertical profiles shown in
Figure 7 confirm that error metrics remain relatively stable with altitude, while correlation metrics vary moderately, indicating robust model performance across different atmospheric layers.
3.3. Ablation Experiment
To examine the contribution of individual variables and model components, two sets of ablation experiments were conducted: input-variable exclusion analysis and stepwise module ablation analysis.
- (1)
Input-variable exclusion analysis
The results for IWC retrieval are summarized in
Table 7. Compared with the full predictor configuration, removing temperature causes the largest degradation in model performance, with RMSE increasing from 0.0092 to 0.0182 and R
2 decreasing from 0.784 to 0.201. This result indicates that temperature provides a dominant thermodynamic constraint for IWC retrieval. Physically, this is consistent with the strong dependence of ice-phase processes on temperature, including saturation vapor pressure, ice nucleation conditions, and depositional growth. Removing radar reflectivity also leads to substantial performance degradation, with RMSE increasing to 0.0192 and R
2 decreasing to 0.244. This suggests that radar reflectivity provides important observational information on hydrometeor scattering and cloud vertical structure.
Excluding relative humidity or height-related information also reduces retrieval performance, although the impact is smaller than that caused by removing temperature or radar reflectivity. The weak performance of the height-only and radar-only configurations further indicates that IWC cannot be reliably retrieved from a single category of predictor. The height-only configuration mainly reflects the climatological vertical background of IWC occurrence, whereas the radar-only configuration captures only radar-observed echo structure without thermodynamic constraints. Therefore, accurate IWC retrieval requires the joint use of radar-observed cloud structure, thermodynamic state, and vertical-context information.
For LWC retrieval, the results summarized in
Table 8 show a different sensitivity pattern. Removing radar reflectivity causes the largest degradation among the variable-exclusion experiments, with RMSE increasing from 0.0591 to 0.0883 and R
2 decreasing from 0.606 to 0.209. This indicates that radar reflectivity provides important scattering-related and structural information on liquid-cloud vertical organization. However, radar reflectivity should not be interpreted as a direct measurement of LWC. Radar reflectivity is strongly weighted toward larger hydrometeors and depends on higher-order moments of the particle size distribution, whereas LWC is more closely related to hydrometeor mass. Therefore, radar reflectivity alone cannot uniquely determine LWC without thermodynamic and vertical-context information.
Removing temperature, relative humidity, or height-related information also degrades LWC retrieval performance, confirming that liquid-water retrieval depends on multiple sources of information. The similar performance of the height-only and radar-only configurations does not imply that radar reflectivity provides little useful information. Instead, it indicates that each single-source predictor captures only a limited aspect of the retrieval problem. The height-only configuration mainly represents the climatological vertical distribution of LWC, such as the tendency for liquid water to occur more frequently in lower cloud layers. In contrast, the radar-only configuration provides information on cloud echo structure but lacks temperature, humidity, and vertical-position constraints. As a result, both configurations show limited skill and cannot adequately represent the actual LWC variability of a specific profile.
The input-variable ablation results indicate that the retrieval skill of the full framework arises from the complementary use of radar reflectivity, thermodynamic predictors, and height-related information. For IWC retrieval, temperature provides a particularly important thermodynamic constraint, whereas radar reflectivity supplies observational information on hydrometeor scattering and cloud vertical structure. For LWC retrieval, removing radar reflectivity causes the largest degradation among the variable-exclusion experiments, but radar reflectivity remains insufficient when used alone. These results support the use of a multisource and vertically structured predictor space rather than relying on any single predictor group.
- (2)
Stepwise module ablation analysis
Module contributions were further evaluated through stepwise ablation experiments, in which the proposed framework was progressively enhanced from the baseline Random Forest model by adding VSE, PAS, and finally PGF + VCR.
For LWC retrieval, the results are summarized in
Table 9. The baseline Random Forest model achieves R
2 = 0.303. Introducing VSE increases R
2 to 0.494, demonstrating that explicit representation of vertical gradients and local profile statistics provides the largest single performance gain. This result indicates that a major limitation of the baseline model lies in its inability to represent vertical structural context when each height level is treated independently.
Adding PAS further increases R2 to 0.507, suggesting that the profile-aware stacking strategy mainly improves generalization by reducing variance through ensemble integration. Although the improvement relative to VSE is smaller, PAS provides a more stable prediction framework across heterogeneous cloud regimes.
The final framework, including PGF and VCR, increases R2 to 0.606 and reduces RMSE to 0.0591, indicating that profile-level structural priors and refinement constraints improve vertical consistency beyond pointwise accuracy alone. These modules, therefore, contribute primarily to the vertical structural coherence of the predicted LWC profiles rather than only to the local regression fit.
The scatter-density distributions in
Figure 8 visually confirm these improvements. Compared with the baseline model, the final framework exhibits a tighter concentration around the 1:1 reference line, reduced spread of residuals, and noticeably less prediction dispersion, indicating both improved accuracy and improved stability.
For IWC retrieval, the module-ablation results summarized in
Table 10 show a similar but even clearer pattern. The baseline model achieves R
2 = 0.412. Adding VSE increases R
2 to 0.688, again indicating that explicit encoding of vertical structural information provides the largest contribution to performance improvement. This large gain suggests that IWC retrieval benefits strongly from local vertical gradients and profile context, which are closely linked to the layered structure of ice clouds.
Incorporating PAS further improves R2 to 0.744, showing that ensemble diversity and profile-aware meta-learning improve model robustness. The complete framework, including PGF and VCR, achieves the best performance with R2 = 0.784 and RMSE = 0.0092. This final improvement indicates that the PGF variables and VCR step help suppress nonphysical oscillations and enhance profile-scale coherence in the retrieved IWC fields.
The corresponding scatter-density distributions in
Figure 9 show a noticeably tighter alignment with the 1:1 reference line, confirming improved agreement between predicted and observed IWC values. Relative to the baseline model, the final framework reduces both spread and bias in the high-density prediction region, indicating that the added modules improve not only average error statistics but also the structural fidelity of the retrieved profiles.
The ablation experiments demonstrate that the largest performance gain comes from explicit representation of vertical structure through VSE, while PAS improves generalization through ensemble diversity, and PGF + VCR further enhances profile-level structural consistency. Together, these results provide direct empirical support for the hierarchical design of the proposed retrieval framework.
3.4. Boundary-Region Evaluation
To further examine whether the proposed framework improves retrieval performance in structurally complex cloud regions, an additional boundary-region evaluation was conducted. Cloud-boundary samples were defined using radar-derived cloud-base and cloud-top heights. Samples located within three radar range gates, approximately 90 m, from either cloud base or cloud top were classified as near-boundary samples, while the remaining in-cloud samples were classified as cloud-interior samples. The baseline random forest model and the proposed framework were then evaluated separately in these two regions for both IWC and LWC retrieval. The results are summarized in
Table 11.
For IWC retrieval, the proposed framework shows a substantial improvement over the baseline RF model in the near-boundary region. The baseline RF model performs poorly near cloud boundaries, with an RMSE of 0.0090 g m−3, a negative R2 of −1.716, and a low correlation coefficient of 0.131. In contrast, the proposed method reduces RMSE to 0.0038 g m−3 and increases R2 and Corr. to 0.524 and 0.729, respectively. This result indicates that the proposed framework is more effective in representing IWC variability near cloud boundaries, where cloud structure and hydrometeor content may change rapidly with height.
In the cloud-interior region, the proposed method also improves IWC retrieval performance, reducing RMSE from 0.0181 to 0.0111 g m−3 and increasing R2 from 0.253 to 0.719. Although improvements are observed in both regions, the near-boundary result is particularly important because it demonstrates that the proposed framework does not only improve vertically averaged statistics, but also enhances retrieval performance in boundary regions that are more challenging for pointwise models.
For LWC retrieval, the proposed method also outperforms the baseline RF model in both near-boundary and cloud-interior regions. Near cloud boundaries, RMSE decreases from 0.0929 to 0.0642 g m−3, while R2 increases from 0.205 to 0.619 and Corr. increases from 0.475 to 0.789. In the cloud-interior region, RMSE decreases from 0.0797 to 0.0590 g m−3, and R2 increases from 0.223 to 0.574. These results suggest that the proposed framework improves the retrieval of liquid water content not only in relatively stable cloud interiors, but also near cloud boundaries where liquid water occurrence and vertical structure can be more variable.
3.5. Cross-Site Evaluation
Additional experiments were conducted using observations from two independent Cloudnet stations, Munich and Jülich, to assess model performance under different atmospheric conditions. In this setup, the model was trained exclusively on the Lindenberg dataset and then directly applied to the Munich and Jülich datasets without retraining or parameter adjustment. The cross-site evaluation results are summarized in
Table 12.
For IWC retrieval, the model achieves an RMSE of 0.0573 g m−3, an R2 of 0.680, and a correlation coefficient of 0.824 for the Munich dataset. When applied to the Jülich dataset, the RMSE increases to 0.0880 g m−3, while the R2 decreases to 0.444 and the correlation coefficient to 0.670.
For LWC retrieval, the Munich dataset yields an RMSE of 0.3882 g m−3, an R2 of 0.376, and a correlation coefficient of 0.640. In contrast, the Jülich dataset shows improved performance, with an RMSE of 0.2420 g m−3, an R2 of 0.547, and a correlation coefficient of 0.748. These results indicate that cross-site performance degradation is reflected not only in increased prediction errors but also in reduced explained variance and weakened statistical consistency between predicted and reference values. The cross-site results demonstrate that the proposed framework exhibits limited generalization capability across independent observational environments, with the degradation being more pronounced for LWC retrieval. The substantial increase in RMSE and the reduction in correlation, particularly for the Munich dataset, suggest that the learned feature–target relationships are sensitive to site-specific atmospheric conditions. This implies that a model trained on a single site cannot fully capture the variability associated with different local cloud regimes and thermodynamic environments. This limitation is likely associated with distribution shifts in both thermodynamic conditions and radar–microphysics relationships across sites. This highlights the necessity of incorporating multi-site training data or domain adaptation strategies to improve model transferability across heterogeneous atmospheric conditions.
4. Discussion
The improvement achieved by the proposed framework is mainly associated with its explicit representation of vertical structure. In conventional pointwise models, each height level is treated as an independent sample, and vertical dependencies are only implicitly represented through shared predictors. This assumption is limited for cloud-water retrieval because cloud microphysical properties often vary coherently along the vertical dimension and may change rapidly near cloud boundaries. By introducing gradient-based descriptors and local contextual features, the proposed framework provides the model with additional information on vertical transitions that cannot be fully represented by single-level predictors alone. This explains why the inclusion of VSE produces the largest performance gain in the stepwise ablation experiments.
The importance of vertical structure is further supported by the boundary-region evaluation. Near radar-derived cloud boundaries, the proposed framework reduced IWC RMSE from 0.0090 to 0.0038 g m−3 and increased R2 from −1.716 to 0.524. For LWC, the near-boundary RMSE decreased from 0.0929 to 0.0642 g m−3, while R2 increased from 0.205 to 0.619. These results indicate that the proposed vertical-context and cloud-geometry representations are particularly beneficial in regions where hydrometeor content and radar reflectivity may vary sharply over short vertical distances. The strong degradation of the baseline RF model near cloud boundaries also suggests that neglecting vertical context is an important source of error in pointwise retrieval, rather than merely a consequence of insufficient model complexity.
The different sensitivities of IWC and LWC retrievals reflect the distinct physical controls governing ice- and liquid-phase clouds. IWC retrieval is strongly influenced by temperature because ice nucleation, depositional growth, saturation conditions, and particle habit evolution are closely linked to thermodynamic state. This is consistent with the input-variable ablation results, where removing temperature produces the largest degradation in IWC performance.
In contrast, LWC retrieval shows greater sensitivity to the inclusion of radar reflectivity, indicating that radar observations provide useful information on liquid-cloud echo structure. However, radar reflectivity should not be interpreted as a direct measurement of LWC. Reflectivity is strongly weighted toward larger hydrometeors and depends on higher-order moments of the particle size distribution, whereas LWC is more closely related to hydrometeor mass. The relatively stronger usefulness of radar reflectivity for LWC retrieval in the present experiments is therefore better interpreted from the perspective of scattering uncertainty: liquid droplets are more nearly spherical and have a stronger dielectric response at 35 GHz, whereas ice-particle retrievals are additionally affected by particle habit, density, orientation, and temperature-dependent microphysical processes. Therefore, the role of radar reflectivity in LWC retrieval should be understood as providing structural and scattering-related information that becomes useful only when combined with thermodynamic and vertical-context predictors.
The cross-site evaluation highlights the limited transferability of the current single-site training strategy. When the model trained at Lindenberg was applied directly to Munich and Jülich, retrieval performance degraded, especially for LWC. This suggests that the statistical relationships learned at one site are not fully invariant across different observational environments. Such degradation is likely related to site-specific differences in boundary-layer structure, aerosol conditions, shallow-cloud evolution, cloud-boundary characteristics, and mixed-phase occurrence, all of which can modify the relationship between radar reflectivity, thermodynamic profiles, and Cloudnet-derived hydrometeor products. The relatively better transferability observed for IWC may reflect the stronger dependence of ice-phase processes on larger-scale thermodynamic structure, whereas LWC is more strongly affected by local boundary-layer and shallow-cloud variability. These results indicate that multi-site training, site-aware calibration, or domain adaptation may be necessary for robust application across heterogeneous cloud regimes.
Several limitations should also be noted. First, the framework remains fully data-driven and does not explicitly enforce physical constraints such as mass conservation, phase consistency, or microphysical closure. Although the vertical refinement step reduces small-scale oscillations, it cannot guarantee physically consistent profiles under all cloud conditions, particularly in mixed-phase clouds where microphysical processes are highly nonlinear. Second, the boundary-region evaluation is based on radar-derived cloud-base and cloud-top definitions. It therefore provides direct evidence for improved retrieval near cloud boundaries, but it should not be interpreted as a complete validation of all phase-transition processes. Third, Cloudnet LWC and IWC products are used as reference targets, but they are themselves retrieval-based estimates and contain uncertainties related to instrument sensitivity, retrieval assumptions, and microphysical parameterisations. The reported metrics therefore quantify agreement with Cloudnet reference products rather than absolute physical accuracy. Future work should incorporate stronger physical constraints into the learning process, extend the training dataset to multiple sites and cloud regimes, and investigate domain-adaptation strategies to improve generalization under changing observational conditions.
5. Conclusions
This study formulated cloud liquid water content (LWC) and ice water content (IWC) retrieval as a vertically structured learning problem by integrating Cloudnet hydrometeor products with cloud radar reflectivity and microwave-radiometer-derived thermodynamic profiles. The proposed framework explicitly incorporated vertical-structure-enhanced features, profile-aware stacking, cloud-geometry information, and vertical consistency refinement to improve both pointwise retrieval accuracy and profile-level coherence.
Using year-round observations from the Lindenberg Cloudnet site, the proposed framework consistently outperformed conventional pointwise and general-purpose machine-learning models. For IWC retrieval, RMSE decreased from 0.0152 to 0.0092 g m−3, while R2 increased from 0.412 to 0.784. For LWC retrieval, RMSE decreased from 0.0786 to 0.0591 g m−3, while R2 increased from 0.303 to 0.606. Ablation experiments further showed that vertical-structure-enhanced features made the largest contribution to the overall improvement.
Boundary-region evaluation further confirmed that the proposed framework improves retrieval performance in structurally complex parts of cloud profiles, particularly near radar-derived cloud boundaries where hydrometeor content can change rapidly with height.
The two retrieval tasks showed different sensitivities to input variables. IWC retrieval was strongly constrained by thermodynamic conditions, particularly temperature, reflecting the role of temperature-dependent ice-phase processes. In contrast, LWC retrieval showed stronger sensitivity to the inclusion of radar reflectivity, which provides important information on liquid-cloud echo structure. However, radar reflectivity alone was insufficient for reliable LWC retrieval, highlighting the need to combine radar-observed cloud structure with thermodynamic and vertical-context information.
Cross-site evaluation revealed that the learned relationships were sensitive to site-specific atmospheric conditions. Models trained at Lindenberg showed degraded performance when directly applied to Munich and Jülich, especially for LWC retrieval. This indicates that single-site training is insufficient to fully represent variations in boundary-layer structure, shallow-cloud processes, and local cloud regimes across different observational environments. Future applications should therefore consider multi-site training, site-aware calibration, or domain-adaptation strategies to improve robustness and transferability.
These findings suggest that cloud-water profile retrieval should not be treated only as an independent pointwise regression problem, but also as a vertically organized profile-retrieval task. By incorporating vertical context, cloud-geometry information, and profile-level refinement, the proposed framework provides a practical data-driven approach for improving LWC and IWC profile retrieval from ground-based remote-sensing observations. This perspective may also provide a useful basis for other profile-based atmospheric retrieval tasks, such as thermodynamic profiling, aerosol retrieval, and data-assimilation-oriented cloud analysis.