Next Article in Journal
Adaptive Beamforming Based on Flamingo Search Algorithm with Early-Stop Strategy
Previous Article in Journal
Measurement of Metal Surface Temperature Based on Visible Light Images: A Strategy for On-Site Image Acquisition
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Cloud Point Temperature of Thermoresponsive Systems: A Predictive Approach in Data Scarcity Conditions

by
Marcela Elisabeth Penoff
1,2,
Facundo Ignacio Altuna
1,2 and
Luis Alejandro Miccio
1,*
1
Institute of Materials Science and Technology (INTEMA), National Research Council (CONICET), Colón 10850, Mar del Plata 7600, Argentina
2
Facultad de Ingeniería, National University of Mar del Plata (UNMdP), Juan B. Justo 4302, Mar del Plata 7600, Argentina
*
Author to whom correspondence should be addressed.
Appl. Sci. 2026, 16(5), 2557; https://doi.org/10.3390/app16052557
Submission received: 3 February 2026 / Revised: 3 March 2026 / Accepted: 4 March 2026 / Published: 6 March 2026

Abstract

In this study, we employ machine learning techniques to improve materials in data scarcity conditions. In particular, we focus on the prediction of the cloud point temperatures of polymer–water systems with thermoresponsive behavior. We compare a model trained directly on the available data with a model based on representations learned through an encoder–decoder model, in turn pre-trained on a larger dataset to generate molecular fingerprints. Our results demonstrate that the embedding-based model significantly outperforms the direct model in predicting the cloud point temperature under the data limitations imposed by rigorous curation. This approach highlights the potential of domain-informed representation learning to tackle complex materials science problems with limited data.

1. Introduction

Thermoresponsive polymers represent an increasingly important class of “smart” materials, capable of experiencing abrupt and reversible physical transitions in response to temperature variations [1,2]. These polymers, particularly those displaying lower critical solution temperature (LCST) and upper critical solution temperature (UCST) behaviors, have attracted substantial interest in diverse fields, ranging from drug delivery and biomedical engineering to sensors, coatings, and smart textiles [3,4,5,6,7,8,9,10,11,12]. The trademark of these polymers is their ability to reversibly transition between soluble and insoluble states in response to temperature changes, driven by a thermodynamic phase separation process, in turn related to complex intermolecular and intramolecular interactions including hydrogen bonding, hydrophobic interactions, and changes in their chain topology [13,14,15,16,17]. In particular, LCST polymers are soluble below a certain critical temperature and become insoluble when heated beyond this point, while UCST polymers are soluble above a specific temperature and phase-separate below it [18]. This thermoresponsive characteristic has significant potential in the design of stimuli-responsive drug delivery systems, in which the polymer solution can circulate freely at physiological temperature but undergo phase separation to release therapeutic agents at “target” sites [19]. Furthermore, such temperature-induced changes in polymer solubility and conformation can also be exploited in smart coatings and bioseparations, thus providing enhanced control over material properties and interactions. Despite considerable research, poly(N-isopropylacrylamide) (PNIPAM) has remained the benchmark polymer for thermoresponsive applications due to its critical temperature being close to physiological conditions [20]. However, PNIPAM has several limitations, including concerns about biocompatibility, phase transition hysteresis, and pronounced sensitivity to end-group [21]. As a result, research efforts have diversified towards alternative thermoresponsive polymers with finely tunable response profiles. For example, polymers such as poly(N-vinylcaprolactam) offer enhanced biocompatibility, while poly(oligoethylene glycol methacrylates) are characterized by minimal hysteresis and greater control over molecular end-groups. Additionally, polyphosphazenes and polyphosphoesters have gained attention due to their structural versatility and exceptional biodegradability [19].
An important area of investigation within thermoresponsive polymer materials involves understanding how chemical structure and environmental conditions influence LCST and UCST transitions. As in many other polymer properties, structural attributes such as hydrophilic–hydrophobic balance, polymer topology, side-chain chemistry, and molecular weight [22,23,24] critically determine the temperature responsiveness of polymer systems [19,25]. On the other hand, purely environmental factors (like solvent quality, salt concentration, and pH, among others) further modulate these transitions. For instance, increasing salt concentrations can dramatically alter cloud points by affecting hydration shells around polymer chains, thereby modulating polymer–solvent interactions [26]. As a result, determining critical solution temperatures experimentally is not trivial due to the number of variables to control and also due to the complexity (and variability) in the measurement methods. Some of the most popular experimental approaches include turbidimetry, differential scanning calorimetry (DSC), and dynamic light scattering (DLS) [11,27,28]. However, significant variations arise due to the subjective definitions of cloud points and differences in turbidity thresholds, heating rates, and polymer concentrations [29,30]. These methodological challenges have been reported, and substantial discrepancies resulting from different operational definitions of cloud points, ranging from the onset of turbidity (approximately 10% opacity) to the inflection point on the turbidity curve (approximately 50% opacity), were noticed [26].
Given these experimental challenges, predictive modeling has become a critical area of advancement, offering computational tools to estimate critical solution temperatures from molecular structures and descriptors. Quantitative structure–property relationships (QSPR), artificial neural networks, and polymer-specific fingerprinting approaches have been effectively employed to streamline polymer design processes, thereby reducing time and resource consumption [19,31,32,33]. For example, some research has focused on developing linear topological and quantum-chemical-based QSPR models, achieving high prediction accuracy [19]. These results evidence the effectiveness of quantum chemical norm descriptors, which capture subtle molecular interactions crucial for predicting thermoresponsive behaviors. Moreover, machine learning techniques, employing molecular fingerprints such as the Morgan fingerprints [34,35,36,37], have significantly enhanced predictive capacities. These approaches efficiently represent polymer structures for machine learning by capturing compositional and topological data derived from routine polymer characterization techniques and can achieve high prediction accuracies (and therefore show the potential of ML-driven approaches for rapidly screening novel polymer candidates). However, predictive models must navigate a critical trade-off between dataset size and experimental data quality [38]. Larger datasets can enhance model generalizability but often can also compromise detailed molecular characterization, potentially limiting prediction accuracy for specialized polymer–solvent systems. Conversely, highly curated datasets help to increase prediction accuracy but at the expense of reducing applicability to broader contexts. Thus, meticulous descriptor selection, rigorous data curation, and experimental validation remain essential to developing robust, generalizable models.
In this study, we address the challenge of predicting cloud point temperatures for aqueous homopolymer solutions under conditions of severe data scarcity. We compare two modeling strategies: (a) a direct model trained solely on a curated set of SMILES (simplified Molecular Input Line Entry System) strings and number-average molecular weights (Mn) and (b) a transfer learning approach based on embeddings generated by an encoder–decoder model pre-trained on a larger dataset of organic compounds (logP data). The latter approach exploits learned representations of chemical structure to enhance predictive capacity. By evaluating and contrasting both models, we aim to demonstrate the substantial advantages of embedding-based transfer learning over conventional direct prediction in low-data regimes and to highlight the potential of such techniques for virtual screening of novel thermoresponsive polymers without the need for costly and time-consuming experimental efforts.

2. Methods

In this section we describe the data curation, the model architectures, and the other methodologies employed to develop a neural network capable of predicting cloud point temperatures in dilute water–homopolymer systems. In particular, we outline our direct prediction approach and a data-scarcity-tolerant architecture, which uses the knowledge of a previously trained encoder–decoder network (originally optimized on logP data [38]) to generate chemical structure embeddings (fingerprints). We then detail the integration of these fingerprints into a model capable of predicting cloud point temperatures, along with our evaluation strategy. The general procedure is illustrated in Figure 1.

2.1. Data Collection and Preprocessing

We employ a raw dataset of measured cloud point temperatures for various polymer solutions, accompanied by polymer-specific parameters (such as molar mass, end-group types, and polymer concentrations) and solution conditions (pH, salt composition, among others). We perform a general curation process by using an automated pipeline that removes empty or inconsistent data and keeps entries with the relevant variables (such as SMILES structure of the repeating units and number-average molecular weight of the polymer). This selection also ensures that the models will be trained on the chemical structure of the materials, so that the approach remains flexible enough to be extended to more-complex systems in the future. In order to ensure a consistent and high-quality dataset, a subsequent specific curation is executed. During this specific curation stage, the dataset was refined to include only homopolymers, and entries were filtered by grouping together all cloud point temperature measurements corresponding to a fixed SMILES string and Mn. From these groupings, only the data corresponding to the most diluted solutions (that is, the lowest concentration recorded) were retained (i.e., values ranging from 0.05 to 0.00025 wt%). This decision ensures that the recorded cloud point temperatures reflect the polymer’s intrinsic thermoresponsive behavior at maximum dilution, with minimal interference from any concentration effects. While these filtering criteria result in a highly reliable and internally consistent dataset, they also impose a severe data scarcity condition due to the exclusion of multiple measurements and non-homopolymeric entries.
As usual, the resulting dataset was randomly partitioned into training, validation, and test sets, with the training data used to optimize the model parameters, the validation set employed for hyperparameter tuning and to trigger early stopping, and the test set reserved for external performance evaluation. All molecular structures were represented as SMILES strings, a standard text-based format for encoding molecular information [39,40]. During preprocessing, the SMILES strings were tokenized at the character level, therefore allowing for sequence processing by the models. The obtained sequences were in turn padded to a uniform length (equal to the maximum length in the dataset).

2.2. Direct Model

We defined a reference by training a model directly on the dataset. In this direct approach, several neural networks were trained to predict cloud point temperatures from the curated polymer data (homopolymer SMILES with their associated Mn).
Our network architecture comprises a feed-forward design with an input layer corresponding to the size of the tokenized SMILES vector. This is followed by an embedding layer with several hidden convolutional layers and a fully connected layer (which enable the network to learn hierarchical representations while gradually compressing the feature space). Each hidden layer employs non-linear activation functions (LeakyReLUs) to capture complex interactions among the input features. The training was performed using the Adam optimization algorithm to minimize the mean squared error (MSE) between the predicted and experimental cloud point temperatures. Given the limited size of the dataset produced by our strict curation strategy, a key aspect of the training protocol is the use of early stopping based on the validation set performance, which together with the use of a dropout strategy helps prevent overfitting. Finally, hyperparameter optimization via grid search was used to fine-tune model parameters.
As mentioned, by training a model directly on the curated dataset, it acts as a reference benchmark and provides some insight into the intrinsic relationship between polymer structure and thermoresponsive behavior. Figure 2 shows a schematic picture of the direct model architecture.

2.3. Data Embedding (Encoder–Decoder Model)

To obtain low-dimensional (and yet chemically meaningful) representations of the polymers, we employed an encoder–decoder architecture. In this way, this model learns to understand chemical structures, therefore allowing any subsequent model to just focus on the property at hand, thus resulting in a much better performance than that of the direct model. The encoder–decoder was trained on a large logP dataset, with the aim of capturing important aspects of the chemical structure, molecular solubility, intermolecular interactions and hydrophobicity. As shown in the schematic picture in Figure 3, the architecture is composed by two Long Short-Term Memory (LSTM) networks [41]: (a) the encoder processes the tokenized SMILES sequences and converts them into a latent vector that captures the chemical structure features (it is composed by an LSTM layer with 256 units); (b) the decoder reconstructs the original SMILES sequence from this latent representation (also an LSTM layer with 256 units that takes the latent vector as input); and (c) intermediate dense layers with ReLU activation functions [42] were used to refine the latent representation before feeding it into the decoder.
In order to optimize the difference between the input and reconstructed SMILES sequences, during the training process we employed sparse categorical cross-entropy as the loss function [43]. We trained the model on 14,500 molecules (validation on 800), and we optimized it using the Adam optimizer [44] (learning rate of 0.001 and a batch size of 1000). As before, we employed early stopping on validation loss to prevent overfitting. The validation loss closely followed the training loss curve, therefore suggesting a good generalization to previously unseen data and reaching a minimum validation loss of approximately 0.5 after 208 epochs.

2.4. Model Based on Embeddings

We developed a model for predicting cloud point temperatures from the embeddings generated from the encoder described in the previous subsection. In this step, the tokenized SMILES, the fixed-size embedding vectors and the corresponding Mn are fed to a deep neural network. Optimization is performed using the Adam optimizer, and training is terminated using early stopping based on validation loss stagnation. Data normalization and scaling steps were applied to the input features to ensure consistency and improve training efficiency. Finally, the hyperparameters were optimized via grid search. Figure 4 shows a schematic picture of the architecture.

2.5. Model Evaluation

The models’ performance was evaluated during training on the validation set, and later on an independent test set, using several standard metrics: root mean squared error (RMSE), mean relative error (MRE), mean average error (MAE), mean average percentage error (MAPE), and coefficient of determination (R2). In addition, visual diagnostic tools including hexagonal binning cross-plots of experimental versus predicted cloud point temperatures were used to inspect the distribution of errors and to validate the model’s generalization capabilities. We then compared our results with our reference direct model, observing that the integration of the embeddings with the deep neural network yielded robust cloud point predictions, even under conditions of very limited training data.

3. Results and Discussion

In this section, we compare the effectiveness of our models in the task of predicting cloud point temperatures of diluted water–homopolymer solutions. Our aim is to understand how our embedded representations perform under data-limited conditions and how they compare with traditional direct approaches. We start by describing the direct model approach, and then we focus on the more advanced embedding-based model results.

3.1. Direct Model

A key challenge in developing a robust direct model for predicting critical solution temperatures is the pronounced scarcity of reliable data. Even though several LCST and UCST cloud point data appear in the literature, the diversity of conditions (like pH, molecular weight, salt content, among others) is immense, and therefore systematically measured, high-quality data remain limited. This poses a substantial obstacle to these direct machine-learning pipelines, which require large, balanced datasets to effectively learn the underlying structure–property relationships. After curating and preprocessing the dataset of polymer–water pairs, nearly 150 samples remained, from which only 115 were employed in the training process: 80% training (92 samples)/20% validation (23 samples). A typical training trajectory is shown in Figure S4.
Moreover, all the trained direct models exhibited only moderate-to-low performance. While certain architectures and hyperparameter configurations seemed promising, their results consistently fell short of delivering stable and accurate cloud point temperature predictions across the full range of samples. Error metrics such as MAE and MAPE (calculated directly in °C scale to avoid masking deviations) remained relatively high (see Table S1). Qualitatively, the models showed limited extrapolation ability beyond the training region, often producing large prediction errors for the polymer chemistries and molecular weight inputs within the data. These findings reinforce the hypothesis that data deficiency, rather than model design alone, is the prime limiting factor in direct modeling efforts (resources are going into learning the chemical structure and the correlation with the property, and the data is simply not enough for that).
Considering the reduced dataset’s complexity (i.e., since the dataset includes only homopolymers, it basically removes any variability due to copolymer structure, crossed interactions, among others) the model performance is low. Focusing on these simpler systems allows a more direct correlation between polymer structure and the observed cloud point temperature behavior. However, the small number of experimentally measured samples for homopolymers and the limited variety of chemical structures still lead to insufficient coverage of feature space, and consequently, poor predictive performance. These results highlight that even under simplified conditions, a purely direct prediction approach struggles without a robust, large-scale dataset and comprehensive coverage of polymer chemistries.

3.2. Embedding-Based Model

Since it basically allows building upon a learned representation of polymer structures, embedding-based models offer a substantial advantage for predicting the polymer properties. In this study, the encoder-based fingerprint strategy circumvents the need for the predictive model to extract all molecular features directly from raw descriptors or simplified molecular input lines (SMILES). Instead, the encoder is initially trained on a broader set of polymer structures, generating latent-space representations (or “embeddings”) that capture the main chemical features. This latent representation effectively reduces the dimensionality of the problem, distilling key structural motifs, molecular weights, and other relevant chemical characteristics into a more manageable space. Figure 5 shows the results of hyperparameter optimization during grid training, where most of the trained models widely outperform their equivalent direct model counterparts.
Compared with direct modeling, where the network must simultaneously learn both the features and their relationship to the cloud point temperatures, the encoder-based approach provides a form of inductive bias that guides the subsequent prediction model toward more-relevant features. These compressed, chemically meaningful fingerprints allow the downstream network to focus on the property of interest without expending resources on “understanding” chemical structures, complex intra- and intermolecular interactions, among other variables. This separation of representation learning from property modeling is particularly valuable in the context of data scarcity: the encoder can be trained on a much broader set of data, using structural information beyond the relatively small subset for which high-quality cloud point temperature measurements are available.
As shown in Figure 6, our evaluations reveal that the encoder-based model significantly outperforms the direct model. On the one hand, training trajectories converge faster and to lower losses (MAE), while, on the other hand, other metrics such as RMSE, MAPE, or R2 consistently improve. These results reflect the more robust and stable predictions derived from these learned fingerprints. Figure 6a shows the Predicted vs. True values obtained from the best model during training (model 7 in Table S2). These results are also reflected in the histogram of deviations shown in Figure 6b. Despite the dispersion in the experimental values, the deviations are well below 10 °C. The same results are observed in the validation set, in Figure 6c and Figure 6d, respectively. Finally, the same trend can be observed in the independent test set (see Predicted vs. True values in Figure 6e).
As shown by the different chemical structures in the dataset (see Figure 7 for the test set examples), the model exhibits an enhanced ability to handle variations in polymer chemistry and molecular weight, with fewer outlier predictions even in these underrepresented regions of structure space. Another noticeable advantage arises from how the encoder captures global and local structural features in a single representation. By enforcing a shared latent space during training, the encoder effectively “clusters” similar polymer motifs, allowing the predictive model to recognize analogies among polymer structures that may appear quite distinct based on simple descriptors (e.g., repeating units, functional groups, or side-chain lengths). This kind of chemical cognition is especially valuable in predicting cloud point temperatures, a property that depends on complex, multifactorial interactions among polymer structure, solvent polarity, and hydrogen-bonding capabilities. When compared with the results of the direct model (with MAPE values of no less than 30%), the superiority of the encoder-based approach (MAPE values of about 14%) highlights a broader principle in data-driven materials science: domain-informed representation learning can often mitigate the inherent limitations of small or noisy datasets. The direct model suffers from insufficient training examples to capture the breadth of structure–property relationships, leading to comparatively erratic and less accurate predictions.
These findings underscore the power of representation learning in the chemical and materials domains, particularly for complex, data-scarce problems like predicting phase behavior in polymer–solvent systems.

4. Conclusions

This work presents a comparative analysis of the effectiveness of direct and embedding-based models for predicting the cloud point temperatures of dilute homopolymer solutions in water. We observed that the scarcity of high-quality data significantly limits the performance of direct models, which struggle to learn the complex structure–property relationships. In contrast, embedding-based models, which use previously learned knowledge about polymer structure representations, offer a substantial advantage. Our results demonstrate that the models based on an encoder pre-trained to generate fingerprints consistently outperform the direct models in predicting the cloud point temperature. This is attributed to the encoder’s ability to capture both global and local structural features, allowing the predictive model to recognize analogies between polymer structures. The superiority of the embedding-based approach highlights the importance of domain-informed representation learning to mitigate the inherent limitations of small or noisy datasets in materials science. In conclusion, this study highlights the power of representation learning in the chemical and materials domains, particularly for complex, data-scarce problems such as predicting phase behavior in polymer–solvent systems.

Supplementary Materials

The following supporting information can be downloaded at https://www.mdpi.com/article/10.3390/app16052557/s1; Additional information on the dataset curation, the training of the direct models, and the training of the embedding-based models.

Author Contributions

Conceptualization, M.E.P., F.I.A. and L.A.M.; Methodology, L.A.M.; Software, L.A.M.; Validation, M.E.P.; Formal analysis, M.E.P., F.I.A. and L.A.M.; Investigation, M.E.P., F.I.A. and L.A.M.; Resources, L.A.M.; Data curation, L.A.M.; Writing—original draft, L.A.M.; Writing—review & editing, M.E.P., F.I.A. and L.A.M.; Supervision, L.A.M.; Funding acquisition, L.A.M. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

The data that supports the findings of this study are available within the article and in the Supporting Information File (SI). All code is available at https://github.com/lamiccio/LCST_manuscript.git.

Acknowledgments

We gratefully acknowledge the support of NVIDIA Corporation with the donation of the GPU used for this research and the financial support from CONICET and the UNMdP.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Kotsuchibashi, Y. Recent Advances in Multi-Temperature-Responsive Polymeric Materials. Polym. J. 2020, 52, 681–689. [Google Scholar] [CrossRef] [Scilit]
  2. Ward, M.A.; Georgiou, T.K. Thermoresponsive Polymers for Biomedical Applications. Polymers 2011, 3, 1215–1242. [Google Scholar] [CrossRef] [Scilit]
  3. Fu, X.; Xing, C.; Sun, J. Tunable LCST/UCST-Type Polypeptoids and Their Structure–Property Relationship. Biomacromolecules 2020, 21, 4980–4988. [Google Scholar] [CrossRef] [Scilit]
  4. Zhang, R.; Cheng, H.; Zhang, C.; Sun, T.; Dong, X.; Han, C.C. Phase Separation Mechanism of Polybutadiene/Polyisoprene Blends under Oscillatory Shear Flow. Macromolecules 2008, 41, 6818–6829. [Google Scholar] [CrossRef] [Scilit]
  5. Kawahara, S.; Asada, Y.; Isono, Y.; Muraoka, K.; Minagawa, Y. Lower Critical Solution Temperature Phase Behavior of Natural Rubber/Polybutadiene Blend. Polym. J. 2002, 34, 1–8. [Google Scholar] [CrossRef] [Scilit]
  6. Sakurai, S.; Jinnai, H.; Hasegawa, H.; Hashimoto, T.; Han, C.C. Microstructure Effects on the Lower Critical Solution Temperature Phase Behavior of Deuterated Polybutadiene and Protonated Polyisoprene Blends Studied by Small-Angle Neutron Scattering. Macromolecules 1991, 24, 4839–4843. [Google Scholar] [CrossRef] [Scilit]
  7. Delmas, G.; Patterson, D. The Molecular Weight Dependence of Lower and Upper Critical Solution Temperatures. J. Polym. Sci. Part C Polym. Symp. 1970, 30, 1–8. [Google Scholar] [CrossRef] [Scilit]
  8. García-Peñas, A.; Biswas, C.S.; Liang, W.; Wang, Y.; Yang, P.; Stadler, F.J. Effect of Hydrophobic Interactions on Lower Critical Solution Temperature for Poly(N-Isopropylacrylamide-Co-Dopamine Methacrylamide) Copolymers. Polymers 2019, 11, 991. [Google Scholar] [CrossRef] [Scilit]
  9. Lu, J.; Xu, M.; Lei, Y.; Gong, L.; Zhao, C. Aqueous Synthesis of Upper Critical Solution Temperature and Lower Critical Solution Temperature Copolymers through Combination of Hydrogen-Donors and Hydrogen-Acceptors. Macromol. Rapid Commun. 2021, 42, 2000661. [Google Scholar] [CrossRef] [Scilit]
  10. Taylor, M.; Tomlins, P.; Sahota, T. Thermoresponsive Gels. Gels 2017, 3, 4. [Google Scholar] [CrossRef] [Scilit]
  11. Zhang, Q.; Weber, C.; Schubert, U.S.; Hoogenboom, R. Thermoresponsive Polymers with Lower Critical Solution Temperature: From Fundamental Aspects and Measuring Techniques to Recommended Turbidimetry Conditions. Mater. Horiz. 2017, 4, 109–116. [Google Scholar] [CrossRef] [Scilit]
  12. Flemming, P.; Münch, A.S.; Fery, A.; Uhlmann, P. Constrained Thermoresponsive Polymers—New Insights into Fundamentals and Applications. Beilstein J. Org. Chem. 2021, 17, 2123–2163. [Google Scholar] [CrossRef] [Scilit]
  13. Zhu, Y.; Batchelor, R.; Lowe, A.B.; Roth, P.J. Design of Thermoresponsive Polymers with Aqueous LCST, UCST, or Both: Modification of a Reactive Poly(2-Vinyl-4,4-Dimethylazlactone) Scaffold. Macromolecules 2016, 49, 672–680. [Google Scholar] [CrossRef] [Scilit]
  14. Gharakhanian, E.G.; Deming, T.J. Role of Side-Chain Molecular Features in Tuning Lower Critical Solution Temperatures (LCSTs) of Oligoethylene Glycol Modified Polypeptides. J. Phys. Chem. B 2016, 120, 6096–6101. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  15. Jung, S.-H.; Lee, H.-I. Well-Defined Thermoresponsive Copolymers with Tunable LCST and UCST in Water. Bull. Korean Chem. Soc. 2014, 35, 501–504. [Google Scholar] [CrossRef] [Scilit]
  16. Pasparakis, G.; Tsitsilianis, C. LCST Polymers: Thermoresponsive Nanostructured Assemblies towards Bioapplications. Polymer 2020, 211, 123146. [Google Scholar] [CrossRef] [Scilit]
  17. Papadakis, C.M.; Müller-Buschbaum, P.; Laschewsky, A. Switch It Inside-Out: “Schizophrenic” Behavior of All Thermoresponsive UCST–LCST Diblock Copolymers. Langmuir 2019, 35, 9660–9676. [Google Scholar] [CrossRef] [Scilit]
  18. Seuring, J.; Agarwal, S. Polymers with Upper Critical Solution Temperature in Aqueous Solution. Macromol. Rapid Commun. 2012, 33, 1898–1920. [Google Scholar] [CrossRef] [Scilit]
  19. Wu, J.-Q.; Gong, X.-Q.; Wang, Q.; Yan, F.; Li, J.-J. A QSPR Study for Predicting θ(LCST) and θ(UCST) in Binary Polymer Solutions. Chem. Eng. Sci. 2023, 267, 118326. [Google Scholar] [CrossRef] [Scilit]
  20. Throat, S.; Bhattacharya, S. Macromolecular Poly(N-Isopropylacrylamide) (PNIPAM) in Cancer Treatment and Beyond. Adv. Polym. Technol. 2024, 2024, 1444990. [Google Scholar] [CrossRef] [Scilit]
  21. Roy, D.; Brooks, W.L.A.; Sumerlin, B.S. New Directions in Thermoresponsive Polymers. Chem. Soc. Rev. 2013, 42, 7214–7243. [Google Scholar] [CrossRef] [Scilit]
  22. Sharifi, S.; Bonardd, S.; Miccio, L.A. Dielectric Constant Prediction in Polymers: A Chemical Structure Based Approach. Next Mater. 2025, 8, 100795. [Google Scholar] [CrossRef] [Scilit]
  23. Miccio, L.A.; Borredon, C.; Schwartz, G.A. A Glimpse inside Materials: Polymer Structure—Glass Transition Temperature Relationship as Observed by a Trained Artificial Intelligence. Comput. Mater. Sci. 2024, 236, 112863. [Google Scholar] [CrossRef] [Scilit]
  24. Borredon, C.; Miccio, L.A.; Cerveny, S.; Schwartz, G.A. Characterising the Glass Transition Temperature-Structure Relationship through a Recurrent Neural Network. J. Non-Cryst. Solids X 2023, 18, 100185. [Google Scholar] [CrossRef] [Scilit]
  25. Ethier, J.G.; Casukhela, R.K.; Latimer, J.J.; Jacobsen, M.D.; Shantz, A.B.; Vaia, R.A. Deep Learning of Binary Solution Phase Behavior of Polystyrene. ACS Macro Lett. 2021, 10, 749–754. [Google Scholar] [CrossRef] [Scilit]
  26. Köster, Y.; Kimmig, J.; Zechel, S.; Schubert, U.S. Fingerprint Applicable for Machine Learning Tested on LCST Behavior of Polymers. Cell Rep. Phys. Sci. 2023, 4, 101553. [Google Scholar] [CrossRef] [Scilit]
  27. Lutz, J.-F.; Weichenhan, K.; Akdemir, Ö.; Hoth, A. About the Phase Transitions in Aqueous Solutions of Thermoresponsive Copolymers and Hydrogels Based on 2-(2-Methoxyethoxy)Ethyl Methacrylate and Oligo(Ethylene Glycol) Methacrylate. Macromolecules 2007, 40, 2503–2508. [Google Scholar] [CrossRef] [Scilit]
  28. Lemanowicz, M.; Gierczycki, A.; Kuźnik, W.; Sancewicz, R.; Imiela, P. Determination of Lower Critical Solution Temperature of Thermosensitive Flocculants. Miner. Eng. 2014, 69, 170–176. [Google Scholar] [CrossRef] [Scilit]
  29. Liu, H.Y.; Zhu, X.X. Lower critical solution temperatures of N-substituted acrylamide copolymers in aqueous solutions. Polymer 1999, 40, 6985–6990. [Google Scholar] [CrossRef] [Scilit]
  30. Carrick, B.R.; Seitzinger, C.L.; Lodge, T.P. Unusual Lower Critical Solution Temperature Phase Behavior of Poly(benzyl methacrylate) in a Pyrrolidinium-Based Ionic Liquid. Molecules 2021, 26, 4850. [Google Scholar] [CrossRef] [Scilit]
  31. Miccio, L.A. Understanding Polymers Through Transfer Learning and Explainable AI. Appl. Sci. 2024, 14, 10413. [Google Scholar] [CrossRef] [Scilit]
  32. Borredon, C.; Miccio, L.A.; Schwartz, G.A. Transfer Learning-Driven Artificial Intelligence Model for Glass Transition Temperature Estimation of Molecular Glass Formers Mixtures. Comput. Mater. Sci. 2024, 238, 112931. [Google Scholar] [CrossRef] [Scilit]
  33. Farajzadehahary, K.; Hamzehlou, S.; Ballard, N. Adding Machine Learning to the Polymer Reaction Engineering Toolbox. Prog. Polym. Sci. 2025, 170, 102029. [Google Scholar] [CrossRef] [Scilit]
  34. Ma, R.; Liu, Z.; Zhang, Q.; Liu, Z.; Luo, T. Evaluating Polymer Representations via Quantifying Structure–Property Relationships. J. Chem. Inf. Model. 2019, 59, 3110–3119. [Google Scholar] [CrossRef] [Scilit]
  35. Casado, U.M.; Altuna, F.I.; Miccio, L.A. Towards Sustainable Material Design: A Comparative Analysis of Latent Space Representations in AI Models. Sustainability 2024, 16, 10681. [Google Scholar] [CrossRef] [Scilit]
  36. Tao, L.; Varshney, V.; Li, Y. Benchmarking Machine Learning Models for Polymer Informatics: An Example of Glass Transition Temperature. J. Chem. Inf. Model. 2021, 61, 5395–5413. [Google Scholar] [CrossRef] [Scilit]
  37. Jaeger, S.; Fulle, S.; Turk, S. Mol2vec: Unsupervised Machine Learning Approach with Chemical Intuition. J. Chem. Inf. Model. 2018, 58, 27–35. [Google Scholar] [CrossRef] [Scilit]
  38. Miccio, L.A. Machine Learning-Driven Property Prediction for Materials in Data-Scarcity Scenarios: Ensemble of Experts Approach. Comput. Mater. Sci. 2025, 258, 114092. [Google Scholar] [CrossRef] [Scilit]
  39. Weininger, D. SMILES, a Chemical Language and Information System: 1: Introduction to Methodology and Encoding Rules. J. Chem. Inf. Comput. Sci. 1988, 28, 31–36. [Google Scholar] [CrossRef] [Scilit]
  40. O’Boyle, N.M. Towards a Universal SMILES Representation—A Standard Method to Generate Canonical SMILES Based on the InChI. J. Cheminform. 2012, 4, 22. [Google Scholar] [CrossRef] [Scilit]
  41. Hochreiter, S.; Schmidhuber, J. Long Short-Term Memory. Neural Comput. 1997, 9, 1735–1780. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  42. Lecun, Y.; Bottou, L.; Bengio, Y.; Haffner, P. Gradient-Based Learning Applied to Document Recognition. Proc. IEEE 1998, 86, 2278–2324. [Google Scholar] [CrossRef] [Scilit]
  43. Terven, J.; Cordova-Esparza, D.M.; Romero-González, J.A.; Ramírez-Pedraza, A.; Chávez-Urbiola, E.A. A comprehensive survey of loss functions and metrics in deep learning. Artif. Intell. Rev. 2025, 58, 195. [Google Scholar] [CrossRef] [Scilit]
  44. Kingma, D.P.; Ba, J.L. Adam: A Method for Stochastic Optimization. In Proceedings of the 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, 7–9 May 2015. [Google Scholar]
Figure 1. Conceptual overview of the workflow and data pipeline. A schematic illustrating the data collection, curation, and preprocessing steps for polymer–solvent systems with reported critical solution temperatures.
Figure 1. Conceptual overview of the workflow and data pipeline. A schematic illustrating the data collection, curation, and preprocessing steps for polymer–solvent systems with reported critical solution temperatures.
Applsci 16 02557 g001
Figure 2. Schematic representation of the direct model architecture. The model uses raw descriptors (SMILES-based features and polymer molecular weight) as input and outputs the predicted cloud point temperatures. Red dashed box “A” is just a guide for the eyes of the reader and separates the model architecture from input and output.
Figure 2. Schematic representation of the direct model architecture. The model uses raw descriptors (SMILES-based features and polymer molecular weight) as input and outputs the predicted cloud point temperatures. Red dashed box “A” is just a guide for the eyes of the reader and separates the model architecture from input and output.
Applsci 16 02557 g002
Figure 3. Scheme of the architecture of the encoder–decoder model used to produce fingerprints (“encoded chemical structure”) from molecular SMILES strings.
Figure 3. Scheme of the architecture of the encoder–decoder model used to produce fingerprints (“encoded chemical structure”) from molecular SMILES strings.
Applsci 16 02557 g003
Figure 4. Architecture of the expert model for cloud point temperature prediction. This scheme depicts the model that uses the molecular fingerprints (generated by the encoder) and polymer molecular weight to predict cloud point temperatures.
Figure 4. Architecture of the expert model for cloud point temperature prediction. This scheme depicts the model that uses the molecular fingerprints (generated by the encoder) and polymer molecular weight to predict cloud point temperatures.
Applsci 16 02557 g004
Figure 5. Results of hyperparameter optimization during grid training (MAE and R2 values in color scale, for models 0 to 17 trained under different batch sizes and learning rates). This figure displays the performance comparison, showing that most embedding-based models widely outperform their equivalent direct model counterparts.
Figure 5. Results of hyperparameter optimization during grid training (MAE and R2 values in color scale, for models 0 to 17 trained under different batch sizes and learning rates). This figure displays the performance comparison, showing that most embedding-based models widely outperform their equivalent direct model counterparts.
Applsci 16 02557 g005
Figure 6. (a) Predictedvs. True cloud point temperatures (in °C) during training for the best performing model (model 7). (b) Corresponding absolute (predicted–true) deviation histograms during training. (c) Predicted vs. True cloud point temperatures on the validation set (model 7). (d) Histogram of deviations on the validation set. (e) Predicted vs. True cloud point temperatures on the independent test set.
Figure 6. (a) Predictedvs. True cloud point temperatures (in °C) during training for the best performing model (model 7). (b) Corresponding absolute (predicted–true) deviation histograms during training. (c) Predicted vs. True cloud point temperatures on the validation set (model 7). (d) Histogram of deviations on the validation set. (e) Predicted vs. True cloud point temperatures on the independent test set.
Applsci 16 02557 g006
Figure 7. Predicted vs. True cloud point temperatures on the independent test set with the corresponding chemical structures of the repeating unit and polymer molecular weight.
Figure 7. Predicted vs. True cloud point temperatures on the independent test set with the corresponding chemical structures of the repeating unit and polymer molecular weight.
Applsci 16 02557 g007
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Penoff, M.E.; Altuna, F.I.; Miccio, L.A. Cloud Point Temperature of Thermoresponsive Systems: A Predictive Approach in Data Scarcity Conditions. Appl. Sci. 2026, 16, 2557. https://doi.org/10.3390/app16052557

AMA Style

Penoff ME, Altuna FI, Miccio LA. Cloud Point Temperature of Thermoresponsive Systems: A Predictive Approach in Data Scarcity Conditions. Applied Sciences. 2026; 16(5):2557. https://doi.org/10.3390/app16052557

Chicago/Turabian Style

Penoff, Marcela Elisabeth, Facundo Ignacio Altuna, and Luis Alejandro Miccio. 2026. "Cloud Point Temperature of Thermoresponsive Systems: A Predictive Approach in Data Scarcity Conditions" Applied Sciences 16, no. 5: 2557. https://doi.org/10.3390/app16052557

APA Style

Penoff, M. E., Altuna, F. I., & Miccio, L. A. (2026). Cloud Point Temperature of Thermoresponsive Systems: A Predictive Approach in Data Scarcity Conditions. Applied Sciences, 16(5), 2557. https://doi.org/10.3390/app16052557

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop