Next Article in Journal
Cytotoxic Activity of Sicilian Red- and White-Grape Seed Oils on Human Liver and Colorectal Cancer Cells
Previous Article in Journal
The Effect of Plant-Based Protein Preparations on Quality and Functional Properties of Cream Filling
Previous Article in Special Issue
Comparative Antioxidant Profiling of Phenolic Acids and Flavonoids: Assay-Resolved Structure–Activity Relationships Under Harmonized In Vitro Conditions
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

When Does Machine Learning Add Value over Theory? Predicting API Solubility in Binary Mixtures with COSMO-RS and DOOIT2 Across Diverse and Homogeneous Systems

by
Maciej Przybyłek
1,2,*,
Tomasz Jeliński
1,
Adrian Drużyński
1 and
Piotr Cysewski
1,2,*
1
Department of Physical Chemistry, Faculty of Pharmacy, Collegium Medicum of Bydgoszcz, Nicolaus Copernicus University in Toruń, Kurpińskiego 5, 85-950 Bydgoszcz, Poland
2
Institute of Advanced Studies, Nicolaus Copernicus University in Toruń, Wileńska 4, 87-100 Toruń, Poland
*
Authors to whom correspondence should be addressed.
Molecules 2026, 31(10), 1566; https://doi.org/10.3390/molecules31101566
Submission received: 31 March 2026 / Revised: 4 May 2026 / Accepted: 6 May 2026 / Published: 8 May 2026
(This article belongs to the Special Issue Organic Molecules in Drug Discovery and Development)

Abstract

Predicting the solubility of active pharmaceutical ingredients (APIs) in binary aqueous-organic mixtures is critical for formulation design, yet remains challenging. Physics-based models such as COSMO-RS provide a solid theoretical foundation but often struggle with non-ideal mixing behavior in complex systems. This study asks a practical question: when does machine learning actually add value beyond established theory? We compared COSMO-RS with DOOIT2 (Dual-Objective Optimization with Iterative Feature Pruning), a hybrid COSMO-RS/machine-learning correction workflow, across two complementary datasets: 85 structurally diverse APIs and related formulation-relevant compounds (10,140 data points) and 37 acid-centered solutes (6030 data points). The datasets also incorporate newly measured solubilities of lidocaine, benzocaine, and vanillic acid in aqueous 4-formylmorpholine mixtures. DOOIT2 employs rigorous API-out Structured Group K-Fold validation with fold-specific ensemble models to ensure realistic assessment of generalization to unseen compounds. The obtained results are dataset-dependent. For the homogeneous acid series, COSMO-RS already delivers strong predictive performance (RMSD = 0.321, R2 = 0.925), and DOOIT2 brings no meaningful improvement (RMSD = 0.310, R2 = 0.923). In contrast, for the diverse API set, DOOIT2 reduces RMSD from 0.686 to 0.527 and increases R2 from 0.829 to 0.849. Residual analysis indicates that prediction uncertainty is driven primarily by the low-solubility region rather than by a simple monotonic dependence on molecular weight alone. These findings delineate the practical boundaries of machine-learning assistance in solubility prediction and offer clear guidance for formulation scientists.

1. Introduction

The solubility of active pharmaceutical ingredients (APIs) in binary aqueous-organic solvent mixtures is one of the most practically important physicochemical properties in drug development. It directly governs key processes such as formulation design, selection of crystallization protocols, dissolution behavior in the gastrointestinal tract, and ultimately the bioavailability of new drug candidates. In many cases, the choice of co-solvent system can dramatically enhance the apparent solubility and dissolution rate of poorly water-soluble compounds, thereby improving their therapeutic potential while reducing the risk of formulation-related failures later in development [1,2,3,4,5].
It is therefore no surprise that considerable research efforts are devoted to improving and predicting API solubility in such systems [6,7,8]. However, reliable solubility data across the full composition range are essential not only for optimizing final dosage forms but also for supporting rational solvent selection during early formulation screening and process scale-up [9,10,11].
In the earliest stages of discovery, when hundreds or even thousands of candidate compounds must be evaluated, exhaustive experimental screening across the full composition and temperature range is simply not feasible. As a result, reliable predictive models are no longer a luxury but a practical necessity. They enable rapid prioritization of promising candidates, guide solvent selection, and significantly reduce the enormous material and time costs associated with purely empirical testing [12,13,14,15].
For many years, physics-based approaches, particularly the Conductor-like Screening Model for Real Solvents (COSMO-RS), have been widely used for such predictions [16,17,18]. In pharmaceutical solubility research, COSMO-RS remains an important reference framework because it provides a theoretically grounded route to solubility estimation and is applicable even when experimental information is limited [16,19,20,21]. For this reason, it offers a natural baseline against which the added value of data-driven correction can be evaluated.
However, despite its theoretical rigor, COSMO-RS does not perform uniformly across all systems, particularly in the context of mixed-solvent solubility prediction. Evidence from previous studies demonstrates that, while the model can provide accurate results for certain systems, it may exhibit substantial deviations in others, especially when complex or non-ideal interactions dominate. For example, in the case of ethenzamide, accurate prediction required extension of the COSMO-RS framework through explicit inclusion of intermolecular interactions (COSMO-RS-DARE), leading to a substantial improvement in predictive performance (R2 = 0.886, MAE = 0.034 in log units) [22]. In contrast, for structurally diverse APIs and more complex solvent environments, baseline COSMO-RS predictions were found to be significantly less accurate, with reported performance as low as R2 = 0.327 and MAE = 0.778 [23]. Intermediate cases, such as edaravone (R2 = 0.760, MAE = 0.310) and theophylline (R2 = 0.449, MAE = 0.048), further illustrate that predictive reliability strongly depends on both solute structure and solvent environment [23]. All reported metrics refer to logarithmic mole fraction solubility. Similar observations have been reported across related systems studied within the COSMO-RS framework. These findings indicate that COSMO-RS is not inherently unreliable, but rather that its performance is system-dependent, with limitations becoming apparent in cases involving complex intermolecular interactions or heterogeneous chemical spaces. The research gap addressed in the present study is therefore not whether COSMO-RS is generally useful, but under what conditions it remains sufficient on its own and when additional correction strategies, such as hybrid data-driven approaches, become justified.
In response to these limitations, machine-learning methods have gained increasing attention in recent years. By learning empirical patterns directly from experimental data, ML models can, in principle, identify and correct for systematic deficiencies in purely theoretical descriptions [24,25,26]. Rather than replacing physics-based models, modern approaches often adopt a hybrid strategy, i.e., using COSMO-RS-derived descriptors as physically meaningful inputs while allowing data-driven components to refine predictions where theory alone falls short [27,28,29]. In this context, DOOIT2 (Dual-Objective Optimization with Iterative Feature Pruning) should be understood as a hybrid COSMO-RS/machine-learning correction workflow rather than an independent predictive model. COSMO-RS predictions (RefSol) are retained as the physically grounded baseline, while machine learning is used to capture systematic residual patterns associated with complex solvation behavior that are not fully described by the theoretical model. This broader conceptual framework is also supported by recent published studies specifically focused on solubility prediction, including the combination of molecular thermodynamics with machine learning for drug solubility in solvents, hybrid semi-mechanistic regression for crystallization-oriented solubility modeling, and comparative thermodynamic/hybrid modeling in pharmaceutical solubility systems [30,31,32].
Over the past decade, a variety of machine-learning architectures, from random forests and gradient boosting to deep neural networks and graph convolutional networks, have been applied to aqueous and mixed-solvent solubility prediction [33,34,35]. While many of these studies report impressive statistics on random train-test splits, a fundamental question has remained largely unanswered: under what conditions does machine learning genuinely improve upon well-established theory, and when does it merely reproduce or even degrade theoretical performance?
Answering this question requires more than another benchmark on a single dataset. It demands systematic comparison across chemically distinct solute classes, consistent and rigorous validation protocols, and explicit evaluation of generalization to truly novel compounds, i.e., the most relevant scenario to pharmaceutical research. In the present work we address exactly this challenge by comparing COSMO-RS with the newly developed DOOIT2 workflow on two deliberately contrasting datasets.
The first dataset (API) contains 85 structurally diverse active pharmaceutical ingredients and related bioactive or formulation-relevant compounds spanning multiple therapeutic classes and chemical scaffolds. It yielded 10,140 solubility records and primarily covers aqueous mixtures with polar aprotic and related organic cosolvents. The second dataset (PhAAc) comprises 37 chemically narrower acid-centered solutes, including phenolic acids, substituted benzoic and cinnamic acids, and related carboxylic or acidic drug-like compounds. This dataset yielded 6030 solubility records and provides broader solvent-space coverage, including aqueous and non-aqueous binary mixtures with alcohols, esters, ketones, nitriles, glycols, carboxylic acids, ethers, and related solvent classes. The two datasets therefore provide complementary test cases: a chemically diverse API-centered set and an acid-centered set with broader solvent coverage. Importantly, the present study incorporates new experimental solubility measurements for lidocaine, benzocaine, and vanillic acid in aqueous mixtures of 4-formylmorpholine, a relatively underexplored but chemically informative solvent system that combines features of dipolar aprotic proton-acceptor media with pronounced preferential-solvation behavior in aqueous mixtures. These new data expand the available solubility literature and provide an additional test case for model transferability to a solvent system that was not heavily represented during training.
To ensure fair and realistic assessment of generalization, we introduced DOOIT2, an evolution of our previously published DOOIT methodology [23]. The key innovation lies in the adoption of API-out Structured Group K-Fold (SGKF) cross-validation combined with fold-specific ensemble models. Unlike conventional random splits, which can inadvertently place structurally similar compounds in both training and test sets, API-out validation guarantees that every prediction for a given compound is made by a model that has never seen any data for that compound. This approach mirrors the real-world situation faced by formulation scientists who must predict the behavior of entirely new molecules. By training a separate ensemble for each validation fold, DOOIT2 also provides a more honest estimate of uncertainty and model stability.
In this study we therefore pose three concrete questions. First, does DOOIT2 outperform COSMO-RS, and if so, for which type of solute class? Second, what are the clear boundaries of the model’s applicability domain, that is which compounds are reliably predicted and which fall outside its learned relationships? Third, how does the methodological shift from random splits to API-out validation affect both performance metrics and the confidence we can place in the results?
By systematically addressing these questions, the present work offers (i) new experimental solubility data for aqueous 4-formylmorpholine systems, involving a less-explored dipolar aprotic solvent with relevance to greener solvent-selection and solvent-replacement strategies; (ii) a systematic comparison of COSMO-RS and ML across chemically distinct datasets; (iii) a rigorous methodology (DOOIT2) for developing and validating generalizable quantitative structure-property relationship (QSPR) models; (iv) a clear delineation of the model’s applicability domain, identifying compounds where predictions are reliable and those where they are not; and (v) a practical tool for formulation scientists seeking to reduce experimental screening in early-stage development. Such insights are essential if we are to move beyond the current “ML versus theory” debate toward a more rational, mechanism-informed strategy for solubility prediction in pharmaceutical science.

2. Results and Discussion

2.1. Experimental Solubility

The experimental solubility of lidocaine, benzocaine, and vanillic acid in 4-formylmorpholine (4FM)-water mixtures at 298.15 K is presented in Figure 1. Full numerical solubility data are provided in the Supplementary Materials in Table S1.1. These three compounds were selected because they represent distinct combinations of functional groups relevant to solvation in mixed solvents. The differences in polarity and intermolecular interaction potential are further illustrated by COSMO-RS σ-surfaces shown in Figure 2, which highlight the distinct distribution of charge density regions for each compound. Although lidocaine and benzocaine are both aromatic local anesthetics, they are not structurally equivalent. Lidocaine contains both amide and tertiary amine groups and has a more flexible structure, whereas benzocaine is a simpler aromatic ester with lower polarity. Vanillic acid differs more markedly because it contains carboxyl, hydroxyl, and methoxy groups, which together provide a distinct balance of polarity and intermolecular interactions.
The choice of 4FM also warrants a brief comment. Compared with routinely used dipolar aprotic solvents such as DMSO, DMF, and acetonitrile, 4FM remains less explored in pharmaceutical solubility studies. The greener-solvent aspect of this medium is relevant and has been noted in earlier solvent-selection and synthesis-oriented studies [36,37,38,39,40,41]. However, this was not the sole reason for its selection in the present work. From a physicochemical perspective, 4FM is a high-boiling dipolar aprotic proton-acceptor solvent, and its aqueous mixtures have been shown to display pronounced composition-dependent solvatochromic and preferential-solvation behavior [42]. This made it a useful test medium for the present study, because it extends the solubility mapping of structurally distinct solutes to an underrepresented class of aqueous proton-acceptor mixed solvents.
As shown in Figure 1, the mole fraction solubility increased steadily with increasing x4FM for all three compounds, and no local maximum was observed within the investigated composition range. In pure water, lidocaine and vanillic acid showed comparable solubility, with xsolute values of 3.13 × 10−4 and 2.91 × 10−4, respectively, whereas benzocaine was the least soluble compound, with xsolute = 1.06 × 10−4. In pure 4FM, benzocaine showed the highest solubility, with xsolute = 3.36 × 10−1, followed closely by lidocaine, with xsolute = 3.30 × 10−1, while vanillic acid remained distinctly less soluble, with xsolute = 2.09 × 10−1.

2.2. Dataset Characteristics

To systematically evaluate the predictive capabilities of COSMO-RS and the DOOIT2 machine learning framework, two complementary datasets were assembled, each designed to probe different dimensions of generalization performance. The datasets differ fundamentally in their chemical composition, solvent coverage, and the nature of the generalization challenge they represent. The first dataset, hereafter referred to as API, comprises 85 structurally diverse solutes centered on active pharmaceutical ingredients and related bioactive or formulation-relevant compounds. The set spans multiple chemical classes, including anti-infective agents, sulfonamides and related antimicrobial compounds, xanthines and other polar heterocycles, aromatic amides and analgesic-type compounds, local anesthetics, nitro- and halo-substituted aromatics, triazoles and other N-heterocycles, and selected sugars or polyol-related compounds. This chemical diversity was intentionally selected to challenge the generalization capability of predictive models because the dataset contains compounds with limited structural similarity to one another. Solubility measurements in the API dataset yielded 10,140 records collected for aqueous mixtures with polar aprotic and related organic cosolvents. The solvent space represented in the curated workbook includes 4-formylmorpholine, dimethylformamide, dimethyl sulfoxide, N-methyl-2-pyrrolidone, acetonitrile, 1,4-dioxane, acetone, formamide, N-methylformamide, and tetrahydrofuran in combination with water. Notably, the compiled datasets include new experimental measurements for lidocaine, benzocaine, and vanillic acid in 4-formylmorpholine-water mixtures, reported here for the first time. These data expand the limited literature on 4FM as a pharmaceutical co-solvent and provide an isothermal test case for model performance in a relatively unexplored solvent system, rather than a full thermodynamic characterization of this solvent pair. Besides equilibrium solubility itself, possible solid-state transformations during equilibration may also influence the interpretation of compound-specific behavior. This aspect, however, was beyond the main scope of the present comparative modeling study. The second dataset, referred to as PhAAc, comprises 37 acidic solutes and is chemically narrower than the API set, although it is not limited to a single benzoic- or cinnamic-acid scaffold. In addition to phenolic acids and substituted benzoic or cinnamic acids, it also contains heteroaromatic and dicarboxylic acids, amino- and hydroxy-acids, and selected acidic drug-like compounds. This composition still provides a substantially more constrained solute space than the API dataset and therefore remains suitable for evaluating model behavior within a chemically narrower domain. Solubility measurements in the PhAAc dataset yielded 6030 records collected across a broader set of aqueous and non-aqueous binary solvent systems. The solvent space extends well beyond DMF, DMSO, and 4FM and includes alcohols, esters, ketones, nitriles, glycols, carboxylic acids, ethers, and related mixed-solvent systems. Within the acid-centered PhAAc dataset, the newly measured vanillic-acid system provides the corresponding experimentally characterized 4FM-containing case. Taken together, the API and PhAAc datasets provide complementary testbeds for evaluating the performance of COSMO-RS and DOOIT2 under distinct generalization scenarios.

2.3. COSMO-RS Baseline Performance

Before evaluating the machine learning models, we established a baseline using COSMO-RS predictions. All calculations were performed using the reference solvent approach, wherein experimental solubility values in the neat solvents, specifically both neat components of the binary solvent system, were supplied as input. This approach effectively anchors the predictions to experimental endpoints and isolates the ability of the model to capture non-ideal mixing behavior across the composition range. Moreover, this setup circumvents the need to provide fusion data, including melting temperature, enthalpy of fusion, and heat capacity differences, which are often unavailable, inconsistent, or problematic for many compounds of pharmaceutical interest.
The use of fusion data in traditional solubility modeling presents several well-documented challenges. First, reliable measurements of melting properties require highly pure crystalline samples and specialized techniques such as differential scanning calorimetry, which may not be available for all compounds, particularly during early-stage drug discovery when only milligram quantities are accessible. Second, many compounds exhibit polymorphism, defined as the existence of multiple crystalline forms with distinct thermodynamic properties, and the fusion data obtained may correspond to a metastable polymorph rather than the thermodynamically stable form relevant to solubility measurements. This ambiguity introduces systematic errors when fusion data are used to convert activity coefficients to absolute solubility. Third, for compounds that decompose before melting or exhibit glass transitions rather than true melting, conventional fusion data are either unattainable or physically meaningless. Fourth, literature compilations of thermophysical data, such as those by Acree and Chickos [43,44], have demonstrated that reported fusion properties may exhibit substantial variability across independent studies, particularly for compounds prone to polymorphism or thermal decomposition. Reported melting temperatures can differ by several degrees, and enthalpy values may vary significantly depending on experimental conditions and sample history. Such inconsistencies introduce additional uncertainty when fusion data are used to derive solubility from thermodynamic relationships. This variability is particularly problematic in pharmaceutical systems, where polymorphism and metastable forms are common. By adopting the reference solvent approach, we bypass these complications entirely and anchor predictions directly to experimentally measured solubility in neat solvents at the temperature of interest. This focuses the modeling effort on the critical challenge of capturing non-ideal mixing behavior across the composition range, independent of the uncertainties inherent in fusion property estimation.
As shown in Figure 3 (left panel), COSMO-RS demonstrated strong predictive performance for the 37 phenolic and carboxylic acids. The root mean square deviation (RMSD) across all 6030 data points was 0.321 log units. The coefficient of determination (R2 = 0.925) indicates that 92.5% of the variance in experimental log solubility is explained by the COSMO-RS predictions. The combination of moderate prediction error and high R2 confirms that, for this solute class, the reference-solvent approach captures the overall solvation trends across the full composition range well. The relatively narrow spread of solubility values in this dataset should also be kept in mind when interpreting the correlation coefficient.
In the case of the API dataset, COSMO-RS exhibited substantially weaker performance, as illustrated in Figure 3 (right panel). The RMSD increased to 0.686 log units, reflecting the greater challenge posed by structurally diverse solutes. At the same time, the R2 value decreased to 0.829, and the scatter around the ideal line was considerably larger, with several compounds exhibiting pronounced systematic deviations. This finding is significant because, despite anchoring calculations to experimental neat-solvent solubility endpoints, the physics-based model does not fully capture the complex and non-additive interactions governing solubility in binary mixtures for structurally diverse solutes. This establishes a clear rationale for applying machine learning to the more challenging API dataset.

2.4. Machine Learning DOOIT2 Models Performance

The contrasting results across the two datasets address the central question of this study by showing that the added value of machine learning is dataset-dependent. However, the determining factor is not simply dataset diversity or homogeneity as abstract categories, but rather the underlying chemical scope of the solutes and the complexity of their solvation behavior. The PhAAc dataset represents a chemically narrower acid-centered domain, dominated by phenolic, benzoic, cinnamic, and related carboxylic or acidic compounds with more recurrent interaction patterns. This chemical coherence helps explain why COSMO-RS already captures the main solvation trends with good accuracy. In contrast, the API dataset encompasses a broader range of chemical scaffolds, from relatively simple aromatic amides to heterocyclic and more structurally diverse drug-like molecules. These chemically varied structures present distinct solvation challenges that are not captured uniformly by a single theoretical baseline.

2.4.1. Model Selection and Validation

The DOOIT2 framework was applied to both datasets using API-out Structured Group K-Fold (SGKF) cross-validation to ensure a rigorous assessment of generalization to unseen compounds. This approach represents a substantial methodological advancement over our previous DOOIT implementation [23,45], which relied on repeated random 80/20 splits. Although random split validation is appropriate for assessing interpolative performance within a chemical series, it systematically overestimates generalization to novel compounds because structurally similar compounds may appear in both the training and test sets. By enforcing chemical separation at the solute level and ensuring that all measurements for a given API are assigned exclusively to either the training or the validation fold, DOOIT2 provides a more honest and realistic assessment of predictive capability. The choice of regression algorithm was not fixed a priori. Instead, multiple ensemble methods were benchmarked for each dataset, and the final selection was based on cross-validated performance under the API-out SGKF framework. As a result, LightGBM (version 4.6.0) was selected for the API dataset, whereas XGBoost (version 3.2.0) performed best for the PhAAc dataset.
For the API dataset, LightGBM with the set 3 feature configuration emerged as the optimal model, whereas XGBoost performed best for the PhAAc dataset.

2.4.2. Descriptor Selection for DOOIT2 Models

The hybrid strategy adopted in this work does not alter the COSMO-RS formalism. The COSMO-RS prediction (RefSol) is retained as the baseline, while additional descriptors derived from COSMO-RS outputs are used exclusively as inputs to a separate empirical correction model. These descriptors should therefore be interpreted as features in a machine learning context rather than as modifications of the underlying thermodynamic equations. An analysis of the features selected by the fold-specific XGBoost models for the PhAAc dataset reveals patterns in both consistency and variability across the folds (Table S4.1 in the Supplementary Materials). RefSol, which represents the COSMO-RS-predicted solubility, was retained as the baseline predictor, whereas the remaining COSMO-RS-derived energetic and σ-profile descriptors were used only as physically informed inputs in a separate empirical correction layer. The fact that RefSol was selected in all five folds confirms that the machine learning model was built around the baseline COSMO-RS estimate rather than as a replacement for it. The recurrent selection of descriptors such as d_HH2 should therefore be interpreted as an empirical feature-selection trend within the hybrid model, not as a reformulation of the underlying COSMO-RS formalism. Several other descriptors showed a high selection frequency across the folds. The hydrophobic region descriptor d_HH3 and the hydrogen bond acceptor descriptor d_HBA4 were both retained with an 80% frequency. Additionally, the solute van der Waals energy E1_vdW_sat and the solvent hydrogen bonding energy E_HB_solvent were frequently retained with a 60% frequency. This pattern suggests that successful prediction requires integrating information from three complementary domains. These include a baseline theoretical prediction, solute-specific energetic terms capturing van der Waals and hydrogen bonding contributions, and an explicit characterization of the interaction landscape of the solvent mixture through σ-profile-derived descriptors. Notably, the number of selected features varied substantially across the folds, ranging from 4 to 19 descriptors, with corresponding fold MAE values ranging from 0.187 to 0.249. This variability does not indicate model instability but rather reflects the fold-specific ensemble approach. Each fold model is optimized for the particular set of compounds held out for validation, and different chemical subspaces may require different descriptor combinations to achieve an optimal prediction. The fold with the lowest MAE of 0.187, designated as Fold 4, selected a parsimonious set of 7 features. Conversely, the fold with the highest MAE of 0.249, designated as Fold 1, selected 19 features. This suggests that certain chemical subspaces inherently require more complex descriptor combinations to achieve comparable accuracy.
An analysis of the features selected by the fold-specific LightGBM models for the API dataset reveals both striking similarities and notable differences compared to the PhAAc models (Table S4.2 in the Supplementary Materials). As observed with the PhAAc dataset, the baseline descriptor RefSol, representing the COSMO-RS predicted solubility, was selected in all five folds, thereby confirming its foundational role regardless of the solute class. Several descriptors demonstrated an equally high selection frequency. The solute van der Waals energy E1_vdW_sat as well as the σ-potential-derived descriptors for the hydrogen bond donor region d_HBD1 and the hydrophobic region d_HH1 were each selected in all five folds. This consistent selection across all chemical subspaces indicates that these descriptors capture fundamental physicochemical properties essential for predicting solubility in binary mixtures irrespective of the specific APIs held out in each fold. The differential descriptors dE_HB_sat and dE_vdW_sat, which represent the relative difference between solute and solvent hydrogen bonding and van der Waals energies, respectively, were selected in 4 out of 5 folds, resulting in an 80% frequency. Their near-universal presence suggests that the model learns to weight the balance between solute and solvent interaction strengths, which effectively captures competitive solvation effects. The API-specific σ-potential descriptors, namely API_HBA4, API_HBA2, and API_HH2, showed lower selection frequencies of 60% and below. This indicates that while certain regions of the solute charge density profile are important for some chemical subspaces, they are not universally required across all folds. Notably, the number of selected features across the folds ranged from 6 to 10 descriptors, with corresponding fold MAE values ranging from 0.332 to 0.403. In contrast to the PhAAc dataset, where the fold with the highest MAE of 0.249 selected the largest number of features at 19, the API dataset showed no clear correlation between model complexity and fold error. This suggests that for structurally diverse APIs, the optimal descriptor set size is more constrained, and the prediction difficulty may arise from factors beyond a simple descriptor count.
A comparative analysis of the descriptors used in the PhAAc and API models is provided in Table 1. The table summarizes the selection frequency, defined as the percentage of SGKF folds in which each descriptor category was retained, for the XGBoost model trained on the PhAAc dataset and the LightGBM model trained on the API dataset. Descriptor categories group related features based on their physicochemical interpretation. The contrasting patterns reveal that the API model places greater emphasis on solute-solvent differential terms and hydrogen bond donor descriptors, whereas the PhAAc model relies more heavily on solvent mixture properties. Descriptor definitions are provided in Supplementary Tables S3.1 and S3.2, whereas per-fold feature-selection results are provided in Supplementary Tables S4.1 and S4.2.
The most striking difference lies in the differential descriptors dE_HB_sat and dE_vdW_sat, which capture the balance between solute and solvent interaction strengths. These features were selected in 80% of the folds for the API model but in only 20 to 40% of the folds for the PhAAc model. This finding is physicochemically informative because, for structurally diverse APIs, the model appears to use these descriptors to capture how the relative strength of solute–solvent interactions varies across compounds with different functional groups. For the chemically narrower PhAAc dataset, where solutes share a more coherent acidic domain, this relative balance is more predictable and can be captured by simpler descriptor combinations. Similarly, the consistent selection of the hydrogen-bond donor descriptor d_HBD1 in all five API folds, compared with only a 40% frequency in the PhAAc model, indicates that donor-capacity variation contributes more strongly to the API model. This observation is consistent with the chemical composition of the two datasets: the PhAAc dataset is dominated by acidic compounds with more recurrent donor/acceptor patterns, whereas the API dataset spans a wider range of donor capacities.
The descriptor selection patterns for the API model reveal a learning strategy that complements the PhAAc findings. Both models invariably retain the COSMO-RS baseline as a foundation. However, the API model places greater emphasis on solute-solvent differential terms that capture the competitive balance between interactions. This balance constitutes a critical factor for diverse chemical structures where the matching between the API and the solvent varies widely. Furthermore, the API model emphasizes hydrogen bond donor descriptors that reflect the broad range of donor capacities across the API set, alongside the solute van der Waals energy. The consistent selection of this van der Waals energy descriptor across all folds indicates that dispersion interactions are universally important regardless of the specific API. The lower frequency of solvent-specific descriptors in the API model compared to the PhAAc model suggests that when solute diversity is high, the model prioritizes capturing the variability in solute properties over detailed solvent characterization. This prioritization likely occurs because solvent properties are already encoded in the RefSol baseline and the differential terms.

2.4.3. Accuracy of DOOIT2 Models

The predictive performance of COSMO-RS and DOOIT2 is summarized in Figure 4, which presents a side-by-side comparison of COSMO-RS and DOOIT2 performance. For the PhAAc dataset, DOOIT2 does not provide a meaningful improvement over COSMO-RS, because the RMSD changes only slightly from 0.321 to 0.310, while R2 remains essentially unchanged at 0.925 versus 0.923. This result is consistent with expectation: when solutes share a chemically narrower and more coherent interaction pattern, the quantum-chemical interactions are already captured well by COSMO-RS, leaving little room for additional data-driven refinement. In contrast, for the structurally diverse API dataset, DOOIT2 improves performance relative to COSMO-RS, reducing the RMSD from 0.686 to 0.527 and increasing R2 from 0.829 to 0.849. This improvement remains moderate in absolute terms and should not be overstated. However, it is worth emphasizing that this result was obtained under a rigorous one-API-out validation framework in which the model was tested on compounds that were not included in training. This is important because model performance can appear better when closely related compounds are present in both the training and test sets. For this reason, validation strategies that reduce such overlap are considered more appropriate when the aim is to evaluate performance for previously unseen compounds [46,47]. Accordingly, the present comparison should be understood as a stricter test of whether machine-learning correction provides added value beyond the baseline COSMO-RS model.
The implication is clear in that the value of machine learning in solubility prediction is not determined by dataset diversity per se but rather by the alignment between the chemical structure of the solute and the theoretical framework. For compounds whose solvation behavior is dominated by interactions that COSMO-RS handles well, such as hydrogen bonding in phenolic acids, theory alone is sufficient. For compounds with more complex or varied interaction landscapes, such as active pharmaceutical ingredients spanning multiple chemical classes, machine learning can identify correction patterns that theoretical models overlook. This insight reframes the common narrative. Rather than viewing machine learning as a universal improvement over theory, it should be understood as a complementary tool whose value depends on the chemical characteristics of the target compounds.

2.5. Error Structure and Applicability Domain

To better understand the predictive behavior of the model, the distribution of residuals was analyzed as a function of both molecular properties and experimental solubility. As shown in the left panel of Figure 5, the residuals plotted against molecular weight do not exhibit a clear systematic trend. Although a slight increase in dispersion can be observed for higher molecular weights, the relationship is not sufficiently pronounced to support a direct dependence of prediction error on molecular size alone. This indicates that molecular weight, as an isolated descriptor, is not a primary determinant of model performance. In contrast, a more distinct pattern emerges when residuals are analyzed as a function of experimental solubility, as shown in the right panel of Figure 5. The model exhibits larger variability in the low-solubility region (more negative log(x)), whereas predictions become more tightly distributed at higher solubility values. This behavior reflects the increased difficulty of accurately predicting very low solubility, where small absolute deviations correspond to larger errors on the logarithmic scale.
These observations suggest that the primary limitation of the model is associated with the intrinsic difficulty of predicting low-solubility systems rather than with simple molecular descriptors such as size or flexibility. Consequently, the applicability domain of the model should be interpreted in terms of solubility regime and data representation rather than strictly structural parameters. The residual analysis indicates that the model performs consistently across a wide range of compounds, with increased uncertainty primarily confined to the low-solubility region.

2.6. The Role of 4-Formylmorpholine

The role of 4-formylmorpholine (4FM) in the present study was primarily to provide an experimentally characterized yet underrepresented solvent environment for testing model transferability. The newly measured 4FM-water systems therefore complement the broader literature-derived datasets by extending them toward a solvent that is less commonly represented in pharmaceutical solubility studies. For these systems, the prediction errors remained consistent with the broader behavior observed for the respective datasets and did not indicate any unusual deterioration associated with 4FM. This, in turn, suggests that the molecular features of 4FM are represented sufficiently well in the descriptor space to permit interpolation despite its more limited representation in the training data. Thus, beyond expanding the experimental knowledge base for 4FM as a pharmaceutical co-solvent, these measurements also provide a practical test of model transferability to a solvent system that bridges amide- and ether-like chemical features.

2.7. Limitations and Future Directions

Several limitations of this study should be acknowledged. First, although the training set for the API dataset is diverse and contains 85 compounds, it remains limited relative to the full chemical space of pharmaceutical compounds. The residual analysis indicates that prediction uncertainty is most pronounced in the low-solubility region, and further expansion of the dataset toward underrepresented structural classes would improve model robustness. Second, while the PhAAc dataset covers 57 solvent systems, it remains a chemically narrower acid-centered dataset. Extending this dataset to additional acidic compound families would help determine whether the present findings generalize beyond the current chemical domain. Third, the current descriptor set does not explicitly account for solute ionization, which may be relevant for ionizable APIs in aqueous mixtures. The incorporation of ionization descriptors, including pKa, speciation, and pH-dependent solubility, could improve predictions for compounds whose solubility is strongly affected by protonation state.
Future work will focus on several key directions. Initial efforts will target the strategic expansion of the API dataset to include more compounds from underperforming classes. This involves adding large and flexible molecules like peptides and macrocycles alongside compounds with basic nitrogen functionalities to fill gaps in the chemical space. Subsequent efforts will extend the PhAAc dataset to include other carboxylic acid families, which will enable an assessment of whether the findings for phenolic acids generalize to aliphatic and dicarboxylic acids. Another objective involves the integration of ionization descriptors to capture pH-dependent solubility effects that are critical for many APIs. Additionally, the transferability of the DOOIT2 framework to other properties such as permeability, stability, and dissolution rate, as well as to other solvent systems like alcohols, polyols, and lipid-based excipients, will be explored in subsequent studies.

3. Materials and Methods

3.1. Materials

Vanillic acid (97%, CAS 121-34-6), 4-formylmorpholine (99%, CAS 4394-85-8), lidocaine (≥98%, CAS 137-58-6), and benzocaine (≥99%, CAS 94-09-7) were obtained from Sigma-Aldrich (St. Louis, MO, USA). Analytical-grade methanol was purchased from Pol-Aura (Morąg, Poland).

3.2. Experimental Determination of Solute Solubility

Solubility was determined using a shake-flask procedure adapted for binary 4FM-water media. The investigated solvent compositions covered the entire composition range of the solute-free binary mixture, from pure water to pure 4FM, with intermediate 4FM mole fractions spaced by 0.1. For each composition, an excess amount of the appropriate solid was introduced into test tubes containing the pre-prepared solvent mixture. The suspensions were equilibrated for 24 h at 298.15 K in an ES-20/60 Orbital Shaker Incubator (Biosan, Riga, Latvia) operated at 60 rpm. The temperature setting accuracy was 0.1 °C, and the variation during the equilibration period did not exceed 0.5 °C.
After incubation, the saturated samples were filtered through PTFE syringe filters with a pore size of 0.22 µm. To reduce the risk of precipitation during sample handling, all accessories coming into contact with the equilibrated solutions, including test tubes, pipette tips, syringes, and filters, were pre-equilibrated at the same temperature as the samples before filtration. Appropriate aliquots of the clear filtrates were then analyzed spectrophotometrically. Three independent saturated samples were prepared and measured for each solvent composition.
Quantification was based on individual calibration curves prepared separately for vanillic acid, benzocaine, and lidocaine. The calibration ranges were 0.00090–0.06370 mg/mL for vanillic acid, 0.00066–0.01340 mg/mL for benzocaine, and 0.02137–0.68384 mg/mL for lidocaine. Absorbance was measured at 295 nm for vanillic acid, 292 nm for benzocaine, and 263 nm for lidocaine using an A360 spectrophotometer (AOE Instruments, Shanghai, China). The corresponding linear calibration equations were A = 33.4789C − 0.0265 for vanillic acid, A = 123.8788C − 0.0201 for benzocaine, and A = 1.4853C − 0.0031 for lidocaine. In all cases, the coefficient of determination (R2) was 0.999. Here, A denotes absorbance and C is the concentration expressed in mg/mL. Each calibration point represented the mean of three measurements. The limits of detection (LOD) and quantification (LOQ), calculated as 3.3σ/S and 10σ/S, respectively, were 0.00010 and 0.00029 mg/mL for vanillic acid, 0.00018 and 0.00056 mg/mL for benzocaine, and 0.00326 and 0.00987 mg/mL for lidocaine.
The density of each saturated solution was determined gravimetrically by weighing 1 mL aliquots transferred with an Eppendorf Reference 2 pipette (Eppendorf AG, Hamburg, Germany) into 10 mL volumetric flasks. The systematic pipette error was 6 μL. Mass measurements were performed using a RADWAG AS 110.R2 PLUS analytical balance (RADWAG, Radom, Poland) with a readability of 0.1 mg. The concentrations determined from the UV-Vis measurements were subsequently converted into mole fraction solubilities using the experimentally determined solution densities and the molar masses of the solute, 4FM, and water according to Equation (1):
X S o l u t e = C S o l u t e M S o l u t e C S o l u t e M S o l u t e + 1000 · ρ C S o l u t e x 4 F M M 4 F M + x H 2 O M H 2 O
where XSolute is the mole fraction solubility of the solute, CSolute is the solute concentration expressed in mg/mL, ρ is the density of the saturated solution expressed in g/mL, MSolute, M4FM, and MH2O are the molar masses of the solute, 4FM, and water, respectively, and x4FM and xH2O denote the mole fractions of 4FM and water in the solute-free binary solvent mixture. All experiments were performed in triplicate.

3.3. COSMO-RS Computations

Quantum-chemical calculations were performed using COSMO-RS theory as implemented in COSMOtherm (version 2024, COSMOlogic GmbH & Co. KG, Leverkusen, Germany). For each solute and solvent, geometry optimizations were carried out at the BP86/TZVP level of theory using the TURBOMOLE (version 7.8) program package. The resulting COSMO files were used to generate σ-profiles and σ-potentials. Prior to COSMO calculations, all structures were fully geometry-optimized at the BP86/TZVP level of theory, ensuring proper pre-optimization before σ-profile generation.
Solubility predictions were obtained using the reference solvent approach, which anchors the calculation to experimental solubility values in neat solvents. For each solute, the experimental solubility in pure water and pure organic solvent at the relevant temperature was supplied as input. The computed solubility in binary mixtures was then obtained via interpolation based on COSMO-RS interaction energies. This approach effectively removes the error associated with solute fusion properties and focuses the prediction entirely on mixture effects.
A comprehensive set of molecular descriptors was extracted from the COSMO-RS output and used as model inputs in the DOOIT2 workflow, whereas the prediction target was the experimental decadal logarithm of mole fraction solubility, log(xexp). The baseline theoretical predictor was RefSol, i.e., the COSMO-RS decadal logarithm of mole fraction solubility obtained with the reference-solvent approach. As summarized in Table 2, Set 1 comprised energetic and chemical-potential descriptors derived from COSMO-RS for the solute, the solvent mixture, and their relative differences under saturated conditions.
Set 2 extended this representation by adding the σ-potential descriptor block. The standard COSMO-RS σ-profile, comprising 61 data points across a charge density range from –0.03 to +0.03 e·Å−2, was condensed by averaging the values over intervals of 0.005 e·Å−2. This produced a 12-step function covering the hydrogen bond donor (HBD), hydrophobic (HH), and hydrogen bond acceptor (HBA) regions. Descriptors were then defined for the solute, the solvent mixture, and their relative differences, as detailed in Supplementary Tables S3.1 and S3.2.
Set 3 does not represent a separate descriptor-generation step. Instead, it denoted the reduced feature configuration retained after iterative pruning and model selection within the DOOIT2 procedure. Accordingly, the final PhAAc model was selected from the Set 2 representation, whereas the final API model was selected from the model trained on features selected from Set 2 and reduced to Set 3 during feature selection.

3.4. Datasets

Two complementary datasets were used in this study, and both were curated with explicit source tracking at the level of the solute-solvent pair. Details of the content are provided in the Supporting Materials (see Section S2). This procedure was introduced to standardize compound names, harmonize solvent labels, identify duplicated or order-inverted solvent-pair entries across literature sources, and resolve obvious bibliographic formatting inconsistencies while preserving transparent provenance for every record retained in the final descriptor table used for model development. The division into the API and PhAAc datasets was operational and based on the intended chemical scope of each set rather than on a formal external taxonomy. The API dataset was designed to maximize structural diversity and therefore grouped active pharmaceutical ingredients together with related bioactive or formulation-relevant solutes measured mainly in aqueous mixtures with polar aprotic and related organic cosolvents. In contrast, the PhAAc dataset was designed as a chemically narrower acid-centered set composed mainly of phenolic acids and related carboxylic acids, while also retaining selected heteroaromatic, dicarboxylic, amino-, and hydroxy-acids that preserved this broader acidic chemical domain. The final API dataset comprised 85 unique solutes and 10,140 solubility records, whereas the final PhAAc dataset comprised 37 unique solutes and 6030 solubility records. In both datasets, the full composition range between the two neat components of a given binary system was retained whenever available, and records obtained at comparable temperatures were merged only after consistency checking. For transparency, the temperature range covered for each solute–solvent system is provided in Supplementary Tables S2.1 and S2.2. The complete pair-level provenance was preserved in the curated workbook used for descriptor generation and model development.
The PhAAc dataset was assembled from literature data for phenolic acids, substituted benzoic and cinnamic acids, heteroaromatic and dicarboxylic acids, amino- and hydroxy-acids, and selected acidic drug-like compounds. In addition to classical phenolic and benzoic-acid derivatives, this dataset also covered compounds such as mesalazine, ibuprofen, naproxen, ketoprofen, indomethacin, artesunate, zaltoprofen, d-histidine, gamma-aminobutyric acid, 4-aminobutyric acid, hydroxyacetic acid, succinic acid, malonic acid, maleic acid, adipic acid, and related systems [48,49,50,51,52,53,54,55,56,57,58,59,60,61,62,63,64,65,66,67,68,69,70,71,72,73,74,75,76,77,78,79,80,81,82,83,84,85,86,87,88,89,90,91,92,93,94,95,96,97,98,99,100,101,102,103,104,105,106,107]. Furthermore, this dataset incorporates the new experimental measurements for vanillic acid in mixtures of 4-formylmorpholine and water reported in the present study.
The API dataset was assembled to maximize structural diversity and included anti-infective agents, sulfonamides and related antimicrobial compounds, xanthines and related polar heterocycles, aromatic amides and analgesic-type compounds, local anesthetics, nitro- and halo-substituted aromatic compounds, triazoles and other N-heterocycles, neutral bioactive compounds, sugars and polyol-related compounds, and several other formulation-relevant solutes retained in the curated workbook. This set included, among others, allopurinol, dapsone, sulfadiazine, sulfamethazine, sulfamethizole, sulfamethoxazole, sulfamethoxypyridazine, sulfapyridine, sulfisomidine, sulfanilamide, carbendazim, ketoconazole, clotrimazole, lamotrigine, maraviroc, ribavirin, acipimox, amrinone, acetanilide, paracetamol, phenacetin, caffeine, theobromine, theophylline, benzamide, benzenesulfonamide, salicylamide, ethenzamide, griseofulvin, gliclazide, isotretinoin, triclocarban, celecoxib, coumarin, chrysin, naringenin, salicin, l-fucose, adenosine, doxofylline, d-histidine, and several nitroaromatic or agrochemical-like compounds [9,54,58,60,68,74,81,87,91,96,98,107,108,109,110,111,112,113,114,115,116,117,118,119,120,121,122,123,124,125,126,127,128,129,130,131,132,133,134,135,136,137,138,139,140,141,142,143,144,145,146,147,148,149,150,151,152,153,154,155,156,157,158,159,160,161,162,163,164,165,166,167,168,169,170,171,172,173,174,175,176,177,178,179,180,181,182]. Furthermore, this dataset incorporates the new experimental measurements for lidocaine and benzocaine in mixtures of 4-formylmorpholine and water reported in the present study.

3.5. Machine Learning Framework: DOOIT2

Predictive models were developed using the DOOIT2 framework, which integrates feature selection, hyperparameter optimization, and model validation into a unified workflow. DOOIT2 represents a significant methodological advancement over our previously published DOOIT framework [23,45,74] by transitioning from repeated random train-test splits to API-out Structured Group K-Fold (SGKF) validation combined with fold-specific ensemble models.

3.5.1. Core Algorithms and Data Preprocessing

Based on preliminary benchmarking, LightGBM (Light Gradient Boosting Machine, version 4.6.0) was selected as the core regressor for the APIs dataset, whereas XGBoost (Extreme Gradient Boosting, version 3.2.0) was employed for the PhAAc dataset. Both algorithms are tree-based ensemble methods capable of capturing complex and non-linear relationships while maintaining interpretability through feature importance metrics. Prior to model training, all descriptors were standardized by removing the mean and scaling to unit variance using the StandardScaler from the scikit-learn library (version 1.7.2). This procedure ensures that no single descriptor disproportionately influences the model due to differences in scale.

3.5.2. API-Out Validation with Fold-Specific Ensemble Models

To rigorously assess generalization to unseen chemical entities, all models were validated using API-out SGKF cross-validation. In this scheme, the dataset was partitioned by solute identity to ensure that all measurements for a given compound were assigned exclusively to either the training or the validation fold. For the APIs dataset, 85 unique solutes were distributed across five folds, with each fold holding out approximately 17 APIs for validation. For the PhAAc dataset, 37 solutes were similarly distributed across five folds.
Crucially, DOOIT2 adopts a fold-specific ensemble modeling approach. Rather than producing a single champion model trained on a fixed training set, the framework independently develops a separate model for each SGKF fold. Each fold-specific model is trained exclusively on data from the remaining four folds, utilizing all solutes except those assigned to the held-out fold, and is optimized to predict solubility for the compounds in its corresponding held-out fold. The result is an ensemble of five models, each with its own potentially distinct feature set and hyperparameters. For any given compound, the prediction is made by the specific fold model for which that compound was held out during training. This ensures that every prediction is genuinely out-of-sample with respect to compound identity.

3.5.3. Dual-Objective Optimization and Iterative Feature Pruning

Model optimization was formulated as a dual-objective problem balancing predictive accuracy against model complexity. These objectives included the minimization of the Mean Absolute Error (MAE) evaluated via 5-fold cross-validation alongside the minimization of model complexity. For tree-based models, this complexity was quantified as the total number of trees multiplied by the average tree depth. Optimization was performed using the Optuna framework (version 3.2), utilizing the Tree-structured Parzen Estimator (TPE) sampler. For each independent run, 2000 optimization trials were conducted to ensure a comprehensive exploration of the hyperparameter space. This process yielded a Pareto front of non-dominated solutions that represent the best achievable trade-off between accuracy and complexity.
Feature selection was integrated into the optimization workflow through iterative backward pruning. Starting with the full descriptor set, a specific sequence of steps was repeated until a minimum feature count was reached. First, dual-objective optimization was performed on the current feature set. A candidate model was then selected from the Pareto front using the one-standard-error (1-SE) rule. This rule identifies the simplest model whose cross-validated MAE falls within one standard error of the best-performing model. Following this selection, feature importance was computed using permutation importance with 10 repetitions, and the least impactful feature was subsequently eliminated. This iterative process generated a family of candidate models at each level of complexity.

3.5.4. Model Selection and Stability Analysis

To ensure the selection of a robust and reproducible model, the entire DOOIT2 procedure was repeated 15 independent times using a different random seed for data partitioning and optimization in each iteration. A multi-criteria selection framework was subsequently applied to identify the final champion model. Architectural stability was assessed by requiring descriptor counts to appear in at least 30% of the independent runs, and the optimal descriptor count was identified as the simplest architecture meeting this threshold. From the architecturally stable group, a specific model instance was selected using a composite scoring system that balanced predictive accuracy with a 50% weight, explanatory power measured by R2 with a 30% weight, and generalization stability based on the train-test performance gap with a 20% weight. The final selected models, specifically LightGBM for the APIs dataset and XGBoost for the PhAAc dataset, were subjected to comprehensive residual analysis and applicability domain assessment to confirm their reliability.

3.5.5. Methodological Advantages

The transition to API-out SGKF with fold-specific ensemble models provides three key benefits. First, it delivers enhanced methodological validity. By enforcing chemical separation at the solute level, DOOIT2 provides a more conservative but fundamentally more honest estimate of predictive power for novel compounds than random split validation. Second, it improves robustness and stability analysis. Evaluating feature stability across fold-specific models ensures that the selected descriptors reflect generalizable physicochemical relationships rather than dataset-specific artifacts. The stability metric of 0.909 for the APIs dataset directly quantifies this reproducibility. Third, it explicitly aligns with real-world applications. Formulation scientists encounter new compounds rather than randomly selected data points, and DOOIT2 validates models under conditions that mirror this reality.

3.5.6. Implementation

The DOOIT2 framework was implemented as a fully automated pipeline in Python 3.10 utilizing the scikit-learn library (version 1.3) for preprocessing and model evaluation, Optuna (version 3.2) for hyperparameter optimization, and pandas (version 2.0) for data management. For each dataset, the pipeline was executed with a 5-fold SGKF to generate five-fold-specific models. The complete code, which includes the implementation of the fold-specific ensemble logic, and the data necessary to reproduce all results are provided in the Supplementary Materials.

3.6. Descriptor Characterization

In line with the methodology detailed in our prior study [23,45], all molecular descriptors were computed from first-principles data using the COSMO-RS (Conductor-like Screening Model for Real Solvents) approach [183,184,185]. The workflow for generating these descriptors proceeds through three defined phases. Initially, a conformational analysis is conducted, involving extensive conformer searches for each solute and solvent molecule via COSMOconf under default settings. To adequately map the conformational space, no more than ten of the most stable conformers are kept for follow-up thermodynamic analysis. Geometry refinement and COSMO output file creation are then performed using TURBOMOLE [186] at the RI-BP/TZVP level, with the TZVPD-FINE basis set (BP_TZVPD_FINE_24.ctd) as per established practices [20,187,188]. Subsequently, the COSMO and energy files required by COSMOtherm [189] to determine interaction energies and chemical potentials are assembled for each solute–solvent system considered in this study. The final mixture calculation outputs provide the descriptors used in the present machine-learning workflow. Their definitions are summarized in Supplementary Tables S3.1 and S3.2.
Apart from energetic and chemical-potential descriptors, the final set of molecular descriptors was augmented with values derived from σ-potential distributions. The standard COSMO-RS output consists of 61 data points covering the charge density range of −0.03e/Å2 + 0.03e/Å2. Consistent with prior machine learning applications, this data was reduced by averaging values over 0.005 intervals. This process resulted in a 12-step function defining three characteristic regions of the σ-potential: hydrogen bond donor (HBD1-4, −0.03e/Å2 to −0.01e/Å2), hydrophobicity, (HH1-4, from −0.01e/Å2 to +0.01e/Å2), and acceptability (Hydrogen Bond Acceptor, HBA1-4, from +0.01e/Å2 to +0.03e/Å2). Consequently, four descriptors were generated for each region, leading to 24 descriptors of this type for the solute, the solvent, and the relative difference between them (Supplementary Table S3.2).

3.7. Performance Metrics

Model performance was evaluated using three complementary metrics describing predictive accuracy, explanatory power, and model reproducibility. The root mean square deviation (RMSD) was used as the primary measure of prediction error, the coefficient of determination (R2) was used to quantify the proportion of variance explained by the model, and a stability metric was used to assess the reproducibility of descriptor selection across independent optimization runs.
The root mean square deviation (RMSD) was calculated as follows:
R M S D = 1 n i = 1 n ( y i y ^ i ) 2
where yi and ŷi are the experimental and predicted log mole fraction solubilities, respectively, and n is the number of observations.
The coefficient of determination (R2) was calculated as follows:
R 2 = 1 i = 1 n ( y i y ^ i ) 2 i = 1 n ( y i y ¯ ) 2
To complement the predictive metrics, model stability for the APIs dataset was quantified as the average similarity of selected feature sets across independent runs, where values approaching 1.0 indicate high reproducibility. For the final API model selected in this study, the stability metric was 0.909. The stability metric was defined as follows:
S t a b i l i t y = 2 k ( k 1 ) i < j F i   F j F i F j
where Fi and Fj represent the feature sets selected in independent runs i and j, and k denotes the total number of runs. All reported performance metrics are based on API-out SGKF cross-validation to ensure that the results reflect true generalization to unseen compounds.

4. Conclusions

This study provides a systematic evaluation of when machine learning adds value beyond COSMO-RS for predicting API solubility in binary mixtures. The results demonstrate that COSMO-RS performs well for chemically homogeneous systems, such as phenolic and carboxylic acids, where solvation behavior is governed by well-defined interaction patterns. In contrast, for structurally diverse APIs, the hybrid DOOIT2 framework provides improved predictive performance, although the magnitude of improvement remains moderate.
Importantly, these findings show that the benefit of machine learning is not universal but depends on the chemical diversity of the solute space and the complexity of solvation interactions. The use of API-out validation highlights that even modest improvements under strict generalization conditions are meaningful. At the same time, residual analysis indicates that prediction uncertainty increases in the low-solubility regime, defining a practical applicability domain.
These results therefore support a complementary strategy in which machine learning is applied selectively, as a refinement of physics-based models, rather than as a universal replacement.

Supplementary Materials

The following supporting information can be downloaded at: https://www.mdpi.com/article/10.3390/molecules31101566/s1, Table S1.1: Newly determined experimental solubility data for lidocaine, benzocaine, and vanillic acid in binary 4-formylmorpholine (4FM)-water mixtures at 298.15 K, obtained in this study. Solubility is expressed as Csolute (mg/mL) and xSolute as a function of the mole fraction of 4FM in the solute-free mixed solvent (x4FM). SD, standard deviation; Table S2.1: Characteristics of the API dataset used for model development. N denotes the number of solubility records taken from the cited source, and temperature range refers to the experimental temperature range covered by the corresponding solute–solvent system; Table S2.2: Characteristics of the PhAAc dataset used for model development. N denotes the number of solubility records taken from the cited source, and temperature range refers to the experimental temperature range covered by the corresponding solute–solvent system; Figure S2.1: Dataset diversity illustrated by the relationship between hydrogen-bond donor character (HBD) and (a) hydrogen-bond acceptor character (HBA) or (b) hydrophobic character (HH) for the API and PhAAc datasets. HBD = HBD1 + HBD2 + HBD3 + HBD4, HBA = HBA1 + HBA2 + HBA3 + HBA4, and HH = HH1 + HH2 + HH3 + HH4; Table S3.1: Detailed explanations of the descriptors used in the study. The acronyms are consistent with the terminology used in the spreadsheet in the Supplementary Materials; Table S3.2: Detailed explanations of the σ-potential related descriptors used in the study. The acronyms are consistent with the terminology used in the spreadsheet in the Supplementary Materials; Table S4.1: Fold-wise feature selection results across SGKF folds for the PhAAc dataset (XGBoost); Table S4.2: Fold-wise feature selection results across SGKF folds for the API dataset (LightGBM).

Author Contributions

Conceptualization, P.C.; methodology, P.C.; formal analysis, P.C.; investigation, M.P., T.J., A.D. and P.C.; resources, T.J. and P.C.; writing—original draft preparation, M.P., T.J. and P.C.; writing—review and editing, M.P., T.J. and P.C.; supervision, P.C.; project administration, T.J. and P.C. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

The original contributions presented in this study are included in the article/Supplementary Materials. Further inquiries can be directed to the corresponding authors.

Acknowledgments

We gratefully acknowledge Polish high-performance computing infrastructure PLGrid (HPC Centers: ACK Cyfronet AGH, WCSS) for providing computer facilities and support within computational grants no. PLG/2025/018825 (Piotr Cysewski) and no. PLG/2026/019188 (Tomasz Jeliński).

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Kawabata, Y.; Wada, K.; Nakatani, M.; Yamada, S.; Onoue, S. Formulation design for poorly water-soluble drugs based on biopharmaceutics classification system: Basic approaches and practical applications. Int. J. Pharm. 2011, 420, 1–10. [Google Scholar] [CrossRef] [Scilit]
  2. Amidon, G.L.; Lennernäs, H.; Shah, V.P.; Crison, J.R. A Theoretical Basis for a Biopharmaceutic Drug Classification: The Correlation of In Vitro Drug Product Dissolution and In Vivo Bioavailability. Pharm. Res. Off. J. Am. Assoc. Pharm. Sci. 1995, 12, 413–420. [Google Scholar] [CrossRef] [Scilit]
  3. Takagi, T.; Ramachandran, C.; Bermejo, M.; Yamashita, S.; Yu, L.X.; Amidon, G.L. A provisional biopharmaceutical classification of the top 200 oral drug products in the United States, Great Britain, Spain, and Japan. Mol. Pharm. 2006, 3, 631–643. [Google Scholar] [CrossRef] [Scilit]
  4. Ghadi, R.; Dand, N. BCS class IV drugs: Highly notorious candidates for formulation development. J. Control. Release 2017, 248, 71–95. [Google Scholar] [CrossRef] [Scilit]
  5. Samineni, R.; Chimakurthy, J.; Konidala, S. Emerging Role of Biopharmaceutical Classification and Biopharmaceutical Drug Disposition System in Dosage form Development: A Systematic Review. Turk. J. Pharm. Sci. 2022, 19, 706–713. [Google Scholar] [CrossRef] [Scilit]
  6. Martínez, F.; Jouyban, A.; Acree, W.E. Pharmaceuticals solubility is still nowadays widely studied everywhere. Pharm. Sci. 2017, 23, 1–2. [Google Scholar] [CrossRef] [Scilit]
  7. Savjani, K.T.; Gajjar, A.K.; Savjani, J.K. Drug solubility: Importance and enhancement techniques. ISRN Pharm. 2012, 2012, 195727. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  8. Coltescu, A.R.; Butnariu, M.; Sarac, I. The importance of solubility for new drug molecules. Biomed. Pharmacol. J. 2020, 13, 577–583. [Google Scholar] [CrossRef] [Scilit]
  9. Cysewski, P.; Jeliński, T.; Przybyłek, M. Finding the Right Solvent: A Novel Screening Protocol for Identifying Environmentally Friendly and Cost-Effective Options for Benzenesulfonamide. Molecules 2023, 28, 5008. [Google Scholar] [CrossRef] [Scilit]
  10. Kolář, P.; Shen, J.-W.; Tsuboi, A.; Ishikawa, T. Solvent selection for pharmaceuticals. Fluid Phase Equilibria 2002, 194–197, 771–782. [Google Scholar] [CrossRef] [Scilit]
  11. González-Miquel, M.; Díaz, I. Green solvent screening using modeling and simulation. Curr. Opin. Green Sustain. Chem. 2021, 29, 100469. [Google Scholar] [CrossRef] [Scilit]
  12. Panapitiya, G.; Girard, M.; Hollas, A.; Sepulveda, J.; Murugesan, V.; Wang, W.; Saldanha, E. Evaluation of Deep Learning Architectures for Aqueous Solubility Prediction. ACS Omega 2022, 28, 40. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  13. Komura, H.; Watanabe, R.; Mizuguchi, K. The Trends and Future Prospective of In Silico Models from the Viewpoint of ADME Evaluation in Drug Discovery. Pharmaceutics 2023, 15, 2619. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  14. Tayyebi, A.; Alshami, A.S.; Rabiei, Z.; Yu, X.; Ismail, N.; Talukder, M.J.; Power, J. Prediction of organic compound aqueous solubility using machine learning: A comparison study of descriptor-based and fingerprints-based models. J. Cheminform. 2023, 15, 99. [Google Scholar] [CrossRef] [Scilit]
  15. Llompart, P.; Minoletti, C.; Baybekov, S.; Horvath, D.; Marcou, G.; Varnek, A. Will we ever be able to accurately predict solubility? Sci. Data 2024, 11, 303. [Google Scholar] [CrossRef] [Scilit]
  16. Klamt, A.; Eckert, F.; Arlt, W. COSMO-RS: An Alternative to Simulation for Calculating Thermodynamic Properties of Liquid Mixtures. Annu. Rev. Chem. Biomol. Eng. 2010, 1, 101–122. [Google Scholar] [CrossRef] [Scilit]
  17. Silva, F.; Veiga, F.; Rodrigues, S.P.J.; Cardoso, C.; Paiva-Santos, A.C. COSMO models for the pharmaceutical development of parenteral drug formulations. Eur. J. Pharm. Biopharm. 2023, 187, 156–165. [Google Scholar] [CrossRef] [Scilit]
  18. Klajmon, M. Purely Predicting the Pharmaceutical Solubility: What to Expect from PC-SAFT and COSMO-RS? Mol. Pharm. 2022, 19, 4212–4232. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  19. Klamt, A.; Schüürmann, G. COSMO: A new approach to dielectric screening in solvents with explicit expressions for the screening energy and its gradient. J. Chem. Soc. Perkin Trans. 1993, 2, 799–805. [Google Scholar] [CrossRef] [Scilit]
  20. Eckert, F.; Klamt, A. Fast solvent screening via quantum chemistry: COSMO-RS approach. AIChE J. 2002, 48, 369–385. [Google Scholar] [CrossRef] [Scilit]
  21. Loschen, C.; Klamt, A. COSMOquick: A Novel Interface for Fast σ-Profile Composition and Its Application to COSMO-RS Solvent Screening Using Multiple Reference Solvents. Ind. Eng. Chem. Res. 2012, 51, 14303–14308. [Google Scholar] [CrossRef] [Scilit]
  22. Cysewski, P. Prediction of ethenzamide solubility in organic solvents by explicit inclusions of intermolecular interactions within the framework of COSMO-RS-DARE. J. Mol. Liq. 2019, 290, 111163. [Google Scholar] [CrossRef] [Scilit]
  23. Cysewski, P.; Jeliński, T.; Giniewicz, J.; Kaźmierska, A.; Przybyłek, M. Duality of Simplicity and Accuracy in QSPR: A Machine Learning Framework for Predicting Solubility of Selected Pharmaceutical Acids in Deep Eutectic Solvents. Molecules 2025, 30, 4361. [Google Scholar] [CrossRef] [Scilit]
  24. Boobier, S.; Hose, D.R.J.; Blacker, A.J.; Nguyen, B.N. Machine learning with physicochemical relationships: Solubility prediction in organic solvents and water. Nat. Commun. 2020, 11, 5753. [Google Scholar] [CrossRef] [Scilit]
  25. Lovrić, M.; Pavlović, K.; Žuvela, P.; Spataru, A.; Lučić, B.; Kern, R.; Wong, M.W. Machine learning in prediction of intrinsic aqueous solubility of drug-like compounds: Generalization, complexity, or predictive ability? J. Chemom. 2021, 35, e3349. [Google Scholar] [CrossRef] [Scilit]
  26. Sodaei, Z.; Ekrami, S.; Hashemianzadeh, S.M. Machine learning analysis of molecular dynamics properties influencing drug solubility. Sci. Rep. 2025, 15, 26955. [Google Scholar] [CrossRef] [Scilit]
  27. Mac Fhionnlaoich, N.; Zeglinski, J.; Simon, M.; Wood, B.; Davin, S.; Glennon, B. A hybrid approach to aqueous solubility prediction using COSMO-RS and machine learning. Chem. Eng. Res. Des. 2024, 209, 67–71. [Google Scholar] [CrossRef] [Scilit]
  28. Oliveira, G.; Wegner, P.H.; de Lima Carvalho, P.V.; Voll, F.A.P.; de Paula Scheer, A.; de Pelegrini Soares, R.; Farias, F.O. Machine learning-enhanced COSMO-SAC for accurate solubility predictions. Fluid Phase Equilibria 2026, 600, 114535. [Google Scholar] [CrossRef] [Scilit]
  29. Wang, W.; Cooley, I.; Alexander, M.R.; Wildman, R.D.; Croft, A.K.; Johnston, B.F. A case study on hybrid machine learning and quantum-informed modelling for solubility prediction of drug compounds in organic solvents. Digit. Discov. 2026, 5, 716–733. [Google Scholar] [CrossRef] [Scilit]
  30. Quilló, G.L.; Bhonsale, S.S.; Collas, A.; Van Impe, J.F.M.; Xiouras, C. Hybrid Semi-mechanistic and Machine Learning Solubility Regression Modeling for Crystallization Process Development. Cryst. Growth Des. 2025, 25, 1111–1127. [Google Scholar] [CrossRef] [Scilit]
  31. Ge, K.; Ji, Y. Novel Computational Approach by Combining Machine Learning with Molecular Thermodynamics for Predicting Drug Solubility in Solvents. Ind. Eng. Chem. Res. 2021, 60, 9259–9268. [Google Scholar] [CrossRef] [Scilit]
  32. Yu, Y.; Sun, C.; Jiang, W. A comprehensive study of pharmaceutics solubility in supercritical solvent through diverse thermodynamic and hybrid Machine learning approaches. Int. J. Pharm. 2024, 664, 124579. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  33. Amiri, M.; Khaleseh, F. Predicting drug solubility in binary solvent mixtures using graph convolutional networks: A comprehensive deep learning approach. Sci. Rep. 2025, 15, 45711. [Google Scholar] [CrossRef] [Scilit]
  34. Bao, Z.; Tom, G.; Cheng, A.; Watchorn, J.; Aspuru-Guzik, A.; Allen, C. Towards the prediction of drug solubility in binary solvent mixtures at various temperatures using machine learning. J. Cheminform. 2024, 16, 117. [Google Scholar] [CrossRef] [Scilit]
  35. Cenci, F.; Diab, S.; Ferrini, P.; Harabajiu, C.; Barolo, M.; Bezzo, F.; Facco, P. Predicting drug solubility in organic solvents mixtures: A machine-learning approach supported by high-throughput experimentation. Int. J. Pharm. 2024, 660, 124233. [Google Scholar] [CrossRef] [Scilit]
  36. Kumar, A.; Jad, Y.E.; El-Faham, A.; de la Torre, B.G.; Albericio, F. Green solid-phase peptide synthesis 4. γ-Valerolactone and N-formylmorpholine as green solvents for solid phase peptide synthesis. Tetrahedron Lett. 2017, 58, 2986–2988. [Google Scholar] [CrossRef] [Scilit]
  37. Liao, J.; Zhang, R.; Jia, X.; Wang, M.; Li, C.; Wang, J.; Tang, R.; Huang, J.; You, H.; Chen, F.-E. Green solvent mixture for ultrasound-assisted solid-phase peptide synthesis: A fast and versatile method and its applications in flow and natural product synthesis. Green Chem. 2024, 26, 10549–10557. [Google Scholar] [CrossRef] [Scilit]
  38. Jordan, A.; Hall, C.G.J.; Thorp, L.R.; Sneddon, H.F. Replacement of Less-Preferred Dipolar Aprotic and Ethereal Solvents in Synthetic Organic Chemistry with More Sustainable Alternatives. Chem. Rev. 2022, 122, 6749–6794. [Google Scholar] [CrossRef] [Scilit]
  39. Murray, P.M.; Bellany, F.; Benhamou, L.; Bučar, D.-K.; Tabor, A.B.; Sheppard, T.D. The application of design of experiments (DoE) reaction optimisation and solvent selection in the development of new synthetic chemistry. Org. Biomol. Chem. 2016, 14, 2373–2384. [Google Scholar] [CrossRef] [Scilit]
  40. Wegner, K.; Barnes, D.; Manzor, K.; Jardine, A.; Moran, D. Evaluation of greener solvents for solid-phase peptide synthesis. Green Chem. Lett. Rev. 2021, 14, 152–163. [Google Scholar] [CrossRef] [Scilit]
  41. Hardegger, L.A.; Mallet, F.; Bianchi, B.; Cai, C.; Grand-Guillaume Perrenoud, A.; Humair, R.; Kaehny, R.; Lanz, S.; Li, C.; Li, J.; et al. Toward a Scalable Synthesis and Process for EMA401, Part I: Late Stage Process Development, Route Scouting, and ICH M7 Assessment. Org. Process Res. Dev. 2020, 24, 1743–1755. [Google Scholar] [CrossRef] [Scilit]
  42. Pasham, F.; Jabbari, M.; Farajtabar, A. Solvatochromic Measurement of KAT Parameters and Modeling Preferential Solvation in Green Potential Binary Mixtures of N-Formylmorpholine with Water, Alcohols, and Ethyl Acetate. J. Chem. Eng. Data 2020, 65, 5458–5466. [Google Scholar] [CrossRef] [Scilit]
  43. Acree, W.; Chickos, J.S. Phase Transition Enthalpy Measurements of Organic and Organometallic Compounds. Sublimation, Vaporization and Fusion Enthalpies From 1880 to 2015. Part 1. C1–C10. J. Phys. Chem. Ref. Data 2016, 45, 033101. [Google Scholar] [CrossRef] [Scilit]
  44. Acree, W.; Chickos, J.S. Phase Transition Enthalpy Measurements of Organic and Organometallic Compounds and Ionic Liquids. Sublimation, Vaporization, and Fusion Enthalpies from 1880 to 2015. Part 2. C11–C192. J. Phys. Chem. Ref. Data 2017, 46, 013104. [Google Scholar] [CrossRef] [Scilit]
  45. Cysewski, P.; Jeliński, T.; Przybyłek, M.; Gliniewicz, N.; Majkowski, M.; Wąs, M. Navigating the Deep Eutectic Solvent Landscape: Experimental and Machine Learning Solubility Explorations of Syringic, p-Coumaric, and Caffeic Acids. Int. J. Mol. Sci. 2025, 26, 10099. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  46. Tropsha, A. Best Practices for QSAR Model Development, Validation, and Exploitation. Mol. Inform. 2010, 29, 476–488. [Google Scholar] [CrossRef] [Scilit]
  47. Tanoli, Z.; Schulman, A.; Aittokallio, T. Validation guidelines for drug-target prediction methods. Expert Opin. Drug Discov. 2025, 20, 31–45. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  48. Marden, J.W.; Dover, M.V. The solubilities of several substances in mixed nonaqueous solutions. J. Am. Chem. Soc. 1916, 38, 1235–1245. [Google Scholar] [CrossRef] [Scilit]
  49. Aydi, A.; Claumann, C.A.; Wüst Zibetti, A.; Abderrabba, M. Differential Scanning Calorimetry Data and Solubility of Rosmarinic Acid in Different Pure Solvents and in Binary Mixtures (Methyl Acetate + Water) and (Ethyl Acetate + Water) from 293.2 to 313.2 K. J. Chem. Eng. Data 2016, 61, 3718–3723. [Google Scholar] [CrossRef] [Scilit]
  50. Manrique, Y.J.; Pacheco, D.P.; Martínez, F. Thermodynamics of Mixing and Solvation of Ibuprofen and Naproxen in Propylene Glycol + Water Cosolvent Mixtures. J. Solut. Chem. 2008, 37, 165–181. [Google Scholar] [CrossRef] [Scilit]
  51. Noubigh, A.; Akermi, A. Solubility and Thermodynamic Behavior of Syringic Acid in Eight Pure and Water + Methanol Mixed Solvents. J. Chem. Eng. Data 2017, 62, 3274–3283. [Google Scholar] [CrossRef] [Scilit]
  52. Gantiva, M.; Martínez, F. Thermodynamic analysis of the solubility of ketoprofen in some propylene glycol+water cosolvent mixtures. Fluid Phase Equilibria 2010, 293, 242–250. [Google Scholar] [CrossRef] [Scilit]
  53. Haq, N.; Siddiqui, N.A.; Shakeel, F. Solubility and molecular interactions of ferulic acid in various (isopropanol + water) mixtures. J. Pharm. Pharmacol. 2017, 69, 1485–1494. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  54. Ruidiaz, M.A.; Delgado, D.R.; Martínez, F.; Marcus, Y. Solubility and preferential solvation of indomethacin in 1,4-dioxane+water solvent mixtures. Fluid Phase Equilibria 2010, 299, 259–265. [Google Scholar] [CrossRef] [Scilit]
  55. Rodríguez, G.A.; Delgado, D.R.; Martínez, F.; Jouyban, A.; Acree, W.E. Solubility of naproxen in ethyl acetate+ethanol mixtures at several temperatures and correlation with the Jouyban–Acree model. Fluid Phase Equilibria 2012, 320, 49–55. [Google Scholar] [CrossRef] [Scilit]
  56. Hu, Y.; Wang, L.; Meng, Z.; Yang, W. Measurement and correlation of the solubility of maleic acid in acetone–ethyl acetate mixtures. Thermochim. Acta 2012, 538, 75–78. [Google Scholar] [CrossRef] [Scilit]
  57. Yu, X.; Shen, Z.; Sun, Q.; Qian, N.; Zhou, C.; Chen, J. Solubilities of Adipic Acid in Cyclohexanol + Cyclohexanone Mixtures and Cyclohexanone + Cyclohexane Mixtures. J. Chem. Eng. Data 2016, 61, 1236–1245. [Google Scholar] [CrossRef] [Scilit]
  58. Liang, J.; Ma, J.; Han, J.; Zheng, M.; Zhao, H. Solubility Determination, Modeling, and Preferential Solvation of Terephthalaldehydic Acid Dissolvend in Aqueous Solvent Mixtures of Methanol, Ethanol, Isopropanol, and N-Methyl-2-pyrrolidone. J. Chem. Eng. Data 2019, 64, 1791–1801. [Google Scholar] [CrossRef] [Scilit]
  59. Zhao, K.; Yang, P.; Du, S.; Li, K.; Li, X.; Li, Z.; Liu, Y.; Lin, L.; Hou, B.; Gong, J. Determination and correlation of solubility and thermodynamics of mixing of 4-aminobutyric acid in mono-solvents and binary solvent mixtures. J. Chem. Thermodyn. 2016, 102, 276–286. [Google Scholar] [CrossRef] [Scilit]
  60. Dali, I.; Aydi, A.; Alberto, C.C.; Wüst, Z.A.; Manef, A. Correlation and semi-empirical modeling of solubility of gallic acid in different pure solvents and in binary solvent mixtures of propan-1-ol + water, propan-2-ol + water and acetonitrile + water from (293.2 to 318.2) K. J. Mol. Liq. 2016, 222, 503–519. [Google Scholar] [CrossRef] [Scilit]
  61. Mahali, K.; Guin, P.S.; Roy, S.; Dolui, B.K. Solubility and solute–solvent interaction phenomenon of succinic acid in aqueous ethanol mixtures. J. Mol. Liq. 2017, 229, 172–177. [Google Scholar] [CrossRef] [Scilit]
  62. Cantillo, E.A.; Delgado, D.R.; Martinez, F. Solution thermodynamics of indomethacin in ethanol+propylene glycol mixtures. J. Mol. Liq. 2013, 181, 62–67. [Google Scholar] [CrossRef] [Scilit]
  63. Oliveira, M.L.N.; Franco, M.R. Solubility of 1,4-butanedioic acid in aqueous solutions of ethanol or 1-propanol. Fluid Phase Equilibria 2012, 326, 50–53. [Google Scholar] [CrossRef] [Scilit]
  64. Noubigh, A. Stearic acid solubility in mixed solvents of (water + ethanol) and (ethanol + ethyl acetate): Experimental data and comparison among different thermodynamic models. J. Mol. Liq. 2019, 296, 112101. [Google Scholar] [CrossRef] [Scilit]
  65. Shen, B.; Wang, Q.; Wang, Y.; Ye, X.; Lei, F.; Gong, X. Solubilities of Adipic Acid in Acetic Acid + Water Mixtures and Acetic Acid + Cyclohexane Mixtures. J. Chem. Eng. Data 2013, 58, 938–942. [Google Scholar] [CrossRef] [Scilit]
  66. Tangirala, R.; De, D.; Aniya, V.; Satyavathi, B.; Thella, P.K.; Srinivasan, M.P.; Parthasarathy, R. Solubility Measurement, Modeling, and Thermodynamic Functions for para-Methoxyphenylacetic Acid in Pure and Mixed Organic and Aqueous Systems. J. Chem. Eng. Data 2018, 63, 3369–3381. [Google Scholar] [CrossRef] [Scilit]
  67. Yang, W.; Lei, Z.; Hu, Y.; Chen, X.; Fu, S. Investigations of the Thermal Properties, Nucleation Kinetics, and Growth of γ-Aminobutyric Acid in Aqueous Ethanol Solution. Ind. Eng. Chem. Res. 2010, 49, 11170–11175. [Google Scholar] [CrossRef] [Scilit]
  68. Bustamante, P. Enthalpy–entropy compensation for the solubility of drugs in solvent mixtures: Paracetamol, acetanilide, and nalidixic acid in dioxane–water. J. Pharm. Sci. 1998, 87, 1590–1596. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  69. AbouEllef, E.M.; Gomaa, E.A.; Mashaly, M.S. Thermodynamic Solvation Parameters for Saturated Benzoic Acid and some of its Derivatives in Binary Mixtures of Ethanol and Water. J. Biochem. Technol. 2018, 9, 42–47. [Google Scholar] [CrossRef] [Scilit]
  70. Sun, R.; Wan, Y.; He, H.; Sha, J.; Li, T.; Ren, B. Solubility of Zaltoprofen in Five Binary Solvents at Various Temperatures: Data Determination and Thermodynamic Modeling. J. Chem. Eng. Data 2020, 65, 2053–2067. [Google Scholar] [CrossRef] [Scilit]
  71. Martínez, J.; Manrique, F. Solubility of Ibuprofen in Some Ethanol + Water Cosolvent Mixtures at Several Temperatures. Lat. Am. J. Pharm. 2007, 26, 344–354. [Google Scholar]
  72. Pacheco, D.P.; Martínez, F. Thermodynamic analysis of the solubility of naproxen in ethanol + water cosolvent mixtures. Phys. Chem. Liq. 2007, 45, 581–595. [Google Scholar] [CrossRef] [Scilit]
  73. Martínez, F.; Peña, M.Á.; Bustamante, P. Thermodynamic analysis and enthalpy–entropy compensation for the solubility of indomethacin in aqueous and non-aqueous mixtures. Fluid Phase Equilibria 2011, 308, 98–106. [Google Scholar] [CrossRef] [Scilit]
  74. Cysewski, P.; Jeliński, T.; Rozalski, R.; Lesniewski, F.; Przybyłek, M. Evaluating the Effectiveness of Reference Solvent Solubility Calculations for Binary Mixtures Based on Pure Solvent Solubility: The Case of Phenolic Acids. Molecules 2025, 30, 4444. [Google Scholar] [CrossRef] [Scilit]
  75. Yuan, Y.; Leng, Y.; Shao, H.; Huang, C.; Shan, K. Solubility of dl-malic acid in water, ethanol and in mixtures of ethanol+water. Fluid Phase Equilibria 2014, 377, 27–32. [Google Scholar] [CrossRef] [Scilit]
  76. Galvão, A.C.; Robazza, W.S.; Bianchi, A.D.; Matiello, J.A.; Paludo, A.R.; Thomas, R. Solubility and thermodynamics of vitamin C in binary liquid mixtures involving water, methanol, ethanol and isopropanol at different temperatures. J. Chem. Thermodyn. 2018, 121, 8–16. [Google Scholar] [CrossRef] [Scilit]
  77. Holguín, A.R.; Rodríguez, G.A.; Cristancho, D.M.; Delgado, D.R.; Martínez, F. Solution thermodynamics of indomethacin in propylene glycol+water mixtures. Fluid Phase Equilibria 2012, 314, 134–139. [Google Scholar] [CrossRef] [Scilit]
  78. Wu, Y.; Qin, Y.; Bai, L.; Kang, Y.; Zhang, Y. Solubility of 3-methyl-4-nitrobenzoic acid in binary solvent mixtures of {(1,4-dioxane, N-methyl-2-pyrrolidone, N,N-dimethylformamide) + methanol} from T = (283.15 to 318.15) K: Experimental determination and thermodynamic modelling. J. Chem. Thermodyn. 2017, 105, 165–172. [Google Scholar] [CrossRef] [Scilit]
  79. Zhang, Y.; Guo, X.; Tang, P.; Xu, J. Solubility of 2,5-Furandicarboxylic Acid in Eight Pure Solvents and Two Binary Solvent Systems at 313.15–363.15 K. J. Chem. Eng. Data 2018, 63, 1316–1324. [Google Scholar] [CrossRef] [Scilit]
  80. Yu, X.; Wu, Y.; Wang, J.; Ulrich, J. Experimental Assessment and Modeling of the Solubility of Malonic Acid in Different Solvents. Chem. Eng. Technol. 2018, 41, 1098–1107. [Google Scholar] [CrossRef] [Scilit]
  81. Shakeel, F.; Haq, N.; Salem-Bekhit, M.M.; Raish, M. Solubility and dissolution thermodynamics of sinapic acid in (DMSO + water) binary solvent mixtures at different temperatures. J. Mol. Liq. 2017, 225, 833–839. [Google Scholar] [CrossRef] [Scilit]
  82. Pacheco, D.P.; Manrique, Y.J.; Martínez, F. Thermodynamic study of the solubility of ibuprofen and naproxen in some ethanol+propylene glycol mixtures. Fluid Phase Equilibria 2007, 262, 23–31. [Google Scholar] [CrossRef] [Scilit]
  83. Huang, Q.; Xie, C.; Li, Y.; Su, N.; Lou, Y.; Hu, X.; Wang, Y.; Bao, Y.; Hou, B. Thermodynamic equilibrium of hydroxyacetic acid in pure and binary solvent systems. J. Chem. Thermodyn. 2017, 108, 76–83. [Google Scholar] [CrossRef] [Scilit]
  84. Li, Z.; He, L.; Yu, X. Solubility Measurements and the Dissolution Behavior of Malonic Acid in Binary Solvent Mixtures of (2-Propanol + Ethyl Acetate) by IKBI Calculations. J. Solut. Chem. 2019, 48, 427–444. [Google Scholar] [CrossRef] [Scilit]
  85. Wüst Zibetti, A.; Aydi, A.; Claumann, C.A.; Eladeb, A.; Adberraba, M. Correlation of solubility and prediction of the mixing properties of rosmarinic acid in different pure solvents and in binary solvent mixtures of ethanol + water and methanol + water from (293.2 to 318.2) K. J. Mol. Liq. 2016, 216, 370–376. [Google Scholar] [CrossRef] [Scilit]
  86. Noubigh, A.; Akremi, A. Solution thermodynamics of trans-Cinnamic acid in (methanol + water) and (ethanol + water) mixtures at different temperatures. J. Mol. Liq. 2019, 274, 752–758. [Google Scholar] [CrossRef] [Scilit]
  87. Imran, S.; Hossain, A.; Mahali, K.; Guin, P.S.; Datta, A.; Roy, S. Solubility and peculiar thermodynamical behaviour of 2-aminobenzoic acid in aqueous binary solvent mixtures at 288.15 to 308.15 K. J. Mol. Liq. 2020, 302, 112566. [Google Scholar] [CrossRef] [Scilit]
  88. Shakeel, F.; Haq, N.; Alanazi, F.K.; Alanazi, S.A.; Alsarra, I.A. Solubility of sinapic acid in various (Carbitol + water) systems: Computational modeling and solution thermodynamics. J. Therm. Anal. Calorim. 2020, 142, 1437–1446. [Google Scholar] [CrossRef] [Scilit]
  89. Takebayashi, Y.; Sue, K.; Furuya, T.; Yoda, S. Solubilities of Organic Semiconductors and Nonsteroidal Anti-inflammatory Drugs in Pure and Mixed Organic Solvents: Measurement and Modeling with Hansen Solubility Parameter. J. Chem. Eng. Data 2018, 63, 3889–3901. [Google Scholar] [CrossRef] [Scilit]
  90. Moradi, M.; Mazaher Haji Agha, E.; Hemmati, S.; Martinez, F.; Kuentz, M.; Jouyban, A. Solubility of 5-aminosalicylic acid in {N-methyl-2-pyrrolidone + ethanol} mixtures at T = (293.2 to 313.2) K. J. Mol. Liq. 2020, 306, 112774. [Google Scholar] [CrossRef] [Scilit]
  91. Li, W.; Farajtabar, A.; Xing, R.; Zhu, Y.; Zhao, H. Solubility of d-Histidine in Aqueous Cosolvent Mixtures of N, N-Dimethylformamide, Ethanol, Dimethyl Sulfoxide, and N-Methyl-2-pyrrolidone: Determination, Preferential Solvation, and Solvent Effect. J. Chem. Eng. Data 2020, 65, 1695–1704. [Google Scholar] [CrossRef] [Scilit]
  92. Shakeel, F.; Haq, N.; Alam, P.; Jouyban, A.; Ghoneim, M.M.; Alshehri, S.; Martinez, F. Solubility of sinapic acid in some (ethylene glycol + water) mixtures: Measurement, computational modeling, thermodynamics, and preferential solvation. J. Mol. Liq. 2022, 348, 118057. [Google Scholar] [CrossRef] [Scilit]
  93. Noubigh, A.; Aydi, A.; Mgaidi, A.; Abderrabba, M. Measurement and correlation of the solubility of gallic acid in methanol plus water systems from (293.15 to 318.15) K. J. Mol. Liq. 2013, 187, 226–229. [Google Scholar] [CrossRef] [Scilit]
  94. Noubigh, A.; Jeribi, C.; Mgaidi, A.; Abderrabba, M. Solubility of gallic acid in liquid mixtures of (ethanol + water) from (293.15 to 318.15) K. J. Chem. Thermodyn. 2012, 55, 75–78. [Google Scholar] [CrossRef] [Scilit]
  95. Ribeiro Neto, A.C.; Pires, R.F.; Malagoni, R.A.; Franco, M.R. Solubility of Vitamin C in Water, Ethanol, Propan-1-ol, Water + Ethanol, and Water + Propan-1-ol at (298.15 and 308.15) K. J. Chem. Eng. Data 2010, 55, 1718–1721. [Google Scholar] [CrossRef] [Scilit]
  96. Matsuda, H.; Kaburagi, K.; Matsumoto, S.; Kurihara, K.; Tochigi, K.; Tomono, K. Solubilities of salicylic acid in pure solvents and binary mixtures containing cosolvent. J. Chem. Eng. Data 2009, 54, 480–484. [Google Scholar] [CrossRef] [Scilit]
  97. Xu, R.; Han, T.; Shen, L.; Zhao, J.; Lu, X. Solubility Determination and Modeling for Artesunate in Binary Solvent Mixtures of Methanol, Ethanol, Isopropanol, and Propylene Glycol + Water. J. Chem. Eng. Data 2019, 64, 755–762. [Google Scholar] [CrossRef] [Scilit]
  98. Jouyban, K.; Mazaher Haji Agha, E.; Hemmati, S.; Martinez, F.; Kuentz, M.; Jouyban, A. Solubility of 5-aminosalicylic acid in N-methyl-2-pyrrolidone + water mixtures at various temperatures. J. Mol. Liq. 2020, 310, 113143. [Google Scholar] [CrossRef] [Scilit]
  99. Mazaher Haji Agha, E.; Barzegar-Jalali, M.; Adibkia, K.; Hemmati, S.; Kuentz, M.; Martinez, F.; Jouyban, A. Solubility of mesalazine in {1-propanol/water} mixtures at different temperatures. J. Mol. Liq. 2020, 301, 112436. [Google Scholar] [CrossRef] [Scilit]
  100. Rezaei, H.; Jouyban, A.; Martinez, F.; Barzegar-Jalali, M.; Hemmati, S.; Rahimpour, E. Solubility and thermodynamic profile of mesalazine in carbitol + ethanol mixtures at different temperatures. J. Mol. Liq. 2021, 324, 114763. [Google Scholar] [CrossRef] [Scilit]
  101. Sheikhi-Sovari, A.; Jouyban, A.; Martinez, F.; Hemmati, S.; Rahimpour, E. Solubility of mesalazine in ethylene glycol + water mixtures at different temperatures. J. Mol. Liq. 2021, 323, 114597. [Google Scholar] [CrossRef] [Scilit]
  102. Jouyban-Gharamaleki, V.; Jouyban, A.; Kuentz, M.; Hemmati, S.; Martinez, F.; Rahimpour, E. A laser monitoring technique for determination of mesalazine solubility in propylene glycol and ethanol mixtures at various temperatures. J. Mol. Liq. 2020, 304, 112714. [Google Scholar] [CrossRef] [Scilit]
  103. Mazaher Haji Agha, E.; Barzegar-Jalali, M.; Adibkia, K.; Hemmati, S.; Martinez, F.; Jouyban, A. Solubility and thermodynamic properties of mesalazine in {2-propanol + water} mixtures at various temperatures. J. Mol. Liq. 2020, 301, 112474. [Google Scholar] [CrossRef] [Scilit]
  104. Noubigh, A.; Akrmi, A. Temperature dependent solubility of vanillic acid in aqueous methanol mixtures: Measurements and thermodynamic modeling. J. Mol. Liq. 2016, 220, 277–282. [Google Scholar] [CrossRef] [Scilit]
  105. Jiménez, D.M.; Muñoz, M.M.; Rodríguez, C.J.; Cárdenas, Z.J.; Martínez, F. Solubility and preferential solvation of some non-steroidal anti-inflammatory drugs in methanol + water mixtures at 298.15 K. Phys. Chem. Liq. 2016, 54, 686–702. [Google Scholar] [CrossRef] [Scilit]
  106. Zhang, Y.; Guo, F.; Cui, Q.; Lu, M.; Song, X.; Tang, H.; Li, Q. Measurement and Correlation of the Solubility of Vanillic Acid in Eight Pure and Water + Ethanol Mixed Solvents at Temperatures from (293.15 to 323.15) K. J. Chem. Eng. Data 2016, 61, 420–429. [Google Scholar] [CrossRef] [Scilit]
  107. Sandeepa, K.; Ravi Kumar, K.; Neeharika, T.S.V.R.; Satyavathi, B.; Thella, P.K. Solubility Measurement and Thermodynamic Modeling of Benzoic Acid in Monosolvents and Binary Mixtures. J. Chem. Eng. Data 2018, 63, 2028–2037. [Google Scholar] [CrossRef] [Scilit]
  108. Cysewski, P.; Przybyłek, M.; Rozalski, R. Experimental and theoretical screening for green solvents improving sulfamethizole solubility. Materials 2021, 14, 5915. [Google Scholar] [CrossRef] [Scilit]
  109. Przybyłek, M.; Kowalska, A.; Tymorek, N.; Dziaman, T.; Cysewski, P. Thermodynamic Characteristics of Phenacetin in Solid State and Saturated Solutions in Several Neat and Binary Solvents. Molecules 2021, 26, 4078. [Google Scholar] [CrossRef] [Scilit]
  110. Cysewski, P.; Jeliński, T.; Cymerman, P.; Przybyłek, M. Solvent Screening for Solubility Enhancement of Theophylline in Neat, Binary and Ternary NADES Solvents: New Measurements and Ensemble Machine Learning. Int. J. Mol. Sci. 2021, 22, 7347. [Google Scholar] [CrossRef] [Scilit]
  111. Osorio, I.P.; Martínez, F.; Peña, M.A.; Jouyban, A.; Acree, W.E. Solubility, dissolution thermodynamics and preferential solvation of sulfadiazine in (N-methyl-2-pyrrolidone + water) mixtures. J. Mol. Liq. 2021, 330, 115693. [Google Scholar] [CrossRef] [Scilit]
  112. Li, H.; Xie, Y.; Li, Z.; Zhao, H. 2-Methoxy-4-nitroaniline Solubility in Several Aqueous Solvent Mixtures: Determination, Modeling, and Preferential Solvation. J. Chem. Eng. Data 2020, 65, 2673–2682. [Google Scholar] [CrossRef] [Scilit]
  113. Cysewski, P.; Jeliński, T.; Przybyłek, M.; Nowak, W.; Olczak, M. Solubility Characteristics of Acetaminophen and Phenacetin in Binary Mixtures of Aqueous Organic Solvents: Experimental and Deep Machine Learning Screening of Green Dissolution Media. Pharmaceutics 2022, 14, 2828. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  114. Delgado, D.R.; Caviedes-Rubio, D.I.; Ortiz, C.P.; Parra-Pava, Y.L.; Peña, M.Á.; Jouyban, A.; Mirheydari, S.N.; Martínez, F.; Acree, W.E. Solubility of sulphadiazine in (acetonitrile + water) mixtures: Measurement, correlation, thermodynamics and preferential solvation. Phys. Chem. Liq. 2020, 58, 381–396. [Google Scholar] [CrossRef] [Scilit]
  115. Jeliński, T.; Stasiak, D.; Kosmalski, T.; Cysewski, P. Experimental and Theoretical Study on Theobromine Solubility Enhancement in Binary Aqueous Solutions and Ternary Designed Solvents. Pharmaceutics 2021, 13, 1118. [Google Scholar] [CrossRef] [Scilit]
  116. Shakeel, F.; Haq, N.; Alshehri, S.; Alenazi, M.; Alwhaibi, A.; Alsarra, I.A. Solubility and Thermodynamic Analysis of Isotretinoin in Different (DMSO + Water) Mixtures. Molecules 2023, 28, 7110. [Google Scholar] [CrossRef] [Scilit]
  117. Jeliński, T.; Bugalska, N.; Koszucka, K.; Przybyłek, M.; Cysewski, P. Solubility of sulfanilamide in binary solvents containing water: Measurements and prediction using Buchowski-Ksiazczak solubility model. J. Mol. Liq. 2020, 319, 114342. [Google Scholar] [CrossRef] [Scilit]
  118. Cysewski, P.; Przybyłek, M.; Kowalska, A.; Tymorek, N. Thermodynamics and intermolecular interactions of nicotinamide in neat and binary solutions: Experimental measurements and COSMO-RS concentration dependent reactions investigations. Int. J. Mol. Sci. 2021, 22, 7365. [Google Scholar] [CrossRef] [Scilit]
  119. Rahimpour, E.; Mazaher Haji Agha, E.; Martinez, F.; Barzegar-Jalali, M.; Jouyban, A. Solubility study of acetaminophen in the mixtures of acetonitrile and water at different temperatures. J. Mol. Liq. 2021, 324, 114708. [Google Scholar] [CrossRef] [Scilit]
  120. Jiménez, D.M.; Cárdenas, Z.J.; Delgado, D.R.; Peña, M.T.; Martínez, F. Solubility temperature dependence and preferential solvation of sulfadiazine in 1,4-dioxane+water co-solvent mixtures. Fluid Phase Equilibria 2015, 397, 26–36. [Google Scholar] [CrossRef] [Scilit]
  121. Cysewski, P.; Przybyłek, M.; Jeliński, T. Predicting sulfanilamide solubility in the binary mixtures using a reference solvent approach. Polim. Med. 2024, 54, 27–34. [Google Scholar]
  122. Przybyłek, M.; Miernicka, A.; Nowak, M.; Cysewski, P. New Screening Protocol for Effective Green Solvents Selection of Benzamide, Salicylamide and Ethenzamide. Molecules 2022, 27, 3323. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  123. Blanco-Márquez, J.H.; Ortiz, C.P.; Cerquera, N.E.; Martínez, F.; Jouyban, A.; Delgado, D.R. Thermodynamic analysis of the solubility and preferential solvation of sulfamerazine in (acetonitrile + water) cosolvent mixtures at different temperatures. J. Mol. Liq. 2019, 293, 111507. [Google Scholar] [CrossRef] [Scilit]
  124. Delgado, D.R.; Peña Fernández, M.Á.; Martínez, F. Preferential solvation of some sulfonamides in 1,4-dioxane + water co-solvent mixtures at 298.15 K according to the inverse Kirkwood-Buff integrals method. Rev. Acad. Colomb. Cienc. Exactas Físicas Nat. 2014, 38, 104–114. [Google Scholar] [CrossRef] [Scilit]
  125. Blanco-Márquez, J.H.; Caviedes Rubio, D.I.; Ortiz, C.P.; Cerquera, N.E.; Martínez, F.; Delgado, D.R. Thermodynamic analysis and preferential solvation of sulfamethazine in acetonitrile + water cosolvent mixtures. Fluid Phase Equilibria 2020, 505, 112361. [Google Scholar] [CrossRef] [Scilit]
  126. Sun, J.; Liu, X.; Fang, Z.; Mao, S.; Zhang, L.; Rohani, S.; Lu, J. Solubility Measurement and Simulation of Rivaroxaban (Form I) in Solvent Mixtures from 273.15 to 323.15 K. J. Chem. Eng. Data 2016, 61, 495–503. [Google Scholar] [CrossRef] [Scilit]
  127. Nozohouri, S.; Shayanfar, A.; Cárdenas, Z.J.; Martinez, F.; Jouyban, A. Solubility of celecoxib in N-methyl-2-pyrrolidone+water mixtures at various temperatures: Experimental data and thermodynamic analysis. Korean J. Chem. Eng. 2017, 34, 1435–1443. [Google Scholar] [CrossRef] [Scilit]
  128. Yu, S.; Xu, X.; Xing, W.; Xue, F.; Cheng, Y. Solubility, thermodynamic parameters, and dissolution properties of gliclazide in seventeen pure solvents at temperatures from 278.15 to 318.15 K. J. Mol. Liq. 2020, 312, 113425. [Google Scholar] [CrossRef] [Scilit]
  129. Chen, X.; Farajtabar, A.; Jia, W.; Zhao, H. Solvent effect on solubility and preferential solvation analysis of buprofezin dissolved in aqueous co-solvent mixtures of N,N-dimethylformamide, ethanol, acetonitrile and isopropanol. J. Chem. Thermodyn. 2019, 138, 179–188. [Google Scholar] [CrossRef] [Scilit]
  130. Chen, J.; Chen, G.; Cong, Y.; Du, C.; Zhao, H. Solubility modelling and preferential solvation of paclobutrazol in co-solvent mixtures of (ethanol, n-propanol and 1,4-dioxane) + water. J. Chem. Thermodyn. 2017, 112, 249–258. [Google Scholar] [CrossRef] [Scilit]
  131. Yu, S.; Cheng, Y.; Xing, W.; Xue, F. Solubility determination and thermodynamic modelling of gliclazide in five binary solvent mixtures. J. Mol. Liq. 2020, 311, 113258. [Google Scholar] [CrossRef] [Scilit]
  132. Xu, R.; Du, Y.; Wang, J.; Farajtabar, A.; Zhao, H. Solubility modelling, solvent effect and preferential solvation of carbendazim in aqueous co-solvent mixtures of N,N-dimethylformamide, methanol, ethanol and n-propanol. J. Chem. Thermodyn. 2019, 128, 87–96. [Google Scholar] [CrossRef] [Scilit]
  133. Wang, L.; Yang, W.; Song, Y.; Gu, Y. Solubility Measurement, Correlation, and Molecular Interactions of 3-Methyl-6-nitroindazole in Different Neat Solvents and Mixed Solvents from T = 278.15 to 328.15 K. J. Chem. Eng. Data 2019, 64, 3260–3269. [Google Scholar] [CrossRef] [Scilit]
  134. Li, W.; Ji, P.; Xu, Y.; Farajtabar, A.; Li, X.; Zhao, H. Maraviroc in aqueous co-solvent solutions of n-propanol, ethanol, dimethyl sulfoxide and N,N-dimethylformamide: Solubility determination, preferential solvation and solvent effect analysis. J. Chem. Thermodyn. 2020, 143, 106044. [Google Scholar] [CrossRef] [Scilit]
  135. Zhao, X.; Farajtabar, A.; Han, G.; Zhao, H. Griseofulvin dissolved in binary aqueous co-solvent mixtures of N,N-dimethylformamide, methanol, ethanol, acetonitrile and N-methylpyrrolidone: Solubility determination and thermodynamic studies. J. Chem. Thermodyn. 2020, 151, 106250. [Google Scholar] [CrossRef] [Scilit]
  136. Caviedes-Rubio, D.I.; Ortiz, C.P.; Martinez, F.; Delgado, D.R. Thermodynamic Assessment of Triclocarban Dissolution Process in N-Methyl-2-pyrrolidone + Water Cosolvent Mixtures. Molecules 2023, 28, 7216. [Google Scholar] [CrossRef] [Scilit]
  137. Chen, J.; Chen, G.; Cheng, C.; Cong, Y.; Li, X.; Zhao, H. Equilibrium solubility, dissolution thermodynamics and preferential solvation of adenosine in aqueous solutions of N,N-dimethylformamide, N-methyl-2-pyrrolidone, dimethylsulfoxide and propylene glycol. J. Chem. Thermodyn. 2017, 115, 52–62. [Google Scholar] [CrossRef] [Scilit]
  138. Cysewski, P.; Jeliński, T.; Przybyłek, M. Exploration of the Solubility Hyperspace of Selected Active Pharmaceutical Ingredients in Choline- and Betaine-Based Deep Eutectic Solvents: Machine Learning Modeling and Experimental Validation. Molecules 2024, 29, 4894. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  139. Shao, D.; Yang, Z.; Zhou, G.; Chen, J.; Zheng, S.; Lv, X.; Li, R. Improving the solubility of acipimox by cosolvents and the study of thermodynamic properties on solvation process. J. Mol. Liq. 2018, 262, 389–395. [Google Scholar] [CrossRef] [Scilit]
  140. Li, X.; Cong, Y.; Li, W.; Yan, P.; Zhao, H. Thermodynamic modelling of solubility and preferential solvation for ribavirin (II) in co-solvent mixtures of (methanol, n-propanol, acetonitrile or 1,4-dioxane) + water. J. Chem. Thermodyn. 2017, 115, 74–83. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  141. Song, S.; Guo, J.; Qiu, J.; Liu, J.; An, M.; Yi, D.; Wang, P.; Zhang, H. Solid–Liquid Equilibrium of Isomaltulose in Five Pure Solvents and Four Binary Solvents from (283.15 to 323.15) K. J. Chem. Eng. Data 2019, 64, 963–971. [Google Scholar] [CrossRef] [Scilit]
  142. Feng, X.; Farajtabar, A.; Lin, H.; Chen, G.; Wang, Z.; Li, X.; Zhao, H. Experimental solubility evaluation and thermodynamic analysis of biologically active D-tryptophan in aqueous mixtures of N,N-dimethylformamide and several alcohols. J. Chem. Thermodyn. 2019, 128, 34–44. [Google Scholar] [CrossRef] [Scilit]
  143. Wu, S.; Shi, Y.; Zhang, H. Solubility Measurement and Correlation for Amrinone in Four Binary Solvent Systems at 278.15–323.15 K. J. Chem. Eng. Data 2020, 65, 4108–4115. [Google Scholar] [CrossRef] [Scilit]
  144. Li, C.; Li, Y.; Gao, X.; Lv, H. Rutaecarpine dissolved in binary aqueous solutions of methanol, ethanol, isopropanol and acetone: Solubility determination, solute-solvent and solvent-solvent interactions and preferential solvation study. J. Chem. Thermodyn. 2020, 151, 106253. [Google Scholar] [CrossRef] [Scilit]
  145. Jeliński, T.; Kubsik, M.; Cysewski, P. Application of the Solute–Solvent Intermolecular Interactions as Indicator of Caffeine Solubility in Aqueous Binary Aprotic and Proton Acceptor Solvents: Measurements and Quantum Chemistry Computations. Materials 2022, 15, 2472. [Google Scholar] [CrossRef] [Scilit]
  146. Bao, Y.; Farajtabar, A.; Zheng, M.; Zhao, H.; Li, Y. Thermodynamic solubility modelling, solvent effect and preferential solvation of naftopidil in aqueous co-solvent solutions of (n-propanol, ethanol, isopropanol and dimethyl sulfoxide). J. Chem. Thermodyn. 2019, 133, 161–169. [Google Scholar] [CrossRef] [Scilit]
  147. Li, X.; Zhu, Y.; Zhang, X.; Farajtabar, A.; Zhao, H. Solubility, Preferential Solvation, and Solvent Effect of Micoflavin in Aqueous Mixtures of Dimethylsulfoxide, Isopropanol, Propylene Glycol, and Ethanol. J. Chem. Eng. Data 2020, 65, 1976–1985. [Google Scholar] [CrossRef] [Scilit]
  148. Tooski, H.F.; Jabbari, M.; Farajtabar, A. Solubility and Preferential Solvation of the Flavonoid Naringenin in Some Aqueous/Organic Solvent Mixtures. J. Solut. Chem. 2016, 45, 1701–1714. [Google Scholar] [CrossRef] [Scilit]
  149. Guo, H.Q.; Li, Y.X.; Chai, X.X. Solubility of 1-methyl-4-nitropyrazole in three binary solvent mixtures from 283.15 K to 323.15 K at 0.1 MPa. J. Mol. Liq. 2019, 291, 111211. [Google Scholar] [CrossRef] [Scilit]
  150. Tinjacá, D.A.; Martínez, F.; Almanza, O.A.; Jouyban, A.; Acree, W.E. Solubility of meloxicam in aqueous binary mixtures of formamide, N-methylformamide and N,N-dimethylformamide: Determination, correlation, thermodynamics and preferential solvation. J. Chem. Thermodyn. 2021, 154, 106332. [Google Scholar] [CrossRef] [Scilit]
  151. Delgado, D.R.; Mogollon-Waltero, E.M.; Ortiz, C.P.; Peña, M.Á.; Almanza, O.A.; Martínez, F.; Jouyban, A. Enthalpy-entropy compensation analysis of the triclocarban dissolution process in some {1,4-dioxane (1) + water (2)} mixtures. J. Mol. Liq. 2018, 271, 522–529. [Google Scholar] [CrossRef] [Scilit]
  152. Fan, J.-P.; Liao, D.-D.; Zhen, B.; Xu, X.-K.; Zhang, X.-H. Measurement and Modeling of the Solubility of Genistin in Water + (Ethanol or Acetone) Binary Solvent Mixtures at T = 278.2–313.2 K. Ind. Eng. Chem. Res. 2015, 54, 12981–12986. [Google Scholar] [CrossRef] [Scilit]
  153. Ortíz, C.P.; Cardenas-Torres, R.E.; Caviedes-Rubio, D.I.; Polania-Orozco, S.D.J.; Delgado, D.R. Thermodynamic analysis and preferential solvation of sulfanilamide in different cosolvent mixtures. Phys. Chem. Liq. 2022, 60, 9–24. [Google Scholar] [CrossRef] [Scilit]
  154. Li, Y.; Li, C.; Gao, X.; Lv, H. Equilibrium solubility, preferential solvation and solvent effect study of clotrimazole in several aqueous co-solvent solutions. J. Chem. Thermodyn. 2020, 151, 106255. [Google Scholar] [CrossRef] [Scilit]
  155. Zhu, C.; Farajtabar, A.; Wu, J.; Zhao, H. 5,7-Dibromo-8-hydroxyquinoline dissolved in binary aqueous co-solvent mixtures of isopropanol, N,N-dimethylformamide, 1,4-dioxane and N-methyl-2-pyrrolidone: Solubility modeling, solvent effect and preferential solvation. J. Chem. Thermodyn. 2020, 148, 106138. [Google Scholar] [CrossRef] [Scilit]
  156. Zhu, Y.; Chen, J.; Zheng, M.; Chen, G.; Farajtabar, A.; Zhao, H. Equilibrium solubility and preferential solvation of 1,1″-sulfonylbis(4-aminobenzene) in binary aqueous solutions of n-propanol, isopropanol and 1,4-dioxane. J. Chem. Thermodyn. 2018, 122, 102–112. [Google Scholar] [CrossRef] [Scilit]
  157. Zhou, Y.; Xu, R.; Zhu, C.; Zhao, H. Solubility Determination and Preferential Solvation of 4-Nitrophthalimide in Binary Aqueous Solutions of Acetone, Ethanol, Isopropanol, and N,N-Dimethylformamide. J. Chem. Eng. Data 2020, 65, 4632–4641. [Google Scholar] [CrossRef] [Scilit]
  158. Barzegar-Jalali, M.; Mazaher Haji Agha, E.; Adibkia, K.; Martinez, F.; Kuentz, M.; Jouyban, A. Solubility of ketoconazole in 1,4-dioxane + water mixtures at T = (293.2 to 313.2) K. J. Mol. Liq. 2020, 306, 112830. [Google Scholar] [CrossRef] [Scilit]
  159. Li, W.; Lin, H.; Song, N.; Chen, G.; Li, X.; Zhao, H. Equilibrium solubility investigation and thermodynamic aspects of biologically active gimeracil (form P) dissolved in aqueous co-solvent mixtures of isopropanol, N,N-dimethylformamide, ethylene glycol and dimethylsulfoxide. J. Chem. Thermodyn. 2019, 133, 19–28. [Google Scholar] [CrossRef] [Scilit]
  160. Li, W.; Farajtabar, A.; Xing, R.; Zhu, Y.; Zhao, H. Equilibrium solubility determination, solvent effect and preferential solvation of amoxicillin in aqueous co-solvent mixtures of N,N-dimethylformamide, isopropanol, N-methyl pyrrolidone and ethylene glycol. J. Chem. Thermodyn. 2020, 142, 106010. [Google Scholar] [CrossRef] [Scilit]
  161. Hatefi, A.; Rahimpour, E.; Ghafourian, T.; Martinez, F.; Barzegar-Jalali, M.; Jouyban, A. Solubility of ketoconazole in N-methyl-2-pyrrolidone + water mixtures at T = (293.2 to 313.2) K. J. Mol. Liq. 2019, 281, 150–155. [Google Scholar] [CrossRef] [Scilit]
  162. Jiménez, D.M.; Cárdenas, Z.J.; Delgado, D.R.; Jouyban, A.; Martínez, F. Solubility and Solution Thermodynamics of Meloxicam in 1,4-Dioxane and Water Mixtures. Ind. Eng. Chem. Res. 2014, 53, 16550–16558. [Google Scholar] [CrossRef] [Scilit]
  163. Qiu, J.; Song, S.; Chen, X.; Yi, D.; An, M.; Wang, P. Determination and Correlation of the Solubility of L-Fucose in Four Binary Solvent Systems at the Temperature Range from 288.15 to 308.15 K. J. Chem. Eng. Data 2018, 63, 3760–3768. [Google Scholar] [CrossRef] [Scilit]
  164. Cárdenas, Z.J.; Jiménez, D.M.; Rodríguez, G.A.; Delgado, D.R.; Martínez, F.; Khoubnasabjafari, M.; Jouyban, A. Solubility of methocarbamol in some cosolvent+water mixtures at 298.15K and correlation with the Jouyban–Acree model. J. Mol. Liq. 2013, 188, 162–166. [Google Scholar] [CrossRef] [Scilit]
  165. Li, Y.; Mou, Y.; Zhu, Y.; Liu, J.; Liu, J.; Zhao, H. Equilibrium solubility determination and thermodynamic aspects of aprepitant (form I) in four binary aqueous mixtures of methanol, ethanol, acetone and 1,4-dioxane. J. Chem. Thermodyn. 2020, 149, 106170. [Google Scholar] [CrossRef] [Scilit]
  166. Li, X.; Liu, Y.; Cao, Y.; Cong, Y.; Farajtabar, A.; Zhao, H. Solubility Modeling, Solvent Effect, and Preferential Solvation of Thiamphenicol in Cosolvent Mixtures of Methanol, Ethanol, N,N-Dimethylformamide, and 1,4-Dioxane with Water. J. Chem. Eng. Data 2018, 63, 2219–2227. [Google Scholar] [CrossRef] [Scilit]
  167. Feizi, S.; Jabbari, M.; Farajtabar, A. A systematic study on solubility and solvation of bioactive compound chrysin in some water + cosolvent mixtures. J. Mol. Liq. 2016, 220, 478–483. [Google Scholar] [CrossRef] [Scilit]
  168. Li, X.; Liu, Y.; Zheng, M.; Zhang, N.; Farajtabar, A.; Zhao, H. Solubility modelling, solvent effect and preferential solvation of allopurinol in aqueous co-solvent mixtures of ethanol, isopropanol, N,N-dimethylformamide and 1-methyl-2-pyrrolidone. J. Chem. Thermodyn. 2019, 131, 478–488. [Google Scholar] [CrossRef] [Scilit]
  169. Yuan, Y.; Farajtabar, A.; Kong, L.; Zhao, H. Thermodynamic solubility modelling, solvent effect and preferential solvation of p-nitrobenzamide in aqueous co-solvent mixtures of dimethyl sulfoxide, ethanol, isopropanol and ethylene glycol. J. Chem. Thermodyn. 2019, 136, 123–131. [Google Scholar] [CrossRef] [Scilit]
  170. Elworthy, P.H.; Worthington, H.E.C. The solubility of sulphadiazine in water-dimethylformamide mixtures. J. Pharm. Pharmacol. 1968, 20, 830–835. [Google Scholar] [CrossRef] [Scilit]
  171. Huang, X.; Wang, J.; Bairu, A.G.; Hao, H. Solid–liquid phase equilibrium and mixing thermodynamic analysis of coumarin in binary solvent mixtures. Phys. Chem. Liq. 2019, 57, 204–220. [Google Scholar] [CrossRef] [Scilit]
  172. Shakeel, F.; Alshehri, S.; Imran, M.; Haq, N.; Alanazi, A.; Anwer, M.K. Experimental and Computational Approaches for Solubility Measurement of Pyridazinone Derivative in Binary (DMSO + Water) Systems. Molecules 2019, 25, 171. [Google Scholar] [CrossRef] [Scilit]
  173. Alshahrani, S.M.; Shakeel, F. Solubility Data and Computational Modeling of Baricitinib in Various (DMSO + Water) Mixtures. Molecules 2020, 25, 2124. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  174. Li, W.; Ma, Y.; Yang, Y.; Xu, S.; Shi, P.; Wu, S. Solubility measurement, correlation and mixing thermodynamics properties of dapsone in twelve mono solvents. J. Mol. Liq. 2019, 280, 175–181. [Google Scholar] [CrossRef] [Scilit]
  175. Huang, H.; Qiu, J.; He, H.; Yi, D.; An, M.; Liu, H.; Hu, S.; Han, J.; Guo, Y.; Wei, N.; et al. Determination and Correlation of the Solubility of d (-)-Salicin in Pure and Binary Solvent Systems. J. Chem. Eng. Data 2020, 65, 4485–4497. [Google Scholar] [CrossRef] [Scilit]
  176. Zhu, C.; Xu, R.; Yin, H.; Zhao, H.; Farajtabar, A. 3,5-dibromo-4-hydroxybenzaldehyde dissolved in aqueous solutions of ethanol, n-propanol, acetonitrile and N,N-dimethylformamide: Solubility modelling, solvent effect and preferential solvation investigation. J. Chem. Thermodyn. 2020, 151, 106252. [Google Scholar] [CrossRef] [Scilit]
  177. Wu, Y.; Qin, Y.; Bai, L.; Kang, Y.; Zhang, Y. Determination and thermodynamic modelling for 4-nitropyrazole solubility in (methanol + water), (ethanol + water) and (acetonitrile + water) binary solvent mixtures from T = (278.15 to 318.15) K. J. Chem. Thermodyn. 2016, 103, 276–284. [Google Scholar] [CrossRef] [Scilit]
  178. Yao, G.; Yao, Q.; Xia, Z.; Li, Z. Solubility determination and correlation for o-phenylenediamine in (methanol, ethanol, acetonitrile and water) and their binary solvents from T = (283.15–318.15) K. J. Chem. Thermodyn. 2017, 105, 179–186. [Google Scholar] [CrossRef] [Scilit]
  179. Romero-Nieto, A.M.; Cerquera, N.E.; Martínez, F.; Delgado, D.R. Thermodynamic study of the solubility of ethylparaben in acetonitrile + water cosolvent mixtures at different temperatures. J. Mol. Liq. 2019, 287, 110894. [Google Scholar] [CrossRef] [Scilit]
  180. Barzegar-Jalali, M.; Mazaher Haji Agha, E.; Mirheydari, S.N.; Adibkia, K.; Martinez, F.; Jouyban, A. Measurement and modelling of the solubility for ketoconazole in {acetonitrile + water} mixtures at T = (293.2 to 313.2) K. Phys. Chem. Liq. 2021, 59, 331–344. [Google Scholar] [CrossRef] [Scilit]
  181. Shen, Y.; Liu, W.; Sun, C.; Yao, T.; Bao, Z. Solubility measurement and solvent effect of Doxofylline in pure solvents and mixtures solvents at 278.15–323.15K. J. Mol. Liq. 2020, 307, 112952. [Google Scholar] [CrossRef] [Scilit]
  182. Mirheydari, S.N.; Barzegar-Jalali, M.; Martinez, F.; Jouyban, A. Solubility of lamotrigine in acetonitrile + water mixtures at various temperatures. Phys. Chem. Liq. 2020, 58, 769–781. [Google Scholar] [CrossRef] [Scilit]
  183. Klamt, A.; Eckert, F.; Hornig, M.; Beck, M.E.; Bürger, T. Prediction of aqueous solubility of drugs and pesticides with COSMO-RS. J. Comput. Chem. 2002, 23, 275–281. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  184. Klamt, A. Conductor-like screening model for real solvents: A new approach to the quantitative calculation of solvation phenomena. J. Phys. Chem. 1995, 99, 2224–2235. [Google Scholar] [CrossRef] [Scilit]
  185. Klamt, A. COSMO-RS: From Quantum Chemistry to Fluid Phase Thermodynamics and Drug Design; Elsevier: Amsterdam, The Netherlands, 2005; ISBN 9780444519948. [Google Scholar]
  186. TURBOMOLE GmbH. TURBOMOLE, Version 7.8; TURBOMOLE GmbH: Karlsruhe, Germany, 2023. [Google Scholar]
  187. Hellweg, A.; Eckert, F. Brick by brick computation of the gibbs free energy of reaction in solution using quantum chemistry and COSMO-RS. AIChE J. 2017, 63, 3944–3954. [Google Scholar] [CrossRef] [Scilit]
  188. Klamt, A.; Jonas, V.; Bürger, T.; Lohrenz, J.C.W. Refinement and parametrization of COSMO-RS. J. Phys. Chem. A 1998, 102, 5074–5085. [Google Scholar] [CrossRef] [Scilit]
  189. Dassault Systèmes. COSMOtherm, Version 24.0.0; BIOVIA: San Diego, CA, USA, 2024.
Figure 1. Experimental mole fraction solubility (xsolute) of lidocaine, benzocaine, and vanillic acid in binary 4-formylmorpholine (4FM)-water mixtures at 298.15 K, plotted as a function of the mole fraction of 4FM in the solute-free solvent mixture (x4FM). Error bars denote standard deviations. The complete numerical solubility data are provided in the Supplementary Materials.
Figure 1. Experimental mole fraction solubility (xsolute) of lidocaine, benzocaine, and vanillic acid in binary 4-formylmorpholine (4FM)-water mixtures at 298.15 K, plotted as a function of the mole fraction of 4FM in the solute-free solvent mixture (x4FM). Error bars denote standard deviations. The complete numerical solubility data are provided in the Supplementary Materials.
Molecules 31 01566 g001
Figure 2. Structural representation of lidocaine, benzocaine, and vanillic acid in the form of a 2D sketch and charge density distribution.
Figure 2. Structural representation of lidocaine, benzocaine, and vanillic acid in the form of a 2D sketch and charge density distribution.
Molecules 31 01566 g002
Figure 3. Parity plots of experimental versus COSMO-RS predicted solubility (log mole fraction) for (left) the PhAAc dataset (37 phenolic and carboxylic acids, 6030 data points) and (right) the API dataset (85 diverse APIs, 10,140 data points). The solid line represents ideal agreement (y = x). For the PhAAc dataset, COSMO-RS achieves RMSD = 0.321 and R2 = 0.925, indicating good overall predictive performance. For the API dataset, performance degrades markedly (RMSD = 0.686, R2 = 0.829), with substantial scatter around the ideal line, reflecting the greater challenge posed by structurally diverse solutes.
Figure 3. Parity plots of experimental versus COSMO-RS predicted solubility (log mole fraction) for (left) the PhAAc dataset (37 phenolic and carboxylic acids, 6030 data points) and (right) the API dataset (85 diverse APIs, 10,140 data points). The solid line represents ideal agreement (y = x). For the PhAAc dataset, COSMO-RS achieves RMSD = 0.321 and R2 = 0.925, indicating good overall predictive performance. For the API dataset, performance degrades markedly (RMSD = 0.686, R2 = 0.829), with substantial scatter around the ideal line, reflecting the greater challenge posed by structurally diverse solutes.
Molecules 31 01566 g003
Figure 4. Predictive performance of COSMO-RS and DOOIT2 for the PhAAc and API datasets. Bar height represents RMSD (left axis), and line markers indicate R2 (right axis) for solubility prediction (log mole fraction) in binary aqueous-organic mixtures. For the PhAAc dataset, COSMO-RS and DOOIT2 yielded RMSD values of 0.321 and 0.310 and R2 values of 0.925 and 0.923, respectively. For the API dataset, the corresponding values were 0.686 and 0.527 for RMSD and 0.829 and 0.849 for R2.
Figure 4. Predictive performance of COSMO-RS and DOOIT2 for the PhAAc and API datasets. Bar height represents RMSD (left axis), and line markers indicate R2 (right axis) for solubility prediction (log mole fraction) in binary aqueous-organic mixtures. For the PhAAc dataset, COSMO-RS and DOOIT2 yielded RMSD values of 0.321 and 0.310 and R2 values of 0.925 and 0.923, respectively. For the API dataset, the corresponding values were 0.686 and 0.527 for RMSD and 0.829 and 0.849 for R2.
Molecules 31 01566 g004
Figure 5. Residual analysis of DOOIT2 model predictions. (left) Residuals (predicted − experimental log mole fraction solubility) plotted as a function of molecular weight (MW). Residuals plotted as a function of experimental solubility (log x_exp) (right). Blue points correspond to the API dataset (LightGBM model), and orange points correspond to the PhAAc dataset (XGBoost model). The plots illustrate the distribution of prediction errors across molecular size and solubility range, highlighting increased variability in the low-solubility region and the absence of a clear dependence of residuals on molecular weight.
Figure 5. Residual analysis of DOOIT2 model predictions. (left) Residuals (predicted − experimental log mole fraction solubility) plotted as a function of molecular weight (MW). Residuals plotted as a function of experimental solubility (log x_exp) (right). Blue points correspond to the API dataset (LightGBM model), and orange points correspond to the PhAAc dataset (XGBoost model). The plots illustrate the distribution of prediction errors across molecular size and solubility range, highlighting increased variability in the low-solubility region and the absence of a clear dependence of residuals on molecular weight.
Molecules 31 01566 g005
Table 1. Comparative Feature Selection of DOOIT2 Models for PhAAc and API Datasets.
Table 1. Comparative Feature Selection of DOOIT2 Models for PhAAc and API Datasets.
Descriptor CategoryPhAAc (XGBoost)API (LightGBM)Key Insight
BaselineRefSol (100%)RefSol (100%)Both rely on COSMO-RS foundation
Solute-solvent differencesLow (20–40%)High (80%)Critical for diverse APIs
Hydrogen bond donorModerate (40%)Very High (100%)Donor capacity varies widely in APIs
Hydrophobic regiond_HH2, d_HH3 (80–100%)d_HH1 (100%)Different hydrophobic bins dominate
Solute van der WaalsModerate (60%)Very High (100%)Universal importance for APIs
Solvent descriptorsHigh (60%)Low (20%)API model focuses on solute variability
Table 2. Summary of target variables, model inputs, and feature configurations used in the DOOIT2 workflow.
Table 2. Summary of target variables, model inputs, and feature configurations used in the DOOIT2 workflow.
ItemRole in the WorkflowDescription
log(xexp)Prediction target (output)Experimental decadal logarithm of mole fraction solubility
RefSolBaseline inputCOSMO-RS predicted decadal logarithm of mole fraction solubility obtained with the reference-solvent approach
Set 1Descriptor configurationEnergetic and chemical-potential descriptors derived from COSMO-RS for the solute, solvent mixture, and their relative differences
Set 2Expanded descriptor configurationSet 1 extended with the σ-potential descriptor block covering HBD, HH, and HBA regions for the solute, solvent mixture, and their relative differences
Set 3Final reduced feature configurationReduced feature set retained after iterative pruning and model selection within DOOIT2; used for the final API model
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Przybyłek, M.; Jeliński, T.; Drużyński, A.; Cysewski, P. When Does Machine Learning Add Value over Theory? Predicting API Solubility in Binary Mixtures with COSMO-RS and DOOIT2 Across Diverse and Homogeneous Systems. Molecules 2026, 31, 1566. https://doi.org/10.3390/molecules31101566

AMA Style

Przybyłek M, Jeliński T, Drużyński A, Cysewski P. When Does Machine Learning Add Value over Theory? Predicting API Solubility in Binary Mixtures with COSMO-RS and DOOIT2 Across Diverse and Homogeneous Systems. Molecules. 2026; 31(10):1566. https://doi.org/10.3390/molecules31101566

Chicago/Turabian Style

Przybyłek, Maciej, Tomasz Jeliński, Adrian Drużyński, and Piotr Cysewski. 2026. "When Does Machine Learning Add Value over Theory? Predicting API Solubility in Binary Mixtures with COSMO-RS and DOOIT2 Across Diverse and Homogeneous Systems" Molecules 31, no. 10: 1566. https://doi.org/10.3390/molecules31101566

APA Style

Przybyłek, M., Jeliński, T., Drużyński, A., & Cysewski, P. (2026). When Does Machine Learning Add Value over Theory? Predicting API Solubility in Binary Mixtures with COSMO-RS and DOOIT2 Across Diverse and Homogeneous Systems. Molecules, 31(10), 1566. https://doi.org/10.3390/molecules31101566

Article Metrics

Back to TopTop