1. Introduction
In petroleum engineering practice, fluid density is one of the key parameters that must be strictly controlled for drilling fluids, completion fluids, as well as workover and kill fluids. Appropriate fluid density is essential not only for maintaining wellbore pressure balance and ensuring well control safety but also for influencing wellbore stability, formation protection, and the smooth execution of field operations. Therefore, formulating working fluids with controllable density, stable performance, and minimum possible cost while satisfying engineering requirements has long been a central objective in field applications.
Completion, workover, and kill fluid systems exhibit considerable diversity and are commonly classified into solid-containing and solids-free systems according to their phase composition. In solid-containing systems, fluid density is primarily increased by introducing weighting materials, which provides a wide adjustable density range and a relatively mature application process.
The problems of weighting material sedimentation and insufficient system stability in ultra-high-density drilling fluids. Based on an ultra-high-density weighting material system, a high-solids dispersant (SMS-19) and an ultra-low-viscosity colloidal stabilizer (SML-4) were developed. A stable ultra-high-density weighted drilling fluid formulation was thus established, enabling the drilling fluid density to be increased to above 2.85 kg/L and successfully applied in drilling high-pressure intervals, thereby verifying the engineering feasibility of the proposed weighting system [
1].
The insufficient high-temperature resistance and poor contamination tolerance of conventional drilling fluids by using barite to adjust density and developed a high-density saturated brine drilling fluid with a density of 2.60 g/cm
3. This system met the requirements of natural gas drilling operations in three contract blocks in Turkmenistan and Uzbekistan [
2].
To solve wellbore stability problems encountered during drilling through extremely thick salt–gypsum formations and directional drilling within such intervals. Based on barite, ultrafine iron ore powder was selected as an additional weighting material to establish an ultra-high-density supersaturated brine drilling fluid system, with density reaching 2.60 g/cm
3 [
3].
Trace amounts of surface-treated manganese–titanium mineral complexes could replace pure barite or other conventional materials as ideal weighting agents. By using these materials to adjust system density, an ultra-high-density oil-based drilling fluid system (2.35–2.65 g/cm
3) was developed and successfully applied under ultra-high-temperature and high-pressure conditions [
4].
Increased weighting materials to raise mud density from 1.80 g/cm
3 to 2.55 g/cm
3, successfully maintaining drilling-fluid stability under high-temperature and ultra-high-density conditions [
5]. In the same year, Wang Guisong et al. used barite as the weighting material to formulate a high-density kill fluid with a density above 1.60 g/cm
3, demonstrating good field performance [
6].
Excessive solid content increases the likelihood of solid particles invading the formation, causing plugging and contamination. When the solid concentration exceeds the suspension capacity of the liquid phase, particle agglomeration and sedimentation may occur, resulting in density non-uniformity and degraded rheological and flow properties of the kill fluid. Consequently, kill fluid systems have gradually shifted toward solids-free formulations. In such systems, inorganic salt brines, including sodium chloride, calcium chloride, and sodium bromide solutions, are widely used due to their adjustable density and ease of formulation.
Formulated a high-density, low-damage kill fluid with a density of 1.20–1.50 g/cm
3 in the Gudong area of Shengli Oilfield. Calcium chloride was also used as the density regulator, which effectively improved sand control efficiency and single-well productivity [
7].
In addition, some systems regulate solution density through chemical additives. The density adjustment in such systems is continuous and controllable, allowing fine-tuned formulation design by varying additive dosage, thereby enabling better matching with formation pressure windows.
Developed a novel low-damage kill fluid using powdered polyacrylamide as a viscosifier and trivalent organic chromium ions as a crosslinking agent. The fluid density was adjustable within the range of 1.0–2.2 g/cm
3. The gel exhibited minimal formation damage and could be degraded by strong oxidants. Statistical results showed that after applying the gel-type workover fluid, oil well productivity recovered on average 6–8 days earlier, and the initial water cut after operations was reduced by approximately 20% [
8].
In 2013, Li Meng et al. Investigated high-temperature-resistant, high-density kill fluid systems formulated with various chemical additives to address the poor temperature tolerance and severe formation damage associated with commonly used kill fluids in oilfields. The developed system had a density of 1.85 g/cm
3, and its viscosity remained below 110 mPa·s before and after exposure to 242 °C, while the API fluid loss was maintained below 8.0 mL [
9].
To address complex drilling conditions in fragile formations prone to collapse and severe losses, which require timely and effective borehole strengthening. Two borehole-strengthening agents were introduced into the drilling fluid, allowing field-controllable densities in the low-density range of 1.14–1.50 g/cm
3. The developed drilling fluid system successfully reconciled the contradiction between wellbore stability and low pressure-bearing capacity [
10].
Severe losses during workover operations caused by insufficient formation energy and low formation pressure. A composite foam gel workover fluid system was developed using acrylamide copolymers as temporary plugging agents, chromium acetate as a crosslinker, and additional foaming and foam-stabilizing agents. The density of the workover fluid could be adjusted within the range of 1.0–1.25 g/cm
3 [
11].
Optimized high-density oil-based drilling fluid formulations by employing new treating agents and optimizing the particle size distribution of solids to reduce fluid density. Field tests demonstrated that the combination of managed pressure drilling and density reduction during horizontal drilling operations effectively reduced downhole complexity, increased rate of penetration, lowered overall drilling costs [
12].
To more intuitively illustrate the research approach, a flowchart of the investigation process was drawn, as shown in
Figure 1.
As shown in
Figure 1, with increasingly stringent requirements for drilling fluids, higher-performance organic salt solutions are gradually selected as the base fluid. Among them, formate-based organic salts exhibit superior performance compared with conventional inorganic salt brines in terms of density regulation, thermal stability, material compatibility, and environmental friendliness. In addition, formate brines feature low corrosivity, and their clear solutions maintain good rheological properties even at high concentrations [
13]. As a result, formate brines have gradually become an important development direction for solids-free kill fluid systems.
Overall, existing studies mainly rely on conventional experimental approaches for fluid formulation design, which involve considerable subjectivity and uncertainty. Methods such as full-factorial experiments, orthogonal design, and empirical formulation face inherent limitations, including a large number of possible combinations, difficulty in capturing interactions among variables, and insufficient consideration of cost factors. These limitations become particularly pronounced in formate-based systems, where density optimization depends on the formulation of the formate brine base fluid and multiple formate salts are often blended to achieve a wide adjustable density range. While single-solute systems still require extensive preliminary measurements, multisolute systems are characterized by numerous formulation combinations corresponding to a given target density, each associated with different costs. As a result, exhaustive experimental exploration is impractical, leading to high experimental costs and making it difficult for traditional approaches to rapidly identify minimum-cost formulations that satisfy field density requirements.
To address the limitations of conventional experimental approaches, formulation design can be reformulated as a constrained optimization problem oriented toward target values. By integrating density constraints with a cost-based objective function, mathematical optimization methods provide a systematic pathway for identifying economically efficient formulations.
Develop the Pitzer model to overcome the low-concentration limitations of existing electrolyte theories. By extending Debye–Hückel theory with ionic-strength-dependent virial coefficients to account for short-range interactions, the model provides accurate thermodynamic predictions up to several molal (typically about 3–6 mol·kg
−1), while retaining relatively simple equations and good agreement with experimental data for both single and mixed electrolytes [
14].
Develop the electrolyte NRTL (eNRTL) model to address the difficulty of describing excess Gibbs energy over a wide composition range. By decomposing the excess Gibbs energy into long-range electrostatic and short-range local-composition contributions, the model successfully represents systems from molecular liquids to fused salts [
15].
Develop a SAFT-type equation of state to model associating liquids by extending Wertheim’s theory. By decomposing the Helmholtz energy into Lennard–Jones, chain, and association contributions, the model accurately reproduces molecular simulation results and experimental phase equilibrium data for selected pure compounds [
16].
Addressed the limitation that coalbed methane production studies mainly focused on geological factors while neglecting engineering parameters. By considering engineering stages such as drilling and completion, hydraulic fracturing, and production drainage, they applied grey relational analysis and correlation analysis to establish a data-driven mathematical model relating production to engineering parameters. The results were consistent with field performance, demonstrating the feasibility of the proposed approach [
17].
Proposed a predictive method combining intelligent optimization algorithms with neural networks to address the lack of a general density prediction model for drilling fluids under high-temperature and high-pressure downhole conditions. By training the model with existing experimental data, high-accuracy density prediction was achieved. Compared with other intelligent methods, the proposed model demonstrated superior accuracy and stability, showing good engineering adaptability and application potential [
18].
Proposed an adaptive model combining extreme learning machines and the k-nearest neighbors algorithm to predict the density of oil-based drilling fluids under high-temperature and high-pressure conditions. Trained on experimental data, the model achieved fast learning speed, strong generalization ability, and higher prediction accuracy than conventional models, effectively overcoming limitations such as local optima and low training efficiency [
19].
Proposed an improved recurrent neural network algorithm to address issues of low efficiency, insufficient accuracy, and poor adaptability in petroleum engineering data-driven classification. By introducing rectified linear unit activation functions, multilayer network structures, and adaptive optimization strategies, the stability and classification accuracy of model training were enhanced. Experimental results demonstrated that the proposed method achieved high accuracy and good practical value across multiple application domains [
20].
Well-established thermodynamic models, such as the Pitzer model, eNRTL model, and SAFT-type equations of state, provide physically rigorous descriptions of electrolyte and associating fluid systems, but their application requires strictly satisfied assumptions, detailed system-specific interaction parameters, and extensive calibration. In practical formulation problems, these conditions are often difficult to meet, which significantly limits their applicability and feasibility in engineering optimization. Likewise, machine-learning-based models, including neural networks and support vector machines, rely heavily on large, representative, and high-quality datasets to achieve reliable performance. Such data requirements are rarely fulfilled in field-oriented formulation studies, thereby restricting their practical applicability. When the target density is known, accurately predicting and selecting component proportions depends on identifying and quantifying the relationships among variables. Under these constraints, regression analysis offers transparent modeling procedures, clear physical meaning, and high computational efficiency, and is particularly suitable for engineering problems with well-defined variables and limited sample sizes. Therefore, regression-based analysis is adopted in this study as a feasible and effective optimization approach under specified performance and cost constraints.
Addressed the problems of dosage determination and cost optimization of organic salt solutions under different density requirements. Based on laboratory density data obtained for both single-salt and blended systems, nonlinear multivariate least squares and optimization theory were employed to establish mathematical relationships and optimization models between solution density and dosage. The results demonstrated high model accuracy, enabling the determination of optimal dosages for arbitrary target densities, and showing strong engineering applicability [
21].
Conducted laboratory experiments to address cost control while maintaining the performance of oil and gas well working fluids. The relationships between additive dosage and fluid performance under two treatment agents were determined. By combining nonlinear multivariate least squares with optimization theory, mathematical expressions and optimization models were established. Multiple field applications verified the effectiveness of the method in cost control, yielding significant economic benefits [
22].
Addressed the difficulty of optimizing plastic viscosity in microbubble drilling fluids and the limitations of traditional orthogonal and uniform experimental designs, such as uncertain reagent dosage and unsatisfactory results. A multivariate regression experimental approach was adopted, in which 17 formulations were designed and viscosity was measured to establish a quadratic regression model. The results showed that the method achieved high accuracy with prediction errors below 5%, effectively optimizing the plastic viscosity of microbubble drilling fluids and demonstrating practical applicability [
23].
Investigated the rheological properties of oil-based drilling fluids with typical formulations under high-temperature and high-pressure conditions. Through multivariate nonlinear regression analysis of experimental data, mathematical models were established to predict apparent viscosity, plastic viscosity, and yield stress under high-temperature and high-pressure conditions. The predicted values showed good agreement with measured data, with correlation coefficients exceeding 0.98 [
24].
Established mathematical relationships between conventional logging data and permeability, invasion radius, and skin factor based on intermediate parameters such as median grain size, porosity, and resistivity, providing a methodological basis for quantitatively evaluating formation damage using logging data [
25].
Developed a model-based optimization method for drilling fluid density and viscosity to balance wellbore stability, pressure control, and hole-cleaning efficiency. Based on regression models embedded in wellbore hydraulics simulation software, the method can calculate optimal drilling fluid density, viscosity, and related drilling parameters (such as flow rate and rate of penetration) within seconds while satisfying downhole pressure and operational constraints. By defining different objective functions, such as minimum pressure loss, maximum cutting transport efficiency, or cost functions, the method enables flexible optimization for various geological and operational conditions, significantly improving the accuracy and efficiency of density control. This approach is suitable for automated drilling fluid management systems and enhances drilling safety and economic performance [
26].
Investigated the unclear causes of sudden production decline during the trial production period of Well SHB-X in the Shunbei Oilfield. Reservoir sensitivity experiments and drilling fluid compatibility evaluations indicated that velocity sensitivity, salt sensitivity, and drilling fluid damage were insufficient to cause production shutdown. Subsequently, based on 100 days of production data from seven wells in the same block, data-driven methods were applied to quantitatively analyze the relationship between liquid productivity index and 12 influencing factors. The results showed that wax deposition and asphaltene precipitation induced by large well depth and high production rates were the primary plugging mechanisms. Field analysis demonstrated that controlling daily liquid production is key to maintaining stable production in deep carbonate reservoirs [
27].
Conducted experimental studies on different types of drilling fluids and improved the temperature–pressure correction model provided by API standards by introducing a quadratic temperature term to account for the nonlinear influence of temperature on drilling fluid liquid-phase density. Comparative analysis with experimental data showed that the improved model consistently outperformed the API model in density prediction accuracy [
28].
Addressed the difficulty of using conventional skin factor to quantitatively reflect the relationship between drilling fluid properties and formation damage and its limitations in guiding field optimization. A multiparameter formation damage model based on data-driven methods was developed. Using data from nine wells drilled with velvet-type drilling fluids and six neighboring wells, multivariate regression and peeling algorithms identified apparent viscosity, density, dynamic–plastic ratio, and pH value as the main influencing factors, and quantitatively established their relationships with daily production differences. After optimizing the drilling fluid formulation for Well Yan 5-V1, daily gas production increased by nearly 800 m
3. Field applications demonstrated that the method can accurately identify damage factors and provide a solid theoretical basis for drilling fluid optimization, with strong potential for wider application [
29].
Prepared four oil-based drilling fluids with identical components but different densities based on field formulations and investigated the effects of temperature and pressure on drilling fluid density. A binary regression mathematical model incorporating temperature and pressure was established. Field drilling fluids of different densities were used to validate the model, and the results showed strong agreement between predicted and measured values, with an average prediction accuracy of 97.93%, meeting field application requirements [
30].
To precisely guide drilling fluid density adjustment under underbalanced drilling conditions. By analyzing logging response characteristics of gas wells in buried-hill formations, a logging-parameter-based real-time drilling fluid density adjustment method was proposed. From a statistical perspective, the correlation between bottom-hole pressure difference and gas-logging parameters was analyzed, and a drilling fluid density adjustment model for the Caofeidian M structure was established, enabling timely and accurate density control during drilling operations [
31].
In petroleum engineering, regression analysis and data-driven approaches have been widely applied to a range of problems, demonstrating that when system variables are well defined and experimental data are limited, data-driven modeling can provide efficient and reliable solutions for complex engineering problems. Motivated by these successful applications, the present study extends data-driven methodology to the formulation design and density optimization of working fluids by building upon the work of Ref. [
31]. Their study adopted a fundamental modeling framework describing the relationship between solution density and component dosage and, although limited to binary-solute systems and mathematical relationships between solute and solvent dosages, demonstrated the feasibility of data-driven density optimization as an early application of this methodology in formulation design. On this basis, the present work further extends the approach to more complex ternary-solute systems, enabling higher-dimensional formulation optimization, and introduces a cost objective function to explicitly incorporate economic considerations, thereby achieving simultaneous satisfaction of density requirements and cost minimization [
32].
All experimental measurements were conducted at a constant temperature of 20 °C. In this study, temperature was treated as a fixed condition rather than an independent variable, in order to isolate the effects of formulation composition on solution density and cost. As a result, temperature effects were implicitly embedded in the measured data and not explicitly modeled. The proposed framework is therefore applicable to formulation optimization under ambient-temperature conditions, while extension to elevated-temperature environments would require additional temperature-dependent data.
2. Materials and Methods
The implementation of minimum-cost formate brine formulations under density constraints relies on two major components. The first component involves the collection and organization of density-related data for formate solutions, including both literature review and experimental measurements. The experiments encompass the preparation of formate solutions, relevant materials, and laboratory instruments. The second component concerns the development of predictive models, including the construction of a density prediction model for formate solutions and the application of data-driven inverse calculation methods for formulation optimization.
2.1. Data Acquisition of Formate Solutions
The density–concentration relationships of single-solute solutions are well documented in the literature; therefore, density data for sodium, potassium, and cesium formate solutions were obtained directly from published sources, with 10 data sets collected for each solute.
For data unavailable in the literature, four multi-solute formate brine systems were prepared and measured in the laboratory. A total of 40 experiments were designed. The 40 laboratory experiments were designed to provide a quasi-uniform coverage of the feasible formulation space under practical solubility and operability constraints. Experimental points were systematically distributed across single-, binary-, and ternary-solute systems rather than randomly selected. The experimental design, materials, apparatus, and procedures are detailed as follows.
- (1)
Experimental Design
Multi-solute formate brine solutions were prepared in the laboratory. Three binary systems composed of sodium formate (Macklin Biochemical Co., Ltd., Shanghai, China), potassium formate (Macklin Biochemical Co., Ltd., Shanghai, China), and cesium formate (Macklin Biochemical Co., Ltd., Shanghai, China), along with one ternary system, were investigated, giving a total of four solution systems. For each system, 10 experiments were conducted, resulting in 40 experiments overall. The experimental variables were solute type and solute dosage, while temperature, pressure, and humidity were controlled. During preparation, solute and solvent dosages were recorded, and solution densities were measured after completion.
- (2)
Experimental Materials
Four types of formate brine solutions were prepared in the laboratory. The required materials included sodium formate (CHNaO2; AR, 99.5%; MW, 68.01), potassium formate (CHOK; AR, 98%; MW, 84.12), and cesium formate (CHO2Cs·H2O; AR, 97%; MW, 195.94).
- (3)
Experimental Apparatus
The laboratory preparation of the four formate brine solutions required basic apparatus and consumables, including beakers, graduated cylinders, spoons, weighing paper, and glass rods. Additionally, a Lee’s density bottle (standard volume 25 mL) and an analytical balance with 0.0001 g precision were used, as shown in
Figure 2.
As shown in
Figure 2, the pycnometer is of high-precision grade (A) and calibrated as a 25 mL “To Contain” (TC) standard at 20 °C, meaning that the liquid is filled to the graduation line at 20 °C to reach 25 mL. In practice, its actual usage precision may be lower than the nominal standard. The ten-thousandth-gram precision balance, as the name suggests, has a weighing accuracy of 0.0001 g, suitable for measuring samples that require high-precision mass data, such as chemical reagents, pycnometer mass, and solution mass.
- (4)
Experimental Procedure
Measure the density of the formate solution using the pycnometer. First, add 100 mL of water to a beaker, the density of water is 0 Then add the formate powder and stir with a glass rod. If the powder does not fully dissolve, add more water and continue stirring until the powder is completely dissolved, completing the preparation of the formate solution. Rinse the pycnometer (including the glass stopper) with deionized water and dry it. Weigh the empty pycnometer on the ten-thousandth-gram precision balance to obtain mass m0. Fill the pycnometer with deionized water up to the graduation line, slowly pouring in to avoid air bubbles. Insert the glass stopper, wipe the exterior of the pycnometer to remove water, and weigh the filled pycnometer to obtain total mass m1. Pour out the water and dry the pycnometer. Then, pour the prepared formate solution from the beaker into the pycnometer up to the graduation line, slowly to prevent air bubbles. Insert the glass stopper, wipe the exterior to remove any liquid, and weigh the filled pycnometer to obtain total mass m2.
Calculate the solution density using the formula. The formula for the solution density is shown in Equation (1).
2.2. Definition of Solute-to-Solvent Mass Ratio
The density of a solution can be regarded as the ratio of the total mass of all components to the total volume of all components minus the volume contraction. It can be approximately calculated using Equation (2).
In Equation (2), ms is the mass of the solvent, g; mwi is the mass of each solute, g; Vs is the volume of the solvent, mL; x is the quantified value of the volume contraction that occurs when solute and solvent are mixed, mL, Under specific environmental conditions, for a solution with a particular solute and solvent, the solution density depends only on the masses of the solute and solvent; therefore, x is a function of the solute and solvent masses.
By converting the volume of each component in the denominator of Equation (2) into the ratio of mass to density and then dividing both the numerator and denominator by m
s, Equation (3) can be obtained after simplification.
In Equation (3), wi is the density of the solute, s is the density of the solvent, g/cm3; and α = mwi/ms, α is the mass ratio of each solute to the solvent, dimensionless.
The solute-to-solvent mass ratio (SSMR) is defined as the ratio of the solute mass to the solvent mass in a given solution. SSMR is a dimensionless ratio. It should be noted that using the SSMR does not change the actual physical significance of the solute and solvent quantities; it is essentially a measure of how much solute and solvent are present. Moreover, the SSMR is only used to quantify the relative amounts of each component in a specific solution (with fixed solute and solvent types), and its effect on density is meaningful only when comparing within the same type of solution. The solution density is a function of either the SSMR or the solute mass fraction. The following will demonstrate this using SSMR as an example.
Suppose there are two solutions composed of the same solute and solvent, and the SSMR in both solutions is identical. Their densities are 1 and 2, and their solution masses are m1 and m2, respectively. Due to solution homogeneity and mass conservation, the density of the same type of solution remains unchanged after mixing, and similarly, the density remains unchanged if the solution is divided arbitrarily. Therefore, even if the solution masses differ, it is always possible, through a finite number of divisions, to split the two solutions into multiple portions of equal mass. During this division process, both the density and SSMR remain unchanged. Ultimately, two sets of solutions with equal mass, identical SSMR, and densities 1 and 2 are obtained. Comparing any one portion from each set, the two solutions have equal mass and equal SSMR, which implies equal solute and solvent masses. Consequently, the densities of the two solutions must be equal, i.e., 1 = 2.
In the above comparison, the physical meanings of solute mass fraction and SSMR are generally similar. Mathematically, solute mass fraction and SSMR can be converted into each other. The relationship between solute mass fraction and SSMR is shown in Equation (4).
In Equation (4), wi is the solute mass fraction, dimensionless; αi is the solute SSMR, dimensionless.
Solute content in solutions is commonly expressed using mass fraction or molar concentration. In this study, the model is established based on the preparation of 1 m3 of formate brine, for which the system volume is fixed during model development, allowing both descriptors to be theoretically applicable. However, during the prediction stage, solution density is the unknown target variable, and the total mass of the system cannot be determined a priori, rendering solute mass fractions impractical for prediction. In contrast, molar concentration is defined on a per-unit-volume basis and remains independent of solution density, making it applicable even when density is unknown. Nevertheless, in engineering practice, formulation design and cost evaluation are typically based on mass rather than molar quantities. Therefore, SSMR was selected as the independent variable in this study.
2.3. Establishment of Formic Salt Solution Density Model and Reverse Calculation Method for Minimum-Cost Formulation
After obtaining the formate solution data through literature survey and laboratory experiments, the solute-to-solvent mass ratio (SSMR) was selected as the independent variable. A multiple linear regression was performed to fit the SSMRs of solutes and the solution density, resulting in multiple linear equations. Six solution systems—comprising single-solute and two-solute solutions—were separately fitted with multiple linear regression, yielding six multiple linear equations. The purpose was to analyze the correlation between solute SSMRs and solution density and to assess the accuracy of the selected data. Finally, for the entire formate solution system, a total of 70 data sets were subjected to multiple linear regression to establish a density model for the formate solution. The equation is expressed in the form of Equation (5).
In Equation (5), y is the density of the formate solution; xi is the SSMR of the solute, dimensionless; b represents the intercept term; ε denotes the error term, which follows a standard normal distribution.
Although formate brines at high solute-to-solvent mass ratios exhibit pronounced non-ideal electrolyte behavior, including strong ionic interactions and volume contraction, a linear regression model was adopted in this study for engineering-oriented formulation optimization. Within the investigated concentration and density ranges, the experimental and literature data exhibit an approximately monotonic and near-linear relationship between SSMR and solution density, making linear regression an acceptable first-order approximation. Preliminary testing of quadratic and exponential regressions showed only marginal improvements in predictive accuracy while increasing model complexity and reducing interpretability. Physics-based thermodynamic models were not employed because they require extensive interaction parameters and large, high-quality datasets that are currently unavailable for multicomponent formate systems. Given the objective of rapid formulation screening and cost optimization under practical constraints, the linear model provides a transparent and computationally efficient balance between accuracy and engineering applicability.
Based on the formate solution density model, a cost objective function is introduced to realize the reverse calculation of formate solution formulation under density constraints. For a specified solution density, the formulation with the minimum cost is selected through mathematical methods. This problem is similar to a constrained extremum problem, i.e., finding the minimum cost within a certain range. It can be transformed into a linear programming problem for solution. The cost of water is not considered. The mathematical form is shown in Equation (6).
In Equation (6), E represents the total cost and serves as the objective function in the linear programming problem, CNY; ci denotes the unit mass cost of each component, c4 denotes the unit mass price of industrial water, CNY/kg; y is the solution density, which acts as a constraint in the linear programming model, g/cm3; ai and b are the coefficients and constant term of the regression equation related to SSMR; ms is the mass of the solvent, kg. s.t. (y) represents other constraint conditions.
In Equation (6), the parameter m
s is included, where m
s represents the mass of the solvent in the solution. To eliminate the variable solvent mass, m
s needs to be expressed in terms of known parameters, namely the solution density, SSMR values, and the solution volume. Taking a multi-solute solution system as an example, the corresponding expression is given in Equation (7).
In Equation (7), ms denotes the mass of the solvent, kg; y represents the solution density, g/cm3; V is the solution volume, m3. It should be noted that the solution density is a specified target value rather than a variable and therefore serves as a parameter in the model. In engineering practice, the solution volume is usually predetermined prior to solution preparation; hence, V is also treated as a parameter rather than a variable.
Regarding the additional constraints, the solubility of each solute is finite, which imposes upper limits on the corresponding SSMR values. As a consequence, the achievable solution density is also bounded. Therefore, range constraints on x1, x2, x3, and y must be introduced in the optimization model.
4. Discussion
Equation (8) represents the mathematical algorithm underlying the model. Overall, the model computation consists of two main parts. The first part is the establishment of the relationship between the density of multi-solute solutions and the solute SSMRs. The second part is the formulation of a cost objective function to achieve the reverse calculation of the minimum-cost composition of multi-solute solutions at a specified density.
The second part consists of purely algebraic calculations and introduces no error; all errors originate from the first part. These errors mainly arise from experimental operations and sample purity, as well as from the theoretical error of the multivariate linear regression model.
A total of 21 validation experimental points were designed. During the experiments, all formate solutions were prepared in the laboratory according to the predefined schemes, and their densities were measured experimentally as the true density values for each validation point. Meanwhile, the solute SSMRs were substituted into the fitted equations, and the densities calculated by the respective models were regarded as the observed densities. Differences exist between the true densities and the model-predicted densities. When the error falls within the range [0, 0.5], the model is considered excellent; when the error lies in the range (0.5, 1], the model is considered good; and when the error exceeds 1, the model is considered fair. The detailed error data are presented in
Table 10.
As shown in
Table 10, there are discrepancies between the predicted densities obtained from the density prediction model and the measured densities. Calculations indicate that the mean squared error (MSE) of the model predictions is 0.00885, the root mean squared error (RMSE) is 0.094 g/cm
3, and the mean absolute error (MAE) is 0.07 g/cm
3. By normalizing these errors with respect to the average measured density of 1.44 g/cm
3, the relative mean absolute error (rMAE) and the relative root mean squared error (rRMSE) are approximately 4.9% and 6.5%, respectively. Therefore, the model exhibits good predictive stability for most samples. However, the RMSE is higher than the MAE, indicating the presence of relatively large prediction deviations in a small number of samples. Overall, the model satisfies the requirements of engineering accuracy.
From an engineering perspective, fluid density requirements are defined within an allowable density window rather than as a single target value. In practical operations, this window is typically wider than the average prediction deviation observed in this study. The mean absolute error of 0.07 g/cm3 therefore represents a relatively small fraction of the operational density window and can be readily accommodated through routine on-site adjustment. Consequently, the achieved prediction accuracy is considered reasonable and sufficient for formulation guidance and cost optimization within engineering applications.
An analysis of the sources of prediction errors in cesium formate solution systems. For high-density cesium formate brines, nonlinear behavior becomes more pronounced due to strong electrolyte non-ideality. At elevated concentrations, high ionic strength and strong short-range interactions between Cs+ and formate ions lead to significant deviations from ideal solution behavior. In addition, pronounced volume contraction and changes in hydration structure reduce the linear proportionality between solute addition and density increase. These effects are further amplified near solubility limits, where small composition variations can induce disproportionate density changes. Consequently, density–composition relationships in high-density cesium formate systems inherently exhibit nonlinear characteristics.
To provide a more intuitive comparison between the model prediction errors and the error evaluation criteria, the prediction errors were plotted as a grouped scatter diagram, as shown in
Figure 11.
As shown in
Figure 11, when the model is applied to sodium formate solutions, the mean density error is 0.038 g/cm
3; for potassium formate solutions, the mean density error is 0.062 g/cm
3; for cesium formate solutions, the mean density error is 0.135 g/cm
3.
For sodium formate–potassium formate solutions, the mean density error is 0.145 g/cm3; for sodium formate–cesium formate solutions, it is 0.022 g/cm3; for potassium formate–cesium formate solutions, it is 0.042 g/cm3; and for sodium formate–potassium formate–cesium formate solutions, the mean density error is 0.045 g/cm3. Based on the mean error values, the model performance is classified as excellent for sodium formate and cesium formate single-solute solutions, good for potassium formate single-solute solutions, and excellent for sodium formate–potassium formate–cesium formate multisolute solutions.
5. Conclusions
This study focuses on the minimum-cost formulation problem of formate solutions under density constraints. Model development and experimental validation were conducted, and the expected objectives were basically achieved. Based on the literature investigation and laboratory experiments, 10 experimental groups were designed, yielding a total of 70 data sets. On the basis of these data, formate solution density models were established through regression fitting. By introducing a cost objective function and combining a data-driven optimization method, the minimum-cost formulation under density constraints was back-calculated. Accuracy evaluation experiments were further designed. The main conclusions and recommendations are as follows.
- (1)
Theoretical aspect
Based on literature and laboratory data, the solute-to-solvent mass ratio (SSMR) was proposed as the independent variable for model fitting, and it was demonstrated that the density of a given solution depends only on the SSMR. Although strict theoretical derivation is not mandatory in engineering practice, it is still recommended that future studies further investigate particle interaction mechanisms to provide a deeper theoretical explanation of the relationship between solution density and SSMR.
- (2)
Model development
Density models for sodium formate, potassium formate, and cesium formate single-solute solutions, as well as multisolute formate solutions, were established. The fitted equation is y = 0.0898 x1 + 0.1217 x2 + 0.2271 x3 + 1.1982, with an R2 value of 0.91241, which is close to unity. The model can effectively describe the density variation in formate solutions. Only linear models were considered in this study; although the accuracy is excellent, considering the strong ionic interactions and volume contraction effects in high-concentration formate brines, it is necessary to supplement the training dataset and introduce nonlinear models, such as quadratic or exponential regressions, to further improve prediction accuracy.
- (3)
Model accuracy
When applied to seven types of formate solutions, the mean density prediction errors are 0.038 g/cm3, 0.062 g/cm3, 0.135 g/cm3, 0.145 g/cm3, 0.022 g/cm3, 0.042 g/cm3, and 0.045 g/cm3, respectively. The root mean square error (RMSE) is 0.094 g/cm3, and the mean absolute error (MAE) is 0.07 g/cm3. The relative MAE (rMAE) and relative RMSE (rRMSE) are approximately 4.9% and 6.5%, respectively. Although relatively large deviations occur for a small number of samples, the overall results indicate that the model meets engineering accuracy requirements. The larger errors are mainly attributed to limited training data or the local data concentration; therefore, expanding the training dataset is recommended.
Overall, this study establishes a data-driven framework for minimum-cost formulation design of formate solutions under density constraints by introducing the SSMR as the key independent variable and integrating regression modeling with cost optimization. The proposed approach demonstrates satisfactory predictive accuracy and practical applicability at the engineering level. Nevertheless, limitations related to the dataset scale, linear model assumptions, and empirical characterization indicate that further improvements can be achieved by expanding experimental data, incorporating nonlinear modeling techniques, and strengthening physicochemical interpretation. These efforts are expected to enhance model robustness and support more comprehensive formulation optimization in future applications.