1. Introduction
Sensors have become ubiquitous, as they are extensively used in industrial systems, space exploration, research and development facilities, transportation vehicles, robotic systems, household appliances and personal devices. They are the cornerstone behind Wireless Sensor Networks (WSNs), the Internet of Things (IoT), cyber-physical systems and Industry 4.0. As the world continues to get even more connected, sensors are the key element that enables system interconnection with the real physical world. Nearly all systems/devices are equipped with sensors, with many including tens or hundreds of sensors, that measure and monitor different parameters, which are then handled in microcontroller or processor units.
Most sensors transform the physical quantity to be measured into an electrical quantity, typically a current or a voltage, which can then be converted into the digital domain and processed. For the processor unit to estimate the physical quantity, it needs to know the sensor transfer function and therefore it must be properly calibrated before use. This calibration procedure is crucial, as it determines key sensor aspects, including static and dynamic parameters, and the overall sensor quality is evaluated based on these characteristics. The usefulness of sensors is dependent on parameters such as sensitivity, responsiveness, non-linearity, drift, hysteresis and bias among others [
1]. It is also common and useful for the measurement system to include, in its measuring chain, some analog signal conditioning before digitalization. In this case, the calibration procedure can be applied just to the sensor or to the complete measuring chain. Within the context of this work, the term sensor calibration is used, but it can, without loss of generalization, be applied to the complete measuring system.
The VIM [
2] defines, in paragraph 4.31, the calibration curve as the “
expression of the relation between indication and corresponding measured quantity value” where “
measured quantity value” can be replaced by “
measured value”. Sensors included in measurement systems require the determination of this calibration curve. The calibration process consists of supplying the sensor with a set of inputs and measuring the corresponding outputs to obtain, for an appropriate model, the calibration curve. This process is necessary to minimize systematic measurement errors and to enable the estimation of the physical quantity that the sensor is measuring. The determination of the sensor calibration curve is essential and its optimization is fundamental in metrology.
In [
3], a set of methodologies to calibrate sensors with statistical estimation algorithms is presented and examples are included. The process is divided into three steps: (i) description of the sensing process with a statistical model; (ii) estimation of the model parameters; (iii) processing new measurements to improve the statistics of the estimated parameters.
The optimum selection of the measurement points for the calibration process was analyzed in [
4] for linear, quadratic, cubic and fourth-order input/output sensor relationships. Therein, it is concluded that, for zero-mean normally distributed variables with the same standard deviation: (i) the optimum location of the sensor inputs does not correspond to equally spaced values in the input range; (ii) an increase in the number of measurements reduces the calibration coefficients uncertainty; (iii) the standard deviations of the calibration parameters decrease with increased repetitions of the same sensor inputs, but there is no improvement by indefinitely continuing to add repetitions.
A discussion on the experimental design techniques to optimize the calibration process was published in [
5], where a relative humidity sensor and a heat meter were considered. The process is based on the since-deprecated error propagation method suggested in the 1993 version of the Guide for the Expression of Uncertainty (GUM) [
6], which has been replaced with the law of propagation of uncertainty in updated versions of the GUM [
7].
An A-optimality criterion is used in [
8] to select the ideal measurement points for calibration. The method applicability is demonstrated for the calibration of a differential pressure gauge using a classical Least Squares (LS) method for a second-order polynomial input/output relation. The considered measurement errors are zero-mean normal/Gaussian distributed.
An example based on relative humidity (RH) resistive and capacitive sensors using different saturated salt solutions is used to determine the ideal measurement/calibration points in [
9]. Uncertainty propagation is used to determine the relative humidity uncertainty. Different order polynomials were evaluated using the residual plots together with the
t-value of the highest-order parameters.
In the situations where the sensor input/output relation is a straight-line, a first-order linear regression is commonly used to estimate the offset (or intercept) and slope parameters that define the straight-line analytical equation. Least squares, which was initially used for the estimation of the orbits of comets, was independently developed by Legendre [
10] and Gauss in the early 1800’s. It is widely used in almost all research fields and taught to high-school and undergraduate students. The method minimizes the sum of the squares of the residuals between the measured outputs and the regression estimates. For sensor calibration, the LS method requires a set of sensor input and output values. The parametric values of the straight-line that best fits the set of measured values given the known sensor inputs are estimated by the method. Straight-line linear regression can also be used to determine the linearity or nonlinearity of a sensor, system, or device. However, the method neglects the measurement uncertainties of the input quantities and, in real calibration scenarios, both the input and output quantities have measurement uncertainties.
The technical specification ISO/TS 28037 [
11] describes different straight-line calibration situations. It specifies how to estimate the slope and offset uncertainties as defined in the GUM. However, as explained in the note of paragraph 5.5.1, when the uncertainty propagation approximation is invalid, the propagation of distributions based on the Monte Carlo Method (MCM) should be used. The process of applying the MCM is beyond the scope of the technical specification, and therefore is not covered in the ISO/TS 28037 technical specification. A detailed analysis, presented in [
12], suggests adjustments to the technical specification as, in some cases, it underestimates the calibration coefficient uncertainties. Although it only considers simulation situations and multivariate normal/Gaussian distributions, it also concludes that other assumptions will result in different outcomes and should be evaluated on a case-by-case basis. To cover all cases, it is stated that MCM should be used and its results are the reference for all situations.
MCM was used for straight-line linear calibration in the field of dimensional metrology of a 1-D measuring machine in [
13]. The measurement uncertainty was considered only for the dependent variable as the input uncertainties were disregarded. The Probability Density Function (PDF) of the measured values was considered to be Gaussian/normal and its uncertainty was the same for all the measurement values, i.e., the homoscedastic situation. The MCM estimated the slope and offset average values, standard deviations and their covariance.
This paper discusses in detail the use of MCM for estimation of the straight-line sensor calibration coefficients. It is capable of addressing any situation, including different PDFs for any of the measurement values (system/sensor input or output). It also accounts for covariances between measurements. Furthermore, it effectively handles heteroscedastic conditions that arise intrinsically from the measurements, for example due to the use of different instrument ranges leading to different uncertainty levels. This work is an extension of the preliminary study presented in [
14], with a new experimental measurement setup and more in-depth discussion/analysis. Also added is the estimation of the output correlation in two extreme cases: one with only two measurements, and the other with a very large number of measurements. Additional discussion and analysis is included for the comparison of the MCM results with the ones obtained from the weighted total least squares algorithm that can also take into account sensor/system input uncertainties.
Since the time and cost of a calibration procedure depends on the number of measured calibration points, an MCM simulation-based strategy for the choice of the location of these points is proposed. This strategy avoids measuring points at non-ideal locations, and allows the optimization of the uncertainty of the estimated sensor parameters. The work developed in this paper contributes to filling the gaps in the ISO/TS 28037 technical specification by presenting and analyzing the performance of MCM in realistic straight-line sensor calibration scenarios. It considers different uncertainty PDFs for each input and output measurement values, which are not covered in the technical specification.
In
Section 2, different straight-line sensor calibration algorithms are described together with their strengths and weaknesses, and the application of the Monte Carlo method is also presented.
Section 3 describes the measurement setup used to retrieve the calibration measurements of a power grid voltage sensor-box, which are used throughout the remainder of this paper. The results obtained in different conditions are presented in
Section 4 with discussion and analysis of the multiple situations considered.
Section 5 proposes the selection strategy for the calibration measurement points. The conclusions are summarized in
Section 6.
2. Straight-Line Sensor Calibration
To calibrate a sensor, a finite set (
N) of measurement pairs are performed. Each pair includes a sensor input measurement,
, and a corresponding output measurement,
:
For each of the measurements there is a corresponding uncertainty
for the sensor input measurement, and
for the sensor output measurement.
The determination of the sensor input/output relation is done using an algorithm that, based on the set of measured pairs, determines the coefficients of the equation/model that describes the sensor input/output relation.
Curve fitting regressions are divided into linear and nonlinear regressions. In linear regression models [
15], which include, for example, the basic straight-line and polynomial fitting, the sensor output is modelled by the sum of a set of
K finite terms whose coefficients
are multiplied by known functions
of the input
x:
When a linear regression is unable to model the input/output relation as defined in (
2), a nonlinear regression is required—this case is not considered here but the Monte Carlo method could also be used in those situations.
In this paper, the input/output relation of the sensor is considered to be modeled by a straight-line (in some of the literature, called simple linear regression):
where
m and
b are the straight-line slope and offset, respectively. This is a particular case of (
2) obtained with
,
,
,
and
. The objective of the sensor calibration procedure is the determination of the coefficients
m and
b that best describe the sensor input/output relation.
In Least Squares (LS) linear regression algorithms, the sum of the squared errors between the sensor outputs and the straight-line values at the corresponding input values
is minimized. A simplified graphical representation of this method is shown in
Figure 1. Since the errors/distances are measured vertically, this method does not take into account uncertainties of the measured sensor inputs, i.e., for all measurements it assumes
and the only uncertainties considered are those of the sensor output measurements
.
The Total Least Squares (TLS) method, which is also called orthogonal linear regression, minimizes the sum of the perpendicular distances between the measured points and the straight-line [
16], as represented in
Figure 2A. It takes into account uncertainties in the sensor input measurements. However, this method is also limited, as it considers that all measurements, either sensor inputs or outputs, have the same uncertainty values—the PDFs represented in
Figure 2A near
and
all have the same standard deviation.
Also capable of considering uncertainty in the sensor input measurements, the Deming regression [
17] requires that the measurement uncertainties in both the input and output quantities are independent, have a normal distribution, are homoscedastic, and that there is a known common ratio of their variances as depicted in
Figure 2B, i.e.,
, where
is the relation between sensor input and output measurement uncertainties. However, due to the unique uncertainties of each measurement, which are usually dependent on the measured value and on the used instrument range, these conditions are often unverified.
To consider different uncertainties for each
and
without restrictions, an option is the Weighted Total Least Squares (WTLS) regression described in paragraph 7 of [
11]. This method is also known as the weighted Orthogonal Distance Regression (ODR), Generalized Distance Regression (GDR), or errors-in-variables model. It assigns different weights,
and
, to the orthogonal distances in their contribution to the cost function, which is iteratively minimized. The weight given to each measurement is the inverse of their standard uncertainty,
and
. With these weights, measurements with lower uncertainty are given more relevance, i.e., its orthogonal distance to the straight-line is more penalized. One limitation of this method is that it is unable to consider covariances from the measured data. One method that can include these covariances, is the Generalized Gauss Markov Regression (GGMR), described in paragraph 10 of [
11]. However, as the other methods described in [
11], its uncertainty estimation is based on the propagation of uncertainties.
To consider differently shaped PDFs and the covariances of the measurements (both inputs and outputs), the propagation of the distributions based on the Monte Carlo method (MCM) [
18,
19] must be used. This method randomly generates values of the sensor input and output based on the measurements according to their PDFs/covariances. It then estimates the model outputs, which in this case are the regression parameters slope
m and offset
b obtained using WTLS, and the process is repeated
M times. Measurement covariances can be taken into account when generating the random values associated with the measurements. A statistical analysis of the estimated parameters can be used to assess their PDF and CDF (Cumulative Distribution Function) shapes, average values, standard deviations, correlation and uncertainties. The number of repetitions should be adjusted to achieve convergence, as described in paragraph 7.2 of [
18]. A significant advantage of the MCM approach is that it allows the inclusion of heteroscedasticity situations, as is the case when the individual measurement uncertainties are different, as they depend on the actual measured values and instrument ranges, which is common in most instruments. In addition, it can consider any distribution (PDF) associated with the measurement uncertainties, as exemplified in
Figure 3.
3. Experimental Setup
The experimental setup is represented in
Figure 4. The sensor-box is based on an LEM LV 25-P Hall-effect voltage sensor (LEM, Geneva, Switzerland) [
20], configured for operation in the European nominal 230 V RMS power grid voltage. The sensor input voltage is supplied by a Wavetek 9100 calibrator (Wavetek, Hsinchu Science Park, Taiwan) [
21] controlled by IEEE 488.2 that sequentially and automatically changes the sensor-box input voltage. To avoid interferences with the 50 Hz power grid, the AC source frequency is set to 60 Hz. The sensor-box input voltage is measured with an Agilent 34410A (Agilent Technologies, San Diego, CA, USA) [
22] 6.5-digital multimeter operating as voltmeter—this setup enables the use of the measurements from the multimeter or from the calibrator to assess the uncertainties of the sensor-box input. An Agilent 3458A 8.5-digit multimeter (Keysight Technologies Inc., Santa Rosa, CA, USA) [
23], also operating as voltmeter, measures the sensor-box output voltage. Both multimeters are also controlled by IEEE 488.2 to collect their AC measurements.
In accordance with the GUM [
7] recommendations for Type B uncertainties derived from instrument specifications, the probability density functions (PDFs) of the instruments’ measured values are considered to be uniform, centered on the measured values, and the uniform PDF width is twice the maximum error obtained from the instrument specifications [
21,
22,
23].
A set of sensor input/output measurements were obtained, with input RMS voltage in the 4 V to 300 V range, to estimate the sensor-box straight-line parameters. For each measured value the maximum error was determined using the specifications of the instruments [
21,
22,
23]—notice that the maximum errors (and uncertainties), as specified by the manufacturers, depend on the ranges used for each measured value.
The complete set of measured data and the corresponding maximum errors are available as detailed in the Data Availability Statement.
5. Selection Strategy of Calibration Points
One important aspect that must be considered when performing sensor calibration is the selection of which sensor input values to measure. The relevance of this decision is derived by three issues related with increasing the number of measured values: (i) it usually results in higher costs—for example, a liquid viscosity sensor [
24] requires the availability of multiple costly solutions with different viscosities; (ii) it requires a lengthier calibration process—in the voltage sensor-box example, increasing the number of input voltages requires changing the calibrator voltage, waiting a specified amount of time to stabilize the sensor input and output voltages before executing the measurements; (iii) finally, what is the desired uncertainty of the calibration coefficients. Within this context it is of crucial importance to minimize the number of measured values in the calibration process. This section aims to discuss how to choose the sensor input values to optimize the regression slope estimation uncertainty.
The first analysis includes using only two pairs of measured values, which corresponds to the minimum number of pairs that can be used to estimate the sensor calibration coefficients. The results presented in
Figure 12 correspond to the relative uncertainty of the slope regression estimates, with the red line corresponding to the situation where the last measured pair (with
) is used and the influence of the second pair input voltage is analyzed. As expected, the lowest uncertainty is obtained when the second value is furthest apart, i.e., when the second value is the first measured pair (with
). Notice the very large uncertainty (nearly 10 %) when the second value is near the last pair. In this case, the two pairs used in the regression are too close and this results in a very large uncertainty in the estimated slope. There is a small discontinuity near the pair with
due to a change in the Wavetek 9100 calibrator range, which provides the sensor-box input voltage,
(it switches from the 105 V to the 320 V range).
The results presented in the blue line of
Figure 12 correspond to the situation where the first measured pair is always used (
) and the second pair used in the regression is changed. In this case, there are four discontinuities near
,
,
and
. These discontinuities are caused by the different maximum errors associated with the ranges of the two instruments. The first (near
) and last (near
) discontinuities are caused by the Agilent 3458A (which is measuring the sensor-box output) changing ranges from
to
and then from
to
. The others are caused by the range change from the Wavetek 9100 calibrator (which corresponds to the sensor-box input) from
to
and from
to
. It should be noted that these discontinuities are also present in the red line of
Figure 12 but are not as noticeable. The last noteworthy aspect of these results is that the lowest slope uncertainty is not obtained using the two most extreme measured pairs of values but instead using
and
—which is immediately before the Wavetek calibrator switches from the 105 V to the 320 V range. This is because it has a lower sensor-box input relative uncertainty than the
value (approximately 0.046 % for
compared with 0.056 % for
).
The number of calibration points and their selection can be optimized using the MCM method by simulating, before measuring, the impact of possible new input values on the uncertainty of the estimated parameters. A possible optimization process is proposed and detailed in the flowchart presented in
Figure 13. The process starts with measuring the two most extreme input values and determining, from these measurements, their uncertainties. The next step is to estimate the slope and offset calibration coefficients and their uncertainties using the MCM. At this stage, the slope uncertainty is compared with the desired uncertainty for this parameter—if the current value has reached the desired target, the calibration is stopped (notice that this description is aimed solely for the slope parameter but can be adapted to include the offset as well). If the target uncertainty has not been reached, the process will evaluate each possible input value (added to the already performed measurements) simulating the output values using the updated slope and offset parameters. From the set of simulated input values, the one that produced the smallest slope uncertainty is selected to be measured. The process is then repeated with the inclusion of the new measurement pair. It should be noted that a limit on the number of new measurements can be added to this iterative process as well as a condition to stop if the slope uncertainty does not improve significantly.
The green line in
Figure 14 shows the results obtained with the optimization process. The non-optimized selection of evenly spaced input values is presented in the blue line (this corresponds to the blue line from
Figure 8). The results when using only two measurements are the same but as the number of measurements increases, the uncertainties obtained with the optimization process are lower, thus highlighting the advantage of the optimization process. A notable case is the selection of the third measurement to be performed. While the equally spaced process uses 152 V (half-way between 4 V and 300 V), the proposed optimization selects 104.975 V, which is the value immediately before the calibrator switches ranges, as described in the analysis of the blue line in
Figure 12. Although the differences may seem only incremental, they can, in some cases, present significant improvements. For example, to achieve a 0.01 % slope relative uncertainty target, the equal spacing option (blue line) requires 60 measurement pairs, while the proposed optimization process (green line) requires only 37.
6. Conclusions
This paper presents an analysis of sensor calibration using the Monte Carlo Method (MCM), specifically for scenarios where the sensor input/output relation is modeled by a straight-line relationship. Traditional regression techniques include Least Squares (LS), Total Least Squares (TLS), Deming, and Weighted Total Least Squares (WTLS). These methods have limitations in handling measurement uncertainties, especially when they include sensor input and output measurements with non-Gaussian distributions and/or covariances. The Monte Carlo-based approach provides a flexible and robust alternative that accounts for heteroscedastic measurement uncertainties, non-Gaussian distributions, and PDF propagation including input covariances.
Through an experimental setup using a Hall-effect voltage sensor-box, the method was validated and shown to yield accurate estimates of regression parameters along with statistically sound intervals with a 95 % level of confidence, as shown in
Figure 9. The presented results show that (i) MCM allows the use of any measurement PDFs for both input and output variables; (ii) for the presented example, increasing the number of calibration points reduces the estimation uncertainties, with the relative uncertainty in the slope decreasing approximately with the square root of the number of measured points; (iii) the location and uncertainty characteristics of selected measurement points significantly affects the uncertainty of the calibration coefficients, particularly when using a limited number of points; (iv) the method can be used to assess the correlation of the estimated calibration coefficients, making it ideal for sensor systems operating under diverse real conditions. Overall, the Monte Carlo method is a versatile tool for sensor calibration, particularly in applications that require rigorous uncertainty evaluation. This approach ensures compliance with current metrological guidelines and can be used in a wide range of sensor types and measurement conditions.
An optimization process to select the ideal input measurements is proposed. It uses MCM on simulated calibration points to evaluate which input value should be measured to minimize the slope uncertainty and it is used until a target slope uncertainty is reached. In the described example, for a target relative uncertainty of in the slope parameter, the optimized process required 37 measurement pairs, while using equally spaced measurements needed 60. This result demonstrates that the process can be used to improve cost and time efficiency in sensor calibration.
While the presented example did not incorporate sensor input covariances, which could affect calibration accuracy in systems where such correlations are significant, the Monte Carlo method is capable of handling correlated input measurements.
Future work will focus on extending the Monte Carlo-based calibration approach to nonlinear sensor models and evaluating its performance across a broader range of sensor types and uncertainty distributions.