Next Article in Journal
Nicotine in Fine Particles in Shanghai: Temporal Variations and Influencing Factors
Next Article in Special Issue
Indoor Air Filtration System Performance: Evidence from a Two-Week Office Study Within the EDIAQI Project
Previous Article in Journal
Seasonal Variability in the Particulate Matter Removal Efficiency of Different Urban Plant Communities: A Case Study
Previous Article in Special Issue
Assessment of Sensor Data from an Air Quality Monitoring Network—The Need for Machine Learning-Based Recalibration and Its Relevance in Health Impact Analysis of Local Pollution Events
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Machine Learning Calibration Transfer for Low-Cost Air Quality Sensors: Distance-Based Uncertainty Quantification in a Hybrid Urban Monitoring Network

1
Institute of Information and Communication Technologies, Bulgarian Academy of Sciences, 1113 Sofia, Bulgaria
2
Centre of Excellence in Informatics and Information and Communication Technologies, 1113 Sofia, Bulgaria
*
Author to whom correspondence should be addressed.
Atmosphere 2026, 17(4), 335; https://doi.org/10.3390/atmos17040335
Submission received: 25 February 2026 / Revised: 20 March 2026 / Accepted: 24 March 2026 / Published: 26 March 2026
(This article belongs to the Special Issue Emerging Technologies for Observation of Air Pollution (2nd Edition))

Abstract

Low-cost air quality sensors enable dense urban monitoring networks but require calibration against reference-grade instruments. While machine learning calibration is well-established for co-located sensor pairs, applying these calibrations to sensors deployed far from any reference station—the operational reality for most network sensors—lacks systematic methodology. We address this gap using 24 months of hourly data (August 2023–July 2025) from Sofia, Bulgaria, where five official reference stations (Executive Environmental Agency) operate alongside 22 AirThings low-cost sensors, four of which are co-located. Random Forest models achieved R 2 ( 0.53 , 0.75 ) across PM2.5, PM10, NO2, and O3, representing from 40% (for O3) to 408% (for PM2.5) improvement over Multiple Linear Regression baselines. Using leave-one-station-out spatial cross-validation, we derived pollutant-specific uncertainty growth rates ( α ) from 3.84% to 5.62% per km, characterizing how calibration uncertainty increases with distance from reference stations (statistically significant for PM10 and O3, p < 0.05 ). Applied to 18 non-co-located sensors, the framework generated 1.2 million calibrated hourly measurements with 95% prediction intervals over the study period. Co-location sites spaced 6 km apart achieve a less than 30% uncertainty increase at network midpoints, within EU Air Quality Directive thresholds for indicative monitoring. These empirically derived α parameters enable network planners to predict measurement reliability at arbitrary sensor locations without ground-truth validation, providing evidence-based guidance for cost-effective hybrid monitoring network design.

Graphical Abstract

1. Introduction

Urban air quality monitoring has become a pressing public health priority as cities worldwide grapple with the impacts of pollution on respiratory health, cardiovascular disease, and overall environmental quality [1,2,3]. Traditional monitoring approaches rely on official reference stations that provide highly accurate measurements but face a fundamental constraint: their high cost limits spatial coverage, leaving vast urban areas unmonitored. This sparse coverage makes it difficult to capture the fine-scale spatial variability of air pollution that exists within cities, where concentrations can vary significantly across neighborhoods due to local traffic patterns, industrial sources, and geographic features. Recent advances in sensor technology have introduced low-cost air quality sensors as a potential solution [4], offering the promise of dense monitoring networks at a fraction of the cost of traditional stations [5,6]. However, these sensors present their own challenges, particularly around calibration accuracy and data quality assurance. We analyze 24 months of hourly data (August 2023–July 2025), deliberately spanning two complete winter heating seasons to capture inter-annual variability and the full range of pollution conditions typical in Central European cities. This extended temporal coverage, combined with exceptional sensor data completeness (96.9% average data availability), provides a robust foundation for developing and validating calibration transfer methodologies.
The core challenge lies in integrating two fundamentally different monitoring systems: official reference stations that deliver traceable, quality-assured measurements but with limited spatial coverage, and low-cost sensor networks that enable dense monitoring but suffer from calibration uncertainties and sensor drift. While official stations operated by environmental agencies provide the gold standard for air quality data and are used for regulatory compliance, their expansion is financially prohibitive for most municipalities. Low-cost sensors, on the other hand, can be deployed in large numbers to reveal spatial pollution patterns, but their raw measurements often deviate significantly from reference-grade instruments due to cross-sensitivities, temperature dependencies, and gradual calibration drift. This creates a dilemma for cities seeking to expand their monitoring capabilities: invest heavily in a few additional reference stations, or deploy many low-cost sensors with uncertain data quality. This study addresses this challenge using data from Sofia, Bulgaria, where a unique opportunity exists to examine both systems operating in parallel–5 official reference stations operated by the Executive Environmental Agency (ExEA), located in the Pavlovo, Hipodruma, Nadezhda, Mladost, and Druzhba neighborhoods, alongside 22 AirThings low-cost sensors deployed throughout the city, with 4 sensors co-located within meters of official stations. A unique characteristic of this dataset is that while all reference stations measure PM10, NO2, and O3, only one station (Hipodruma) provides PM2.5 measurements. This creates an opportunity to examine calibration transfer under varying data availability constraints, with PM2.5 calibration relying on a single co-located pair while other pollutants benefit from four independent training locations.
Three critical gaps prevent the effective integration of official and low-cost monitoring systems into validated hybrid networks. First, while co-located sensors can be calibrated against nearby reference stations, the vast majority of low-cost sensors are deployed in locations without proximate reference measurements, creating a calibration transfer problem that has not been systematically addressed in the literature. Second, operational frameworks that guide municipalities on how to strategically deploy and maintain hybrid networks are largely absent, leaving cities without clear implementation pathways. Third, quality assurance protocols for low-cost sensor data remain underdeveloped, particularly regarding uncertainty quantification and transparent communication of data reliability to policy makers and the public.
This study introduces a machine learning-based calibration transfer framework for creating validated hybrid air quality monitoring networks that combine the accuracy of official reference stations with the spatial density of low-cost sensor arrays. Our approach develops calibration models using co-located sensor pairs and then systematically transfers these calibrations to non-co-located sensors through transfer learning techniques that account for spatial distance, meteorological conditions, and measurement uncertainty. We implement and compare two machine learning calibration methods–Multiple Linear Regression as a baseline and Random Forest as an ensemble approach–achieving R 2 ( 0.53 , 0.75 ) across pollutants for operational calibration transfer. A key innovation is our distance-based uncertainty quantification framework that predicts calibration uncertainty as a function of spatial separation from reference stations (uncertainty growth rates: 3.8–5.6% per km; statistically significant for PM10 and O3, with consistent trends for PM2.5 and NO2), enabling practitioners to estimate measurement reliability for sensors deployed without ground-truth validation. This framework addresses a longstanding question in the sensor community: How sparse can co-location infrastructure be while maintaining acceptable calibration quality? Our analysis demonstrates that co-location sites spaced 6 km apart achieve <30% uncertainty increase at network midpoints, enabling cost-effective city-wide monitoring while maintaining transparent quality assurance.
The remainder of this paper is organized as follows. Section 2 reviews the relevant literature on low-cost air quality sensors, machine learning calibration methods, transfer learning applications, hybrid monitoring networks, and quality assurance frameworks. Section 3 describes our methodology, including the Sofia case study setup, data preprocessing, calibration model development, transfer learning implementation, and quality flagging system design. Section 4 presents results from calibration model evaluation, transfer learning performance, and network optimization analysis. Section 5 discusses the implications for urban air quality monitoring policy and practice, addresses limitations, and proposes future research directions. Section 6 concludes with practical recommendations for cities seeking to implement hybrid monitoring networks.

2. Literature Review

2.1. Low-Cost Air Quality Sensors

The past decade has witnessed rapid development in low-cost air quality sensing technology, driven by advances in miniaturized sensors, wireless communication, and affordable computing platforms [5]. These sensors, typically costing between $100 and $1000, measure particulate matter (PM2.5, PM10), gaseous pollutants (NO2, O3, CO), and meteorological variables using optical, electrochemical, or metal-oxide semiconductor technologies. Field deployments have demonstrated that low-cost sensors can capture meaningful spatial and temporal pollution patterns, revealing hotspots and gradients that official monitoring networks miss entirely due to their sparse coverage [6]. However, these sensors face well-documented limitations that complicate their use in scientific and regulatory applications. Cross-sensitivities to humidity, temperature, and interfering gases can introduce significant measurement errors. Sensor drift over time degrades calibration accuracy, with performance typically declining within months of deployment. Inter-sensor variability means that even sensors from the same manufacturer and model batch can produce different readings under identical conditions.
Validation studies comparing low-cost sensors against reference instruments have yielded mixed results, with correlation coefficients ranging from strong agreement ( R 2 > 0.8 ) under controlled conditions to poor performance ( R 2 < 0.5 ) in challenging field environments. The literature reveals that raw sensor measurements often require substantial calibration corrections before they can be considered reliable. Various calibration approaches have been proposed, from simple linear corrections to more sophisticated methods accounting for environmental cofactors [7], but most studies focus exclusively on co-located calibration scenarios where sensors sit directly next to reference stations. This leaves a critical gap: how to calibrate the majority of sensors deployed in locations far from any reference station.

2.2. Machine Learning Calibration Methods

Machine learning has emerged as a powerful tool for improving low-cost sensor accuracy by learning complex, nonlinear relationships between sensor readings, environmental conditions, and true pollutant concentrations [8]. Multiple Linear Regression (MLR) serves as the baseline approach in many studies, using sensor outputs and meteorological variables as predictors of reference measurements [9]. While interpretable and computationally efficient, MLR struggles to capture nonlinear sensor responses and complex interactions between variables. Random Forest models have shown superior performance in numerous calibration studies [10], effectively handling nonlinearities, variable interactions, and outliers without requiring extensive feature engineering. These ensemble methods consistently outperform linear models, achieving correlation improvements of 0.1–0.3 in R 2 values [11].
Neural networks and deep learning approaches represent the cutting edge of sensor calibration research, with architectures ranging from simple feedforward networks to more complex convolutional and recurrent designs [12]. These models can automatically extract relevant features from high-dimensional input data and learn hierarchical representations of sensor behavior [13]. Long Short-Term Memory (LSTM) networks have recently been applied to sensor calibration with the specific goal of modeling temporal dependencies and capturing sensor drift dynamics. LSTMs can learn how sensor accuracy degrades over time and potentially provide adaptive calibration that adjusts as sensors age. However, these advanced methods require substantial training data and computational resources, and their “black box” nature complicates interpretation and diagnostic analysis when calibrations fail.
Systematic algorithm comparisons confirm this hierarchy: gradient tree boosting, AdaBoost, and neural networks all outperform linear baselines in controlled benchmarks, but no single method consistently dominates across sensor types and environments [14]. Recent studies specifically targeting calibration propagation (applying a model trained at one node to geographically separated nodes) report that RF and gradient boosting maintain acceptable accuracy when source and target sites share similar microenvironments, while LSTM networks offer additional gains for pollutants with strong temporal structure [15,16].
Prior work has specifically evaluated supervised and unsupervised machine learning approaches for calibrating low-cost air quality sensors in urban deployments, consistently finding ensemble methods superior to linear baselines [17]. Despite this methodological diversity, a critical limitation persists across the literature: nearly all machine learning calibration studies train and validate models exclusively on co-located sensor-reference pairs. The question of how well these models transfer to sensors deployed elsewhere—spatially separated from training locations–remains largely unexplored. This represents a fundamental gap because the practical value of low-cost sensor networks depends precisely on deploying sensors where reference stations do not exist.
The study in [18] proposes machine learning calibration methods which use a high-cost instrument as a reference to improve the accuracy of networks of low-cost sensors. The authors combine three models for machine learning: linear regression, random forest, and Gradient Boosting Regression. In [19], the integration of machine learning and neural networks for the calibration of a low-cost nitrogen dioxide sensor is proposed. Tastan [20] developed the Internet of Things (IoT)-based air quality monitoring system. The system is tested by applying several machine learning algorithms as follows: Decision Tree (DT), Linear Regression (LR), Random Forest (RF), k-Nearest Neighbors (kNN), AdaBoost (AB), Gradient Boosting (GB), Support Vector Machines (SVM), and Stochastic Gradient Descent (SGD). In [21], the authors apply Linear Regression (LR), Random Forest (RF), Gradient Boosting (GB), k-nearest Neighbors (KNN), and Neural Networks (NN) for calibration of low-cost sensors for monitoring PM2.5. They apply k-fold cross-validation methods to ensure the robustness of the performance of the model.

2.3. Transfer Learning in Sensor Networks

Transfer learning, the process of applying knowledge gained from one task or domain to a different but related task or domain [22,23], has seen limited application in air quality sensor calibration despite its potential relevance. The broader environmental monitoring literature includes examples of spatial transfer, where models trained at one location are applied to another, and temporal transfer, where models trained during one time period are applied to later periods. These studies demonstrate that transfer learning can work reasonably well when source and target domains share similar characteristics, but often requires domain adaptation techniques to account for systematic differences.
In the air quality sensor context, a few recent studies have begun exploring calibration transfer [8,24], typically by training models on co-located sensors and applying them to nearby non-co-located sensors, sometimes with correction factors based on meteorological similarity or spatial proximity. However, these efforts remain exploratory and lack systematic frameworks for quantifying transfer uncertainty, determining when transfer is appropriate, and adapting transferred calibrations to local conditions. No comprehensive methodology exists for implementing calibration transfer across an entire sensor network with varying distances from reference stations, diverse microclimatic conditions, and heterogeneous sensor characteristics. This gap is particularly problematic for operational deployments where most sensors will necessarily be far from reference stations.

2.4. Hybrid Monitoring Networks

The concept of hybrid monitoring networks—integrating official reference stations with low-cost sensor arrays—has gained traction as a pragmatic approach to expanding spatial coverage while maintaining data quality [25,26]. Several pilot deployments in cities worldwide have demonstrated the feasibility of operating both systems in parallel, with reference stations providing calibration anchors and validation data while low-cost sensors fill in spatial gaps. Co-location studies, where low-cost sensors are deliberately placed alongside reference instruments, have become standard practice for developing calibration equations and assessing sensor performance under real-world conditions [27].
However, the literature on hybrid networks reveals a significant gap in operational guidance. While researchers have successfully demonstrated hybrid network concepts in controlled deployments, cities lack clear frameworks for making strategic decisions about network design: how many reference stations are needed as calibration anchors, where to place them for maximum network benefit, how to optimize low-cost sensor placement to maximize coverage while ensuring calibratability, and how frequently sensors require recalibration. Cost-benefit analyses comparing hybrid network expansion strategies against traditional reference-station-only expansion are largely absent from the literature, making it difficult for municipalities to justify investments in hybrid approaches to policymakers and funding agencies.

2.5. Quality Assurance and Uncertainty Quantification

Quality assurance (QA) for low-cost sensor data presents unique challenges distinct from traditional reference monitoring [28]. While reference stations operate under strict QA/QC protocols with regular calibration checks, zero/span tests, and performance audits, comparable protocols for low-cost sensor networks remain underdeveloped. The literature includes various proposals for automated QA procedures: statistical outlier detection, temporal consistency checks, spatial coherence analysis, and comparison against co-located reference data when available [29]. However, these methods typically flag potentially problematic data without providing quantitative uncertainty estimates that would allow users to assess fitness for specific purposes.
Uncertainty quantification–providing confidence intervals or prediction intervals for calibrated measurements–is essential for regulatory and policy applications but is rarely implemented in low-cost sensor deployments [30]. Some studies have explored uncertainty estimation using machine learning prediction intervals, ensemble model variance, or residual analysis from co-located comparisons [31]. A particularly important gap is the need for distance-based uncertainty models that recognize a fundamental reality: measurements from sensors located far from reference stations inherently carry greater calibration uncertainty than those close to reference stations. Such models would enable transparent communication of data reliability, allowing end users to make informed decisions about whether data quality is sufficient for their intended application, whether for community awareness, scientific research, or regulatory compliance.
Furthermore, existing quality flagging systems, when they exist at all, tend to be binary (pass/fail) rather than providing graduated quality levels that reflect the reality that data can have varying degrees of reliability suitable for different purposes. Transparent, automated quality flagging systems that clearly communicate confidence levels based on objective criteria—distance to reference stations, calibration model performance, measurement uncertainty—would significantly enhance the credibility and usability of hybrid network data.

2.6. Research Gaps and Study Objectives

Prior studies have established important foundations for calibration transferability research. Van Zoest et al. [32] evaluated spatial transferability of three calibration methods for NO2 electrochemical sensors across a 35-node network in Eindhoven, finding that calibration parameters do not generalise well across locations. Wei et al. [33] demonstrated the importance of temperature-sensitive corrections for long-term NO2 network calibration. De Vito et al. [34] documented “concept drift” when electrochemical NO2 sensor nodes are relocated to different microenvironments, concluding that static field calibrations degrade with relocation. These three studies collectively establish that calibration transfer for gas-phase sensors is problematic; they address NO2 measured by electrochemical sensors exclusively.
Three interconnected gaps remain. First, calibration transfer for optical particulate matter sensors (PM2.5, PM10) has received far less systematic attention; the different physical measurement principles and interference mechanisms require independent investigation. Second, none of the prior works quantify how much uncertainty grows as a function of spatial separation from reference infrastructure, leaving network planners without actionable co-location spacing guidance. Third, the simultaneous treatment of calibration relocation and cross-node sensor-to-sensor fabrication variance has not been addressed within a unified uncertainty framework.
This study addresses these gaps using data from Sofia, Bulgaria’s hybrid monitoring network. We develop and evaluate RF and MLR calibration models on co-located sensor pairs across four pollutants (PM2.5, PM10, NO2, O3), then introduce a distance-based uncertainty quantification framework that translates spatial separation into calibration uncertainty bounds. We provide practical network design guidance demonstrating that co-location sites spaced 6 km apart achieve less than 30% uncertainty increase at network midpoints, within EU indicative monitoring thresholds.

3. Methodology

3.1. Study Area and Monitoring Network

Sofia, Bulgaria’s capital city with approximately 1.3 million residents, is situated in a geological basin surrounded by mountains, creating conditions favorable for temperature inversions that trap air pollutants, particularly during winter heating seasons, when particulate matter concentrations often exceed EU air quality standards. This topographic configuration makes Sofia an ideal case study for hybrid monitoring network development, as the city experiences significant spatiotemporal variability in pollution levels that sparse reference station networks cannot adequately capture. Recent digital twin modeling of mobile-source emissions across Sofia’s urban area has further characterized this heterogeneous pollution landscape [35], underscoring the need for dense monitoring networks capable of resolving intra-urban concentration gradients.

3.1.1. Reference Monitoring Stations

The Executive Environmental Agency (ExEA) operates five official reference stations across Sofia, located in the Pavlovo, Hipodruma, Nadezhda, Mladost, and Druzhba neighborhoods. These stations employ reference-grade instrumentation following European Standard methods [36]: gravimetric measurement (EN 12341:2023) for particulate matter with beta-attenuation monitors for continuous PM10 monitoring [37], chemiluminescence (EN 14211:2024) for nitrogen dioxide and nitrogen monoxide [38], and ultraviolet photometry (EN 14625:2024) for ozone [39]. A critical limitation of this network is that only the Hipodruma station measures PM2.5, creating substantial data scarcity for fine particulate matter analysis across the city, a significant constraint given that PM2.5 is the pollutant most strongly associated with adverse health outcomes. All reference stations measure PM10, NO2, O3, and CO, along with meteorological variables including temperature, atmospheric pressure, relative humidity, wind speed, and wind direction.

3.1.2. Low-Cost Sensor Network

The Sofia Municipality deployed 22 AirThings low-cost air quality sensors throughout the city as part of an initiative to expand spatial monitoring coverage. These sensors employ optical particle counters for particulate matter measurement and electrochemical cells for gaseous pollutants, measuring PM2.5, PM10, NO2, O3, and CO at hourly resolution. While substantially less expensive than reference instrumentation (approximately $200–500 per sensor versus $20,000–50,000 per reference station), these devices face well-documented accuracy challenges including temperature and humidity cross-sensitivities, calibration drift, and inter-sensor variability [5,6]. Critically, four sensors (AT1, AT3, AT11, AT15) were deliberately co-located within meters of reference stations (Hipodruma, Druzhba, Mladost, and Pavlovo, respectively), creating a unique natural experiment for developing and validating calibration transfer methodologies. The fifth reference station (Nadezhda) lacked a co-located sensor and was therefore excluded from calibration training and spatial cross-validation.
The monitoring stations are Develiot Urban Air Quality Monitoring Stations (UAQMS, v7.3, March 2022 hardware generation; Develiot, Sofia). Gaseous pollutants are measured by electrochemical cells: NO2, O3, and CO operate in the low-concentration urban range (0–1 ppm for NO2/O3; 0–10 ppm for CO) with a resolution of <10 ppb and a manufacturer-specified operational lifespan of 12 months, after which sensor replacement is required to maintain calibration stability. Particulate matter is measured by a laser-scattering optical particle counter (range 0–1000 μg/m3, resolution 0.1 μg/m3, 12-month operational lifespan). The integrated meteorological sensor provides temperature (accuracy ±0.2 °C), relative humidity (±1.5%), and atmospheric pressure (±0.5 hPa). All sensing elements are housed in an IP65-rated enclosure with active climatization, a patented thermal management system that reduces temperature and humidity cross-sensitivities inherent to electrochemical and optical sensing technologies.

3.2. Data Collection and Temporal Coverage

We analyze 24 months of hourly air quality and meteorological data spanning 1 August 2023 through 31 July 2025, yielding 17,521 hourly timestamps. This extended temporal coverage was deliberately selected to capture two complete winter heating seasons (November–March), when Sofia experiences the most severe air quality episodes due to residential heating combined with meteorological conditions that trap pollutants. Prior calibration studies have demonstrated that training data spanning at least one complete annual cycle is essential for capturing seasonal variations in sensor response [11], and multi-year datasets reduce the risk of overfitting to atypical meteorological conditions [40]. The dataset encompasses the full range of seasonal variability typical in Central European continental climates, from winter temperature inversions with stagnant air masses to summer photochemical ozone formation events.
Data completeness exceeded expectations for both monitoring systems. Reference stations achieved 95–99% data availability across pollutants (PM10: 97.2%, NO2: 98.6%, O3: 98.4%), with the exception of PM2.5 at Hipodruma (93.7% complete). These availability rates exceed the 90% data capture requirement specified in the EU Air Quality Directive for fixed measurements [36]. The low-cost sensor network demonstrated exceptional reliability with 96.9% average data availability across all 22 sensors, ranging from 89.9% (AT5) to 99.6% (AT17, AT20), substantially higher than the 50–80% availability commonly reported in low-cost sensor deployments [31].

3.3. Data Preprocessing and Quality Control

All raw data underwent systematic quality control procedures following EPA guidance for low-cost sensor evaluation [41] and implemented in Python using pandas (version 2.0+) and numpy (version 1.24+). Timestamps were standardized to UTC and synchronized to hourly resolution across both monitoring systems, ensuring temporal alignment between sensor and reference measurements–a critical step given that even small timing offsets can introduce spurious calibration errors [7].
Physical range validation flagged measurements outside plausible bounds based on historical extremes observed in Sofia and surrounding regions: PM2.5 and PM10 (0–1000 μg/m3), NO2 andO3 (0–500 ppb), temperature (−20 to 45 °C), relative humidity (0–100%), and atmospheric pressure (900–1050 hPa). The Rate-of-change filters identified implausible spikes (>200 μg/m3/h for PM, >100 ppb/h for gases), which typically indicate sensor malfunctions rather than genuine pollution events [29]. Missing data gaps shorter than three consecutive hours were forward-filled using the last valid observation, a conservative approach that avoids introducing interpolation artifacts while maintaining temporal continuity for time-series features. Longer gaps were preserved as missing values and excluded from model training.
Temporal features were extracted to enable machine learning models to capture diurnal and seasonal pollution patterns that reflect emission source activity and atmospheric dynamics [42]. Three integer-valued temporal features were included: hour of day (0–23), month (1–12), and a binary weekend indicator. These raw integer representations were retained rather than cyclically encoded, as the Random Forest algorithm assigns thresholds through recursive partitioning and does not require the numerical continuity assumptions that motivate sine-cosine transformation in distance-based or linear models.
Co-located sensor pairs were identified using Haversine distance calculations between AirThings sensor and ExEA reference station coordinates. Sensors within 5 m were classified as co-located, consistent with EPA recommendations that co-location distances should not exceed the inlet height of reference instruments to minimize microenvironmental differences [43]. This analysis identified four co-located pairs: Hipodruma ↔ AT1, Pavlovo ↔ AT15, Mladost ↔ AT11, and Druzhba ↔ AT3, yielding 69,064 synchronized hourly measurements across all four pairs, with 16,162 valid PM2.5 observation pairs from Hipodruma ↔ AT1 specifically.

3.4. Calibration Model Development

Calibration models were trained exclusively on the four co-located sensor pairs to learn the relationship between low-cost sensor readings and co-located reference measurements. This co-location-based calibration approach, sometimes termed field calibration, has become standard practice in low-cost sensor research because it captures real-world sensor behavior under ambient conditions rather than controlled laboratory settings [8,11]. The fundamental calibration task is to estimate a function f such that:
C ref = f ( C sensor , M , T ) + ϵ
where C ref is the reference pollutant concentration, C sensor is the low-cost sensor reading, M represents meteorological covariates (temperature T, relative humidity RH, atmospheric pressure P), T represents temporal features (hour, month, day of week), and ϵ is irreducible error arising from measurement noise and unmodeled factors. The inclusion of meteorological covariates is essential because low-cost sensors exhibit well-documented cross-sensitivities: optical particle counters are affected by hygroscopic particle growth at high humidity [44,45], while electrochemical gas sensors show temperature-dependent baseline drift and sensitivity changes [46,47].

3.4.1. Model Specifications

Model A: Multiple Linear Regression (MLR) serves as an interpretable baseline, assuming additive linear relationships between predictors and reference concentrations [9]:
C ref = β 0 + β 1 C sensor + β 2 T + β 3 RH + β 4 P + k β k T k + ϵ
where coefficients β are estimated via ordinary least squares. MLR provides a performance floor against which more complex models can be compared, and its coefficients offer direct physical interpretation (e.g., β 3 quantifies the humidity correction per percentage point RH).
Model B: Random Forest (RF), implemented via scikit-learn (version 1.3+), constructs an ensemble of decision trees trained on bootstrap samples with random feature subsets [48]. This ensemble approach has emerged as the dominant method for low-cost sensor calibration due to its ability to capture nonlinear sensor responses, automatic handling of variable interactions, and robustness to outliers [10,11]:
C ^ ref = 1 N trees i = 1 N trees T i ( C sensor , M , T )
Hyperparameters were selected via 5-fold cross-validation on the training set: N trees = 200 (beyond which out-of-bag error stabilized), maximum tree depth = 15 (balancing model complexity against overfitting), minimum samples per leaf = 5 (preventing trees from memorizing individual observations), and maximum features per split = p where p is the number of predictors (the default for regression tasks that promotes tree diversity) [49]. The feature set for each pollutant model included: raw sensor reading, temperature, relative humidity, pressure, and cyclically-encoded temporal features (hoursin, hourcos, monthsin, monthcos, weekend indicator). Feature importance was quantified using the mean decrease in Impurity (MDI), which measures how much each feature contributes to reducing prediction variance in all decision tree splits [48].

3.4.2. Training and Evaluation Protocol

Data from co-located pairs were pooled across all available stations for each pollutant (four stations for PM10, NO2, O3; one station for PM2.5) and split chronologically: the first 80% of timestamps (19.2 months) for training and the final 20% (4.8 months) for independent testing. This temporal split, rather than random sampling, is critical for realistic performance estimation because it simulates the operational scenario where models trained on historical data must predict future measurements [50]. Random splits would allow information leakage across the temporal boundary, artificially inflating performance estimates.
Three complementary cross-validation strategies assessed model generalization across different dimensions of extrapolation:
  • Temporal cross-validation: A single temporal hold-out fold was used for validation: models were trained on the first 18 months (August 2023–January 2025) and evaluated on the held-out final 6 months (February–July 2025). This fold evaluates whether calibration relationships remain stable over time or degrade due to sensor drift, seasonal regime shifts, or changes in emission sources.
  • Seasonal cross-validation: Leave-one-season-out design (winter: December–February; spring: March–May; summer: June–August; fall: September–November). This tests whether models trained predominantly on one pollution regime (e.g., winter heating episodes) generalize to different seasonal conditions.
  • Spatial cross-validation: Leave-one-station-out design for pollutants with multiple co-location pairs (PM10,NO2, O3). This is the most stringent test, evaluating whether calibration relationships learned at one location transfer to geographically distinct sites with potentially different microenvironments, source mixtures, and sensor-specific characteristics [51]. Note:O3 spatial CV was limited to three stations because the Mladost reference station does not measure ozone.
The production calibration models applied to all 22 sensors are subsequently retrained on the full 24-month co-location dataset; the 18-month split is a validation instrument only and does not limit the training data available for deployment.
The performance of the model was evaluated using three complementary metrics: coefficient of determination ( R 2 ) , root mean square error (RMSE), and mean absolute error (MAE). R 2 captures the proportion of variance explained and enables comparison across pollutants with different concentration ranges; RMSE emphasizes large errors and is sensitive to outliers; MAE provides a robust central tendency measure of prediction error magnitude.

3.5. Calibration Transfer and Uncertainty Quantification

The 18 non-co-located sensors represent the primary operational challenge: how to provide calibrated measurements for sensors deployed 1–6 km from the nearest reference station, where no ground-truth reference data exist for validation. This calibration transfer problem–applying models trained at co-located sites to spatially distant sensors–is conceptually related to spatial prediction problems in geostatistics, where prediction uncertainty typically increases with distance from observation locations [52].

3.5.1. Transfer Methodology

For each non-co-located sensor, calibrated concentrations were generated by applying the trained Random Forest model directly to the sensor’s raw measurements and concurrent meteorological data. The underlying assumption is that sensors of the same model (AirThings) exhibit similar response characteristics, such that calibration relationships learned from co-located sensors transfer to non-co-located sensors of identical hardware [13]. This assumption is strongest when sensors are from the same manufacturing batch and have similar deployment ages, and weakest when inter-sensor variability is high, or sensors have experienced differential aging or damage.

3.5.2. Distance-Based Uncertainty Model

We developed a distance-based uncertainty quantification framework to provide realistic confidence bounds for transferred calibrations. The conceptual basis is that calibration uncertainty should increase with spatial separation from reference stations due to: (1) spatial heterogeneity in pollution source mixtures and atmospheric conditions that the calibration model was not trained on, (2) potential inter-sensor variability not captured by models trained on different sensor units, and (3) microenvironmental differences between co-located sites (often in controlled settings near official stations) and typical deployment locations [42].
We model measurement uncertainty as a linear function of distance:
σ j ( d ) = σ base · 1 + α · d j
where σ j ( d ) is the estimated prediction uncertainty for sensor j, σ base is the baseline calibration uncertainty estimated as the intercept of a linear regression of RMSE against distance from spatial cross-validation results (representing expected uncertainty at zero distance from reference stations), d j is the Haversine distance to the nearest reference station (km), and α is the uncertainty growth rate parameter (km−1). The linear form was chosen for parsimony and interpretability; more complex functional forms (exponential, power-law) were explored but did not significantly improve fit given the limited distance range (1–6 km) in our network.
The parameter α was estimated empirically using the spatial cross-validation results. In leave-one-station-out cross-validation, each held-out station provides an estimate of prediction error at a known distance from the training stations. By regressing absolute prediction residuals against distance across all held-out evaluations, we obtained pollutant-specific α estimates. This approach assumes that performance degradation observed when transferring to held-out co-located stations is representative of degradation expected when transferring to non-co-located sensors at similar distances, a reasonable approximation given that co-located stations span diverse microenvironments across Sofia.

3.5.3. Prediction Intervals

The distance-scaled uncertainty estimates provide 95% prediction intervals for each calibrated measurement:
PI 95 ( j ) = C ^ sensor , j ± 1.96 · σ j ( d )
assuming approximately normal prediction errors, which was verified by examining residual distributions from co-located calibration. These intervals communicate measurement reliability to end users: sensors near reference stations have narrow intervals reflecting high confidence, while peripheral sensors have wider intervals appropriately reflecting greater calibration uncertainty. This transparent uncertainty quantification addresses a key criticism of low-cost sensor deployments, that measurements are often presented without acknowledgment of their substantial uncertainty [30].

4. Results

4.1. Co-Location Calibration Performance

Table 1 summarizes the calibration performance of both models across all four pollutants on the held-out test set. Random Forest consistently outperformed MLR across all pollutants, with R 2 improvements ranging from 40% (O3) to 407% (PM2.5). Better results are in bold.
Figure 1 shows calibration scatter plots comparing Random Forest predictions against reference measurements for each pollutant.
Based on these results, Random Forest was selected as the operational calibration model for transfer to non-co-located sensors.

4.2. Cross-Validation Analysis

To assess model generalizability beyond the training period and locations, we conducted both temporal and spatial cross-validation experiments. The temporal cross-validation design trained models on the first 18 months of data and reserved the final 6 months as an independent test set, simulating the operational deployment scenario where calibration models must extrapolate to future time periods that may differ in pollution regimes and meteorological conditions. Under this temporal holdout design, MLR performance degraded substantially for some pollutants, with NO2 dropping from R 2 = 0.20 to R 2 = 0.08 . Negative R 2 values indicate that predictions performed worse than simply using the training period mean as a constant predictor, suggesting that the linear model failed to capture temporal dynamics that differ between training and test periods. In contrast, O3 maintained a more stable performance ( R 2 = 0.44 ) , indicating that the drivers of ozone calibration are more temporally consistent. These results suggest that the importance of temporal factors varies substantially across pollutants.
Spatial cross-validation using a leave-one-station-out design provided critical insights about calibration transferability across Sofia’s monitoring network. In this experiment, Random Forest models were trained on three of the four co-located pairs and evaluated on the held-out fourth pair, rotating through all combinations. Table 2 summarizes performance when each co-location site is held out for testing.
The variable performance across stations ( R 2 ranging from 0.15 to 0.57 with the full feature set) indicates that calibration transfer is not uniformly successful across all locations. Hipodruma shows particularly poor O3 transfer performance, likely reflecting unique microenvironmental conditions at this central urban site, while Druzhba NO2 also yields negative R 2 , consistent with the pronounced NO2 spatial gradient near this traffic-exposed location. Notably, temporal features (hour, month, weekend indicator) consistently improve spatial transfer performance across all three pollutants and all stations: mean R 2 increases from 0.25 to 0.33 for PM10, from 0.03 to 0.23 for NO2, and from 0.19 to 0.29 for O3 when temporal features are included. This empirical result indicates that, for Sofia’s compact urban network, the shared diurnal and seasonal structure of emission sources across neighborhoods provides a generalizable signal rather than location-specific overfitting, as might be expected in a larger or more spatially heterogeneous city. This variability underscores the importance of the uncertainty quantification framework developed in Section 3.5.
Seasonal cross-validation (Table 3) revealed substantial variation in calibration performance across seasons. O3 achieved the highest seasonal R 2 during summer ( 0.46 ) , reflecting the strong temperature-ozone relationship during photochemically active months. Conversely, PM2.5 showed negative R 2 during summer ( 0.15 ) , suggesting that calibration relationships developed during heating-season pollution episodes may not transfer well to summer conditions with different particle composition. NO2 exhibited consistently modest performance across all seasons ( R 2 ( 0.15 , 0.15 ) ) , indicating challenges in generalizing electrochemical sensor calibrations across seasonal emission and meteorological regimes.
To verify that RF extrapolation is not a concern in practice, we checked what fraction of transfer-site observations fall within the training feature space. Across all three pollutants, 99.5% of transfer observations ( n > 300 , 000 ) lie within the min–max bounds of the co-location training data for every feature; the only gap is 0.5% of temperature readings marginally outside the training range. RF extrapolation is therefore not a material source of error in this network.

4.3. Feature Importance Analysis

Random Forest feature importance analysis (Figure 2) reveals that fundamentally different calibration mechanisms operate across pollutant types. For particulate matter, the sensor readings themselves dominate the calibration model, contributing 60% of the total importance for PM2.5 and 58% for PM10. Temperature and seasonal features (encoded through month) add modest contributions of 10–18%, primarily capturing humidity-related optical interference in winter months. This pattern indicates that the optical particle counters in these low-cost sensors provide informative base signals that require only modest environmental correction to align with reference measurements.
The calibration mechanism for NO2 differs markedly from that of particulate matter. Temporal features dominate the model, with month contributing 28% and hour of day contributing 16% of total importance, together exceeding the 32% contribution from the sensor reading itself. This pattern reflects the strong diurnal and seasonal cycles characteristic of NO2 concentrations, driven by morning and evening traffic rush hours and seasonal variations in atmospheric mixing depth. The implication is that electrochemical NO2 sensors require substantial temporal adjustment, and the calibration model essentially learns to predict typical NO2 patterns for a given time of day and season, with the sensor reading providing secondary refinement.
Ozone calibration is dominated almost entirely by temperature, which accounts for 57% of feature importance while the sensor reading contributes only 16%. This finding reflects the well-documented temperature cross-sensitivity of electrochemical O3 sensors, where the sensing element responds not only to ozone concentration but also strongly to ambient temperature. The calibration model thus primarily corrects for thermal artifacts rather than extracting an ozone-specific signal from the sensor output.
These non-linear sensor behaviours directly explain the large RF–MLR performance gap (average 246% improvement). Multiple Linear Regression can fit only a single slope per feature; it cannot represent the temperature-saturating response of EC O3 sensors, the interaction between humidity and optical PM scattering, or the non-linear diurnal NO2 profile. Random Forest partitions the feature space into local regions where each non-linearity is captured independently, which is precisely why ensemble methods outperform linear baselines so markedly for these sensor types.
These distinct calibration mechanisms have practical implications for sensor network design and transferability. Particulate matter sensors, where the sensor reading dominates calibration, may transfer more reliably across locations because the fundamental sensor response remains consistent. Gaseous sensors, in contrast, require careful attention to local temporal pollution patterns and temperature regimes, and calibrations developed at one location may not transfer as effectively to sites with different diurnal traffic patterns or thermal environments.

4.4. Calibration Transfer to Non-Co-Located Sensors

Having established Random Forest as the preferred calibration approach through the co-location experiments, we applied the trained models to calibrate measurements from the 18 non-co-located sensors distributed across Sofia. Over the 24-month study period, this transfer process generated a total of 1,207,809 calibrated hourly measurements, each accompanied by distance-scaled uncertainty estimates derived from the framework described in Section 3.5.
The calibrated measurements reveal the pollution landscape across Sofia’s urban area. Mean PM2.5 concentrations of 14.5 μg/m3 indicate moderate fine particulate pollution that exceeds the WHO guideline of 5 μg/m3 but remains below the EU annual limit of 25 μg/m3. PM10 levels are elevated at 27.3 μg/m3, reflecting contributions from road dust, construction activity, and residential heating during winter months. NO2 concentrations average 28.9 μg/m3, consistent with traffic-dominated urban environments. Ozone levels are relatively high at 48.7 μg/m3, typical of continental European cities where strong solar radiation during summer drives photochemical production from precursor pollutants transported into the Sofia basin.

4.5. Spatial Uncertainty Quantification

A central question for network planners is how measurement uncertainty grows as sensors are deployed increasingly far from reference stations. To address this, we fitted linear uncertainty growth models relating prediction residual magnitude to distance from the nearest co-location site. Table 4 presents the fitted uncertainty growth parameters for all pollutants, including the baseline uncertainty at co-located sites ( σ base ), the growth rate per kilometer ( α ), and statistical significance of the distance effect.
Figure 3 visualizes the uncertainty growth model, showing how prediction uncertainty increases with distance from co-location sites for each pollutant.
Figure 4 illustrates the spatial distribution of monitoring stations across Sofia and the predicted PM10 calibration uncertainty for each non-co-located sensor. We present PM10 because it exhibits the highest uncertainty growth rate (5.62%/km) among the pollutants with statistically significant distance effects, representing the most conservative (worst-case) scenario for network planning. The map demonstrates the practical application of the uncertainty framework: sensors near reference stations (e.g., AT8 at 6% uncertainty) can provide reliable, calibrated measurements, while sensors at the network periphery (e.g., AT2 at 31% uncertainty) require more cautious interpretation.
The modest uncertainty growth rates, all below 6% per kilometer, have direct implications for practical network design. At 6 km spacing between co-location sites, sensors positioned at the network midpoint (3 km from the nearest reference station) experience less than 30% uncertainty increase relative to co-located sensors. This degradation remains well within acceptable bounds for indicative monitoring applications, suggesting that relatively sparse co-location infrastructure can support city-wide calibration transfer without excessive quality degradation. The stronger statistical significance for PM10 and O3 provides confidence in these design recommendations for coarse particulate and ozone monitoring, while the consistent trends for PM2.5 and NO2, though not statistically significant at conventional thresholds given the limited sample of 18 transfer sensors, point in the same direction.

5. Discussion

5.1. Summary of Key Findings

This study developed and validated a machine learning-based calibration transfer framework for hybrid air quality monitoring networks. Our investigation yielded three primary findings:
First, Random Forest calibration substantially outperformed Multiple Linear Regression across all pollutants ( R 2 ( 0.53 , 0.75 ) vs. R 2 ( 0.14 , 0.53 ) ), demonstrating the value of ensemble learning for capturing non-linear sensor-environment-pollutant relationships. The performance gains were most pronounced for particulate matter (PM2.5: 408% improvement, PM10: 370% improvement).
Second, and most significantly, we quantified spatial uncertainty degradation through a distance-based framework characterized by pollutant-specific uncertainty growth rates ( α parameters): PM10 (5.62%/km), PM2.5 (5.33%/km), NO2 (4.24%/km), and O3 (3.84%/km). These α parameters enable prediction of calibration uncertainty for sensors deployed 1–6 km from reference stations without requiring ground-truth validation.
Third, the modest uncertainty growth rates (<6%/km for all pollutants) support practical network design recommendations: co-location sites spaced 6 km apart achieve <30% uncertainty increase at network midpoints, enabling cost-effective city-wide monitoring.

5.2. Comparison with Literature

Table 5 contextualizes our calibration performance against published studies.
Our PM2.5 calibration ( R 2 = 0.745 ) is notable given reliance on a single co-located training pair (Hipodruma ↔ AT1, n = 16,162 ), whereas benchmark studies typically employ multiple co-location sites. This suggests that the 24-month training duration compensates partially for limited spatial replication. NO2 performance ( R 2 = 0.532 ) aligns with the literature findings that electrochemical NO2 sensors are more challenging to calibrate than optical PM sensors due to cross-sensitivities to O3 and temperature.
The MobiliSense study [53] found median correlations of only from 0.21 to 0.27 between personal exposure and fixed station measurements, with slight improvement as distance to the nearest station decreased. While that study addressed a different question, comparing mobile individuals against fixed monitors rather than fixed sensors at varying distances, it reinforces the broader principle that distance from reference infrastructure degrades measurement reliability. Our α parameters quantify this same phenomenon for fixed sensor networks, providing the operational metrics that network planners need.

5.3. Methodological Contribution

The distance-based uncertainty quantification framework addresses a practical gap in low-cost sensor deployment. While geostatistical methods (kriging, land-use regression) quantify ambient spatial variability of pollutant concentrations [10], and personal exposure studies demonstrate distance-dependent misclassification between fixed stations and mobile individuals [53], neither approach directly addresses the operational question facing network planners: how does calibration uncertainty degrade as sensors are deployed farther from reference stations?
Prior studies of calibration transfer (Van Zoest et al. [32], Wei et al. [33], and De Vito et al. [34]) established that transferring NO2 electrochemical sensor calibrations across locations is unreliable, but did not quantify uncertainty as a function of spatial separation, did not address optical PM sensors, and did not provide actionable network design guidance. The present work extends this body of knowledge in three respects: it covers optical PM sensors (PM2.5, PM10) alongside electrochemical gas sensors; it introduces pollutant-specific α parameters that convert distance into a calibration uncertainty bound; and it simultaneously captures both calibration relocation and cross-node fabrication variance within a single operational metric.
Our framework fills this gap by providing actionable α parameters (3.8–5.6%/km) that quantify calibration transfer uncertainty. This quantity is distinct from ambient concentration variability: it measures how much additional error is introduced by applying a calibration model developed at one location to sensors at another location, rather than how much pollutant concentrations vary spatially. These parameters enable practitioners to predict measurement reliability at any network location without requiring ground-truth validation, supporting cost-effective decisions about co-location infrastructure density. While the specific α values derived here are Sofia-specific, reflecting its basin topography, source mix, and climate, the methodology itself is transferable to other cities implementing hybrid monitoring networks, though local calibration studies would be required to derive city-specific uncertainty parameters.

5.4. Practical Implications

The regulatory context provides important benchmarks for evaluating this framework’s practical utility. The EU Air Quality Directive [36] specifies Data Quality Objectives for indicative measurements at ± 50 % uncertainty for particulate matter. Our framework achieves less than 30% uncertainty increase at network midpoints when co-location sites are spaced 6 km apart, comfortably within these indicative measurement thresholds. This positions calibrated low-cost sensor networks as viable supplements to reference monitoring for indicative air quality assessment, community awareness applications, and preliminary spatial mapping. However, they cannot replace the fixed reference stations required for regulatory compliance monitoring and enforcement actions. It is also worth noting that reference instruments themselves carry measurement uncertainty, typically 10–15% for PM analyzers, and this baseline uncertainty propagates through calibration transfer to the low-cost sensor measurements.
The economic case for hybrid networks is compelling when comparing the costs of equivalent spatial coverage. Reference-grade PM analyzers cost approximately $20,000–50,000 per station, including installation, housing, and utility connections, while low-cost sensors cost $200–500 per unit with minimal infrastructure requirements. For Sofia’s 18-sensor low-cost network, achieving equivalent spatial coverage with reference stations would require approximately $360,000 (18 stations at $20,000 minimum each), compared to approximately $10,000 for the low-cost sensors plus the co-location infrastructure already provided by existing reference stations. This represents a 36-fold cost reduction while providing roughly ten times better spatial resolution than the existing 5-station reference network alone.
The uncertainty growth parameters derived in this study enable evidence-based decisions about co-location spacing. A spacing of 2 km between co-location sites achieves less than 10% uncertainty increase at the midpoint between stations, suitable for applications requiring high confidence. Spacing of 6 km achieves less than 30% uncertainty increase, acceptable for indicative monitoring and spatial screening. For PM10, NO2, and O3, where four co-location pairs support robust spatial cross-validation, these recommendations rest on a solid empirical foundation. The situation for PM2.5 is more constrained because Sofia operates only a single reference station measuring fine particulate matter, and expanding PM2.5 reference coverage should be a priority for municipalities seeking comprehensive hybrid network design.

5.5. Limitations and Future Research

Four limitations constrain our findings, each suggesting specific directions for future research.
The most significant limitation is that our uncertainty estimates at transfer sites are predicted rather than validated. The α parameters quantify how uncertainty grows with distance from co-location sites, but these predictions have not been empirically verified by deploying reference instruments at the transfer locations. Mobile reference campaigns offer a path forward: deploying portable reference-grade instruments at selected transfer sites (for example, at 2 km, 4 km, and 6 km from co-location stations) for one to two week intensive measurement periods could directly validate whether observed prediction errors match the uncertainty estimates our framework provides. Sofia’s Executive Environmental Agency possesses mobile monitoring capabilities that would be suitable for such validation studies.
The distance range over which our linear uncertainty growth model has been validated is limited to 1–6 km, the spatial extent of Sofia’s current sensor network. Whether uncertainty continues to grow linearly beyond this range, or instead saturates at some asymptotic level or accelerates at greater distances, remains unknown. Extended-range validation studies in larger metropolitan areas such as London, Paris, or Berlin could address this gap and determine whether the linear model requires modification for city-wide deployments spanning tens of kilometers.
The α values derived here reflect Sofia’s particular characteristics: a basin topography that traps pollutants, a source mix dominated by residential heating and traffic, and a continental climate with strong seasonal temperature variations. Whether these parameters generalize to other cities is an open question. Multi-city comparative studies applying identical methodology could establish whether α parameters cluster by city typology (basin versus coastal versus flat terrain), climate zone, or dominant pollution sources, potentially enabling transferable lookup tables that would allow network planners in new cities to estimate uncertainty growth rates without requiring local calibration studies.
The statistical significance of the distance–uncertainty relationship varies by pollutant: PM10 ( p = 0.007 ) and O3 ( p = 0.031 ) are significant at α = 0.05 , while PM2.5 ( p = 0.082 ) and NO2 ( p = 0.135 ) are not. This disparity is primarily a power constraint rather than evidence that the distance effect is absent for the non-significant pollutants. With only n = 18 transfer sensors spanning a 1–5.6 km distance range, the regression has limited degrees of freedom; a positive trend can exist and remain physically consistent yet fall short of conventional significance thresholds. The α estimates for PM2.5 (5.33%/km) and NO2 (4.24%/km) are quantitatively consistent with the significant PM10 (5.62%/km) and O3 (3.84%/km) values, and all four point estimates have the same sign. Practitioners should treat the PM2.5 and NO2  α values as indicative rather than prescriptive; validation with larger networks spanning wider distance ranges would narrow the confidence intervals and determine whether the linear model holds for these pollutants.
The inclusion of temporal features (hour, month, weekend indicator) in the calibration models warrants discussion in the context of spatial transfer. It has been argued that temporal features may encode location-specific emission profiles that do not generalize to other sites. We tested this directly by running leave-one-station-out spatial cross-validation with RF models both with and without temporal features (Section 4.2). Empirically, temporal features improved spatial transfer for all three pollutants in Sofia’s network. This outcome reflects the city’s relatively compact geography, where major emission sources (residential heating, traffic) follow broadly similar diurnal and seasonal rhythms across neighborhoods. In networks with greater spatial heterogeneity, for example, in cities with highly localized industrial sources, this assumption may not hold, and temporal features could introduce the location-specific bias that prior literature cautions against.
The spatial cross-validation reveals substantial station-to-station variability, with negative R 2 at specific held-out stations (NO2 at Druzhba, O3 at Hipodruma). This is a structural consequence of the small leave-one-station-out training set: with only four co-location sites, removing one leaves just three spatial training points to span Sofia’s urban variability. When the held-out station occupies a distinctive microenvironment (Druzhba’s predominantly residential NO2 profile is not well represented by the three remaining stations, and Hipodruma’s city-center photochemical O3 dynamics differ from those at the outer-ring stations), the calibration model lacks sufficient spatial context to generalize, producing predictions that are systematically offset from the true concentration mean. This negative R 2 reflects microenvironmental specificity rather than a general failure of the calibration transfer method; the positive R 2 values achieved when the full four-station co-location dataset is used confirm that the approach performs reliably when spatial coverage is adequate.
A further limitation concerns interferent gases. Cross-sensitivity between non-target and target species is well documented in electrochemical gas sensors: O3 sensors are sensitive to NO2 and vice versa. The AirThings sensors used here provide only calibrated concentration outputs through their consumer API; raw electrochemical voltages and cross-channel signals are not accessible without manufacturer cooperation. Interferent correction was therefore not possible within our framework. Calibration accuracy for NO2 and O3 may consequently be reduced at transfer sites where the NO2/O3 ratio differs substantially from co-location training conditions.
Random Forest regressors cannot extrapolate beyond the feature space seen during co-location training; however, a coverage check (Section 4.2) confirms that 99.5% of transfer-site observations lie within the training bounds, so this structural limitation has negligible practical effect in the present network. Practitioners deploying this framework in networks with substantially different pollution regimes or climatic conditions should verify that transfer-site features overlap with training-site distributions before applying the calibration.
Finally, the constraint of having only a single reference station measuring PM2.5 in Sofia means that spatial cross-validation is impossible for this health-critical pollutant. All PM2.5 calibration and uncertainty estimates rest on a single co-located pair, and we cannot assess how well these calibrations transfer across the city’s diverse microenvironments. Expansion of PM2.5 reference monitoring infrastructure would enable robust spatial validation and more reliable uncertainty quantification for fine particulate matter.

6. Conclusions

This study presents a comprehensive framework for creating validated hybrid air quality monitoring networks that combine the accuracy of official reference stations with the spatial density of low-cost sensor arrays.
Random Forest calibration models achieved R 2 values ranging from 0.53 to 0.75 across four pollutants (PM2.5, PM10, NO2, and O3), representing an average improvement of 246% over Multiple Linear Regression baselines. The superior performance of ensemble methods reflects their ability to capture the non-linear relationships between sensor readings, environmental conditions, and true pollutant concentrations that linear models cannot represent.
The central methodological contribution is the empirical derivation of distance-based uncertainty growth rates. These α parameters quantify how calibration uncertainty increases as sensors are deployed farther from reference stations: 5.62% per kilometer for PM10, 5.33% for PM2.5, 4.24% for NO2, and 3.84% for O3. These parameters enable practitioners to predict measurement uncertainty at any network location without requiring ground-truth validation at each sensor site, addressing a longstanding gap in low-cost sensor deployment methodology.
The modest uncertainty growth rates translate directly into practical network design guidance. Co-location sites spaced 6 km apart achieve less than 30% uncertainty increase at network midpoints, well within acceptable bounds for indicative monitoring applications. For PM10, NO2, and O3, where four co-location pairs support robust cross-validation, four to five strategically placed co-location sites can provide calibration anchors for a city-wide sensor network. PM2.5 network design remains constrained by limited reference infrastructure in Sofia, highlighting the importance of expanding fine particulate monitoring capacity.
The operational demonstration of this framework, generating 1.2 million calibrated measurements for Sofia’s 18 non-co-located sensors over 24 months, each accompanied by 95% prediction intervals, establishes that hybrid networks can provide spatially dense, quality-assured air quality data at a fraction of the cost of traditional reference-station expansion. This framework offers municipalities actionable guidance for implementing cost-effective hybrid monitoring networks while maintaining the transparent quality assurance essential for scientific credibility and public trust.

Author Contributions

Conceptualization, P.Z. and S.F.; methodology, P.Z.; software, P.Z.; validation, P.Z.; writing—original draft preparation, P.Z.; writing—review and editing, S.F. All authors have read and agreed to the published version of the manuscript.

Funding

The work was partially supported by the Centre of Excellence in Informatics and ICT under the Grant No. BG16RFPR002-1.014-0018-C01, financed by the Research, Innovation and Digitalization for Smart Transformation Programme 2021–2027 and co-financed by the European Union.

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

The data presented in this study are openly available in https://citylab.gate-ai.eu/sofiasensors/, accessed on 23 March 2026.

Conflicts of Interest

The authors declare no conflicts of interest.

Abbreviations

The following abbreviations are used in this manuscript:
AQAir Quality
EUEuropean Union
ExEAExecutive Environmental Agency (Bulgaria)
EPAU.S. Environmental Protection Agency
LSTMLong Short-Term Memory
MAEMean Absolute Error
MDIMean Decrease in Impurity
MLRMultiple Linear Regression
NO2Nitrogen Dioxide
O3Ozone
PMParticulate Matter
PM2.5Particulate Matter with aerodynamic diameter ≤ 2.5 μm
PM10Particulate Matter with aerodynamic diameter ≤ 10 μm
QA/QCQuality Assurance/Quality Control
RFRandom Forest
RMSERoot Mean Square Error
WHOWorld Health Organization

References

  1. World Health Organization. WHO Global Air Quality Guidelines: Particulate Matter (PM2.5 and PM10), Ozone, Nitrogen Dioxide, Sulfur Dioxide and Carbon Monoxide; World Health Organization: Geneva, Switzerland, 2021; ISBN 978-92-4-003422-8. Available online: https://iris.who.int/handle/10665/345329 (accessed on 1 December 2025).
  2. European Environment Agency. Europe’s Air Quality Status 2023; EEA Briefing No. 05/2023; European Environment Agency: Copenhagen, Denmark, 2023; ISBN 978-92-9480-554-6. [CrossRef]
  3. Fidanova, S.; Zhivkov, P.; Roeva, O. InterCriteria Analysis Applied on Air Pollution Influence on Morbidity. Mathematics 2022, 10, 1195. [Google Scholar] [CrossRef] [Scilit]
  4. Snyder, E.G.; Watkins, T.H.; Solomon, P.A.; Thoma, E.D.; Williams, R.W.; Hagler, G.S.W.; Shelow, D.; Hindin, D.A.; Kilaru, V.J.; Preuss, P.W. The Changing Paradigm of Air Pollution Monitoring. Environ. Sci. Technol. 2013, 47, 11369–11377. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  5. Morawska, L.; Thai, P.K.; Liu, X.; Asumadu-Sakyi, A.; Ayoko, G.; Bartonova, A.; Bedini, A.; Chai, F.; Christensen, B.; Dunbabin, M.; et al. Applications of Low-Cost Sensing Technologies for Air Quality Monitoring and Exposure Assessment: How Far Have They Gone? Environ. Int. 2018, 116, 286–299. [Google Scholar] [CrossRef] [Scilit]
  6. Castell, N.; Dauge, F.R.; Schneider, P.; Vogt, M.; Lerner, U.; Fishbain, B.; Broday, D.; Bartonova, A. Can Commercial Low-Cost Sensor Platforms Contribute to Air Quality Monitoring and Exposure Estimates? Environ. Int. 2017, 99, 293–302. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  7. Spinelle, L.; Gerboles, M.; Villani, M.G.; Aleixandre, M.; Bonavitacola, F. Field Calibration of a Cluster of Low-Cost Commercially Available Sensors for Air Quality Monitoring. Part B: NO, CO and CO2. Sensors Actuators B Chem. 2017, 238, 706–715. [Google Scholar] [CrossRef] [Scilit]
  8. Hagan, D.H.; Isaacman-VanWertz, G.; Franklin, J.P.; Wallace, L.M.M.; Kocar, B.D.; Heald, C.L.; Kroll, J.H. Calibration and Assessment of Electrochemical Air Quality Sensors by Co-Location with Regulatory-Grade Instruments. Atmos. Meas. Tech. 2018, 11, 315–328. [Google Scholar] [CrossRef] [Scilit]
  9. Cordero, J.M.; Borge, R.; Narros, A. Using Statistical Methods to Carry Out in Field Calibrations of Low Cost Air Quality Sensors. Sens. Actuators B Chem. 2018, 267, 245–254. [Google Scholar] [CrossRef] [Scilit]
  10. Zimmerman, N.; Presto, A.A.; Kumar, S.P.N.; Gu, J.; Hauryliuk, A.; Robinson, E.S.; Robinson, A.L.; Subramanian, R. A Machine Learning Calibration Model Using Random Forests to Improve Sensor Performance for Lower-Cost Air Quality Monitoring. Atmos. Meas. Tech. 2018, 11, 291–313. [Google Scholar] [CrossRef] [Scilit]
  11. Malings, C.; Tanzer, R.; Hauryliuk, A.; Kumar, S.P.N.; Zimmerman, N.; Kara, L.B.; Presto, A.A.; Subramanian, R. Development of a General Calibration Model and Long-Term Performance Evaluation of Low-Cost Sensors for Air Pollutant Gas Monitoring. Atmos. Meas. Tech. 2020, 13, 3693–3713. [Google Scholar] [CrossRef] [Scilit]
  12. Devarakonda, S.; Sevusu, P.; Liu, H.; Liu, R.; Iftode, L.; Nath, B. Real-Time Air Quality Monitoring Through Mobile Sensing in Metropolitan Areas. In Proceedings of the 2nd ACM SIGKDD International Workshop on Urban Computing, Chicago, IL, USA, 11 August 2013; pp. 1–8. [Google Scholar]
  13. Kizel, F.; Etzion, D.; Shafran-Nathan, R.; Nahum, V.L.; Broday, D.M. Node-to-Node Field Calibration of Wireless Distributed Air Pollution Sensor Network. Environ. Pollut. 2018, 233, 900–909. [Google Scholar] [CrossRef] [Scilit]
  14. Liang, L.; Daniels, J. What Influences Low-Cost Sensor Data Calibration? A Systematic Assessment of Algorithms, Duration, and Predictor Selection. Aerosol Air Qual. Res. 2022, 22, 220076. [Google Scholar] [CrossRef] [Scilit]
  15. Vajs, I.; Drajic, D.; Cica, Z. Data-Driven Machine Learning Calibration Propagation in A Hybrid Sensor Network for Air Quality Monitoring. Sensors 2023, 23, 2815. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  16. Apostolopoulos, I.D.; Fouskas, G.; Pandis, S.N. Field Calibration of a Low-Cost Air Quality Monitoring Device in an Urban Background Site Using Machine Learning Models. Atmosphere 2023, 14, 368. [Google Scholar] [CrossRef] [Scilit]
  17. Zhivkov, P. Optimization and Evaluation of Calibration for Low-Cost Air Quality Sensors: Supervised and Unsupervised Machine Learning Models. In Proceedings of the 2021 16th Conference on Computer Science and Intelligence Systems (FedCSIS), Online, 2–5 September 2021; pp. 255–258. [Google Scholar]
  18. Sousan, S.; Wu, R.; Popoviciu, C.; Fresquez, S.; Park, Y.M. Advancing low-cost air quality monitor calibration with machine learning methods. Environ. Pollut. 2025, 374, 126191. [Google Scholar] [CrossRef] [Scilit]
  19. Koziel, S.; Pietrenko-Dabrowska, A.; Wojcikowski, M.; B, P. High-performance machine-learning-based calibration of low-cost nitrogen dioxide sensor using environmental parameter differentials and global data scaling. Sci. Rep. 2024, 14, 26120. [Google Scholar] [CrossRef] [Scilit]
  20. Tastan, M. Machine Learning–Based Calibration and Performance Evaluation of Low-Cost Internet of Things Air Quality Sensors. Sensors 2025, 25, 3183. [Google Scholar] [CrossRef] [Scilit]
  21. Yaqoob, I.; Kumar, V.; Chaudhry, S.A. Machine Learning Calibration of Low-Cost Sensor PM2.5 data. In Proceedings of the 2024 IEEE International Symposium on Systems Engineering (ISSE), Perugia, Italy, 16–18 October 2024; pp. 1–8. [Google Scholar]
  22. Pan, S.J.; Yang, Q. A Survey on Transfer Learning. IEEE Trans. Knowl. Data Eng. 2010, 22, 1345–1359. [Google Scholar] [CrossRef] [Scilit]
  23. Weiss, K.; Khoshgoftaar, T.M.; Wang, D. A Survey of Transfer Learning. J. Big Data 2016, 3, 1–40. [Google Scholar] [CrossRef] [Scilit]
  24. Vito, S.D.; Fattoruso, G.; Pardo, M.; Tortorella, F.; Francia, G. A Global Multiunit Calibration as a Method for Large-Scale IoT Particulate Matter Monitoring Systems Deployments. IEEE Trans. Instrum. Meas. 2024, 73, 2501916. [Google Scholar] [CrossRef] [Scilit]
  25. Mead, M.I.; Popoola, O.A.M.; Stewart, G.B.; Landshoff, P.; Calleja, M.; Hayes, M.; Baldovi, J.J.; McLeod, M.W.; Hodgson, T.F.; Dicks, J.; et al. The Use of Electrochemical Sensors for Monitoring Urban Air Quality in Low-Cost, High-Density Networks. Atmos. Environ. 2013, 70, 186–203. [Google Scholar] [CrossRef] [Scilit]
  26. Bigi, A.; Bianchi, M.; Ghermandi, G. Performance Assessment of Low-Cost Sensors for Air Quality Monitoring: The Case Study of a Hybrid Network in the Po Valley. Sensors 2023, 23, 2388. [Google Scholar]
  27. Fishbain, B.; Lerner, D.; Castell, N.; Cole-Hunter, T.; Popoola, O.; Broday, D.M.; Iñiguez, T.M.; Nieuwenhuijsen, M.; Jovasevic-Stojanovic, M.; Topalovic, D.; et al. An Evaluation Tool Kit for Assessment of Low-Cost Sensors for Air Quality Measurement. In Proceedings of the WMO/CIMO Technical Conference, Main, Germany, 24–26 October 2017. [Google Scholar]
  28. Williams, R.; Kilaru, V.; Snyder, E.; Kaufman, A.; Dye, T.; Rutter, A.; Russell, A.; Hafner, H. Air Sensor Guidebook; EPA/600/R-14/159 (NTIS PB2015-100610); U.S. Environmental Protection Agency: Washington, DC, USA, 2014. Available online: https://cfpub.epa.gov/si/si_public_record_Report.cfm?dirEntryId=277996 (accessed on 1 January 2026).
  29. Clements, A.L.; Griswold, W.B.; Rs, A.; Johnston, J.E.; Herting, M.M.; Thorson, J.; Collier-Oxandale, A.; Hannigan, M. Low-Cost Air Quality Monitoring Tools: From Research to Practice (A Workshop Summary). Sensors 2017, 17, 2478. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  30. Rai, A.C.; Kumar, P.; Pilla, F.; Skouloudis, A.N.; Sabatino, S.D.; Ratti, C.; Yasar, A.; Rickerby, D. End-User Perspective of Low-Cost Sensors for Outdoor Air Pollution Monitoring. Sci. Total. Environ. 2017, 607–608. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  31. Liu, H.Y.; Schneider, P.; Haugen, R.; Vogt, M. Performance Assessment of a Low-Cost PM2.5 Sensor for a Near Four-Month Period in Oslo, Norway. Atmosphere 2019, 10, 41. [Google Scholar] [CrossRef] [Scilit]
  32. van Zoest, V.; Osei, F.B.; Stein, A.; Hoek, G. Calibration of Low-Cost NO2 Sensors in an Urban Air Quality Network. Atmos. Environ. 2019, 210, 66–75. [Google Scholar] [CrossRef] [Scilit]
  33. Wei, P.; Sun, L.; Anand, A.; Zhang, Q.; Huixin, Z.; Deng, Z.; Wang, Y.; Ning, Z. Development and Evaluation of a Robust Temperature Sensitive Algorithm for Long Term NO2 Gas Sensor Network Data Correction. Atmos. Environ. 2020, 232, 117509. [Google Scholar] [CrossRef] [Scilit]
  34. Vito, S.D.; Esposito, E.; Castell, N.; Schneider, P.; Bartonova, A. On the Robustness of Field Calibration for Smart Air Quality Monitors. Sens. Actuators B Chem. 2020, 310, 127869. [Google Scholar] [CrossRef] [Scilit]
  35. Zhivkov, P.; Fidanova, S.; Dimov, I. Digital Twin-Based Framework for Real-Time Monitoring and Analysis of Urban Mobile-Source Emissions. Atmosphere 2025, 16, 731. [Google Scholar] [CrossRef] [Scilit]
  36. European Parliament and Council. Directive 2008/50/EC of the European Parliament and of the Council of 21 May 2008 on Ambient Air Quality and Cleaner Air for Europe. Off. J. Eur. Union 2008, L 152/1, 1–44. Available online: https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX:32008L0050 (accessed on 1 September 2025).
  37. EN 12341:2023; Ambient Air—Standard Gravimetric Measurement Method for the Determination of the PM10 or PM2.5 Mass Concentration of Suspended Particulate Matter. European Committee for Standardization: Brussels, Belgium, 2023.
  38. EN 14211:2024; Ambient Air—Standard Method for the Measurement of the Concentration of Nitrogen Dioxide and Nitrogen Monoxide by Chemiluminescence. European Committee for Standardization: Brussels, Belgium, 2024.
  39. EN 14625:2024; Ambient Air—Standard Method for the Measurement of the Concentration of Ozone by Ultraviolet Photometry. European Committee for Standardization: Brussels, Belgium, 2024.
  40. Schneider, P.; Castell, N.; Vogt, M.; Dauge, F.R.; Lahoz, W.A.; Bartonova, A. Mapping Urban Air Quality in Near Real-Time Using Observations from Low-Cost Sensors and Model Information. Environ. Int. 2017, 106, 234–247. [Google Scholar] [CrossRef] [Scilit]
  41. U.S. Environmental Protection Agency. How to Evaluate Low-Cost Sensors by Collecting with Federal Reference Method Monitors. Available online: https://www.epa.gov/sites/default/files/2018-01/documents/collocation_instruction_guide.pdf (accessed on 1 December 2025).
  42. Kelly, K.E.; Whitaker, J.; Petty, A.; Widmer, C.; Dybwad, A.; Sleeth, D.; Martin, R.; Butterfield, A. Ambient and Laboratory Evaluation of a Low-Cost Particulate Matter Sensor. Environ. Pollut. 2017, 221, 491–500. [Google Scholar] [CrossRef] [Scilit]
  43. Duvall, R.M.; Long, R.W.; Beaver, M.; Kronmiller, K.; Wheeler, M.; Szykman, V.J. Performance Testing Protocols, Metrics, and Target Values for Fine Particulate Matter Air Sensors; EPA/600/R-20/280; U.S. Environmental Protection Agency: Washington, DC, USA, 2021.
  44. Jayaratne, E.R.; Liu, X.; Thai, P.; Dunbabin, M.; Morawska, L. The Influence of Humidity on the Performance of a Low-Cost Air Particle Mass Sensor and the Effect of Atmospheric Fog. Atmos. Meas. Tech. 2018, 11, 4883–4890. [Google Scholar] [CrossRef] [Scilit]
  45. Crilley, L.R.; Shaw, M.; Pound, R.; Kramer, L.J.; Price, R.; Young, S.; Lewis, A.C.; Pope, F.D. Evaluation of a Low-Cost Optical Particle Counter (Alphasense OPC-N2) for Ambient Air Monitoring. Atmos. Meas. Tech. 2018, 11, 709–720. [Google Scholar] [CrossRef] [Scilit]
  46. Cross, E.S.; Williams, L.R.; Lewis, D.K.; Magoon, G.R.; Onasch, T.B.; Kaminsky, M.L.; Worsnop, D.R.; Jayne, J.T. Use of Electrochemical Sensors for Measurement of Air Pollution: Correcting Interference Response and Validating Measurements. Atmos. Meas. Tech. 2017, 10, 3575–3588. [Google Scholar] [CrossRef] [Scilit]
  47. Wei, P.; Ning, Z.; Ye, S.; Sun, L.; Yang, F.; Wong, K.C.; Westerdahl, D.; Louie, P.K. Impact Analysis of Temperature and Humidity Conditions on Electrochemical Sensor Response in Outdoor Air Quality Monitoring. Sensors 2018, 18, 59. [Google Scholar] [CrossRef] [Scilit]
  48. Breiman, L. Random Forests. Mach. Learn. 2001, 45, 5–32. [Google Scholar] [CrossRef] [Scilit]
  49. Probst, P.; Wright, M.N.; Boulesteix, A.L. Hyperparameters and Tuning Strategies for Random Forest. Wiley Interdiscip. Rev. Data Min. Knowl. Discov. 2019, 9, e1301. [Google Scholar] [CrossRef] [Scilit]
  50. Roberts, D.R.; Bahn, V.; Ciuti, S.; Boyce, M.S.; Elith, J.; Guillera-Arroita, G.; Hauenstein, S.; Lahoz-Monfort, J.J.; Schroder, B.; Thuiller, W.; et al. Cross-Validation Strategies for Data with Temporal, Spatial, Hierarchical, or Phylogenetic Structure. Ecography 2017, 40, 913–929. [Google Scholar] [CrossRef] [Scilit]
  51. Meyer, H.; Reudenbach, C.; Hengl, T.; Katurji, M.; Nauss, T. Improving Performance of Spatio-Temporal Machine Learning Models Using Forward Feature Selection and Target-Oriented Validation. Environ. Model. Softw. 2018, 101, 1–9. [Google Scholar] [CrossRef] [Scilit]
  52. Cressie, N. Statistics for Spatial Data; John Wiley & Sons: New York, NY, USA, 1993. [Google Scholar]
  53. Bista, S.; Rahmati, M.; Eeftens, M.; Deguen, A.; Jacquemin, B.; Bard, D.; Basagana, X.; Slama, R. Relationships Between Fixed-Site Ambient Measurements of Nitrogen Dioxide, Ozone, and Particulate Matter and Personal Exposures in Grand Paris, France: The MobiliSense Study. Int. J. Health Geogr. 2025, 24, 5. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Random Forest calibration performance on held-out test data. Scatter plots show calibrated sensor predictions versus co-located reference measurements for (a) PM2.5, (b) PM10, (c) NO2, and (d) O3. Point density increases from purple (sparse) to yellow (dense). Dashed black line indicates perfect agreement (1:1); solid red line shows the linear regression fit.
Figure 1. Random Forest calibration performance on held-out test data. Scatter plots show calibrated sensor predictions versus co-located reference measurements for (a) PM2.5, (b) PM10, (c) NO2, and (d) O3. Point density increases from purple (sparse) to yellow (dense). Dashed black line indicates perfect agreement (1:1); solid red line shows the linear regression fit.
Atmosphere 17 00335 g001
Figure 2. Random Forest feature importance by pollutant, showing the relative contribution of each predictor variable to calibration accuracy. Features include: raw sensor reading, meteorological variables (temperature, pressure, relative humidity), and temporal indicators (hour of day, month, and weekend: a binary flag indicating Saturday/Sunday versus weekday). Sensor reading dominates for particulate matter, temperature for O3, and temporal features for NO2, reflecting distinct calibration mechanisms across pollutant types.
Figure 2. Random Forest feature importance by pollutant, showing the relative contribution of each predictor variable to calibration accuracy. Features include: raw sensor reading, meteorological variables (temperature, pressure, relative humidity), and temporal indicators (hour of day, month, and weekend: a binary flag indicating Saturday/Sunday versus weekday). Sensor reading dominates for particulate matter, temperature for O3, and temporal features for NO2, reflecting distinct calibration mechanisms across pollutant types.
Atmosphere 17 00335 g002
Figure 3. Distance-based uncertainty growth for calibrated measurements. Each panel shows how prediction uncertainty increases as sensors are deployed farther from reference stations, with the slope indicating the uncertainty growth rate (%/km). Horizontal dashed lines mark 10%, 30%, and 50% uncertainty thresholds. Shaded regions indicate 95% confidence intervals around the fitted model. Filled circles show actual spatial cross-validation RMSE values (leave-one-station-out RF) plotted at the distance between the held-out co-location station and its nearest training station; these empirical points validate the fitted model at the two distances available in the network (1.8 km and 2.6 km). PM10 and O3 show statistically significant distance effects ( p < 0.05 ); PM2.5 and NO2 trends are consistent but not statistically significant at α = 0.05 .
Figure 3. Distance-based uncertainty growth for calibrated measurements. Each panel shows how prediction uncertainty increases as sensors are deployed farther from reference stations, with the slope indicating the uncertainty growth rate (%/km). Horizontal dashed lines mark 10%, 30%, and 50% uncertainty thresholds. Shaded regions indicate 95% confidence intervals around the fitted model. Filled circles show actual spatial cross-validation RMSE values (leave-one-station-out RF) plotted at the distance between the held-out co-location station and its nearest training station; these empirical points validate the fitted model at the two distances available in the network (1.8 km and 2.6 km). PM10 and O3 show statistically significant distance effects ( p < 0.05 ); PM2.5 and NO2 trends are consistent but not statistically significant at α = 0.05 .
Atmosphere 17 00335 g003
Figure 4. Spatial distribution of Sofia’s hybrid air quality monitoring network showing predicted PM10 calibration uncertainty. Black stars indicate reference stations operated by ExEA; white circles mark co-located low-cost sensors (uncertainty ≈ 0%); colored circles show non-co-located sensors with color gradient representing predicted uncertainty increase based on distance from the nearest reference station (green = low uncertainty, red = high uncertainty). Labels show sensor ID and predicted uncertainty percentage. The uncertainty values are calculated using the PM10 growth rate of 5.62%/km, which represents the highest (most conservative) rate among statistically significant pollutants.
Figure 4. Spatial distribution of Sofia’s hybrid air quality monitoring network showing predicted PM10 calibration uncertainty. Black stars indicate reference stations operated by ExEA; white circles mark co-located low-cost sensors (uncertainty ≈ 0%); colored circles show non-co-located sensors with color gradient representing predicted uncertainty increase based on distance from the nearest reference station (green = low uncertainty, red = high uncertainty). Labels show sensor ID and predicted uncertainty percentage. The uncertainty values are calculated using the PM10 growth rate of 5.62%/km, which represents the highest (most conservative) rate among statistically significant pollutants.
Atmosphere 17 00335 g004
Table 1. Calibration model performance on co-located sensor test set.
Table 1. Calibration model performance on co-located sensor test set.
PollutantModelntestR2RMSE (μg/m3)MAE (μg/m3)
PM2.5MLR24080.14711.527.84
Random Forest24080.7456.304.00
PM10MLR97260.14219.2012.94
Random Forest97260.66811.957.43
NO2MLR98820.19916.9413.21
Random Forest98820.53212.949.77
O3MLR75400.53221.5617.48
Random Forest75400.74515.9112.52
Table 2. Random Forest spatial cross-validation results (leave-one-station-out). Results are shown for the full feature set, including temporal features (hour, month, weekend indicator) and for physical features only (temperature, pressure, humidity), to empirically assess the effect of temporal features on spatial calibration transfer.
Table 2. Random Forest spatial cross-validation results (leave-one-station-out). Results are shown for the full feature set, including temporal features (hour, month, weekend indicator) and for physical features only (temperature, pressure, humidity), to empirically assess the effect of temporal features on spatial calibration transfer.
PollutantTest StationR2 (with Temporal)R2 (Without Temporal)RMSE (μg/m3)
PM10Pavlovo0.4460.35616.41
Hipodruma0.2190.09819.78
Mladost0.4650.41611.72
Druzhba0.1740.12718.87
NO2Pavlovo0.3290.14117.74
Hipodruma0.3970.24313.92
Mladost0.2220.03613.98
Druzhba−0.029−0.28218.72
O3Pavlovo0.4370.38524.87
Hipodruma−0.145−0.31128.92
Druzhba0.5650.50121.32
Table 3. Seasonal cross-validation results (leave-one-season-out).
Table 3. Seasonal cross-validation results (leave-one-season-out).
PollutantTest SeasonntrainntestR2RMSE (μg/m3)MAE (μg/m3)
PM2.5Winter12,04240050.22317.139.74
Spring12,37636710.3455.433.93
Summer11,9114136−0.1517.205.21
Fall11,81242350.1678.055.71
PM10Winter50,72414,1150.14329.2417.25
Spring47,58017,2590.22313.589.35
Summer47,85016,9890.17411.938.75
Fall48,36316,476−0.01716.5811.86
NO2Winter51,19014,6910.14720.8515.79
Spring48,59617,285−0.14815.7812.12
Summer48,71117,1700.09611.868.88
Fall49,14616,735−0.01320.2514.73
O3Winter38,37411,890−0.02621.6517.36
Spring37,34612,9180.11424.0119.52
Summer37,31012,9540.46122.3517.34
Fall37,76212,5020.20023.3918.87
Table 4. Distance-based uncertainty growth parameters.
Table 4. Distance-based uncertainty growth parameters.
Pollutantσbase (μg/m3)α (km−1)Growth (%/km)R2p-Value
PM1020.260.05625.620.3790.0066 **
PM2.59.300.05335.330.1770.0819
NO217.470.04244.240.1340.1350
O324.020.03843.840.2600.0307 *
* p < 0.05, ** p < 0.01; Distance range: 1.0–5.6 km ( n = 18 sensors)
Table 5. Comparison of calibration performance with published literature.
Table 5. Comparison of calibration performance with published literature.
StudyPollutantR2 RangeTraining DataNotes
[10]PM2.50.70–0.85Multi-site, monthsDense co-location network
[11]PM2.50.65–0.80Multi-site, yearGeneral calibration model
This studyPM2.50.745Single pair, 24 monthsLimited co-location data
[10]NO20.50–0.70Multi-site, monthsElectrochemical sensors
This studyNO20.5324 pairs, 24 montsConsistent with literature
[11]O30.60–0.75Multi-site, yearTemperature-corrected
This studyO30.7454 pairs, 24 montsStrong T-dependence
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Zhivkov, P.; Fidanova, S. Machine Learning Calibration Transfer for Low-Cost Air Quality Sensors: Distance-Based Uncertainty Quantification in a Hybrid Urban Monitoring Network. Atmosphere 2026, 17, 335. https://doi.org/10.3390/atmos17040335

AMA Style

Zhivkov P, Fidanova S. Machine Learning Calibration Transfer for Low-Cost Air Quality Sensors: Distance-Based Uncertainty Quantification in a Hybrid Urban Monitoring Network. Atmosphere. 2026; 17(4):335. https://doi.org/10.3390/atmos17040335

Chicago/Turabian Style

Zhivkov, Petar, and Stefka Fidanova. 2026. "Machine Learning Calibration Transfer for Low-Cost Air Quality Sensors: Distance-Based Uncertainty Quantification in a Hybrid Urban Monitoring Network" Atmosphere 17, no. 4: 335. https://doi.org/10.3390/atmos17040335

APA Style

Zhivkov, P., & Fidanova, S. (2026). Machine Learning Calibration Transfer for Low-Cost Air Quality Sensors: Distance-Based Uncertainty Quantification in a Hybrid Urban Monitoring Network. Atmosphere, 17(4), 335. https://doi.org/10.3390/atmos17040335

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop