Skip to Content
AxiomsAxioms
  • Article
  • Open Access

28 February 2026

16 Pages

Functional Data Analysis of Particulate Matter (PM10) in Chungbuk, Korea: An Application of Penalized B-Spline Smoothing and Functional Inference

,
,
and
1
Department of Information Statistics, Chungbuk National University, Cheongju 28644, Republic of Korea
2
Department of STAT, DreamCIS Inc., Daejeon 34944, Republic of Korea
*
Author to whom correspondence should be addressed.
†
These authors contributed equally to this work.

Abstract

Functional data analysis (FDA) provides a framework for representing high-frequency or longitudinal observations as smooth functions, enabling principled dimension reduction and feature extraction. We develop a B-spline-based FDA approach with a total variation penalty to model daily PM10 trajectories, with smoothing parameters selected via AIC and BIC. Functional principal component analysis (FPCA) is applied to identify dominant temporal patterns, including overall levels, seasonal deviations, and episodic peaks, while preserving abrupt changes. This methodology allows for flexible and interpretable summaries of complex time series and comparisons across spatial or temporal domains. We apply this framework to 365-day PM10 curves from 28 monitoring stations in Chungcheongbuk-do, Korea, for 2022. The first four principal components capture over 80% of total variation, reflecting winter peaks, early spring fluctuations, and summer troughs. Urban–rural contrasts examined via functional two-sample t-tests and FPCA scores reveal minimal differences at the functional level. This study illustrates how FDA, combined with penalized B-splines, can concisely summarize complex temporal dynamics, quantify dominant patterns, and offer a flexible framework for analyzing environmental time series and other functional datasets. The approach provides a general strategy for understanding temporally structured processes in various scientific fields.

1. Introduction

Functional data analysis (FDA) is a branch of statistics that treats each observational unit as a function, curve, or trajectory evolving over a continuum such as time, space, or wavelength, rather than as a finite-dimensional vector [1,2]. In FDA, the basic data object is a smooth function, observed on a dense or possibly irregular grid, and the primary inferential targets include the mean function, covariance structure, and low-dimensional representations such as functional principal components scores. This perspective has proved particularly powerful in scientific fields where modern sensing technologies routinely generate high-frequency, longitudinal, or image-type measurements; examples include growth curves, spectrometric data, biomechanical trajectories, and environmental time series [1,3]. By explicitly exploiting the smoothness and functional structure of the data, FDA provides a principled framework for regularization, noise reduction, and dimension reduction that is difficult to achieve with purely multivariate methods.
From a methodological standpoint, FDA combines ideas from nonparametric regression, smoothing splines, and functional approximation theory. In particular, basis expansions—such as B-splines, Fourier series, and wavelets—play a central role in representing functional observations as linear combinations of basis functions with a relatively small number of coefficients [1,2]. This representation facilitates the use of penalized likelihood or penalized least squares criteria, where roughness penalties enforce smoothness and prevent overfitting. Among various basis systems, B-splines are widely used in practice due to their excellent numerical properties, local support, and flexibility in handling complex designs and boundary behavior. Recent research has further developed penalized B-spline estimators that incorporate total variation (TV) penalties on derivatives, providing an attractive compromise between smooth approximation and the ability to capture sharp changes or structural breaks in functional relationships [4].
In parallel with these methodological advances, there has been growing interest in applying FDA to environmental and public-health problems, where the underlying processes are naturally viewed as time-varying curves or surfaces [3]. Ambient air pollution, and in particular particulate matter (PM), is a representative example. Particulate matter with aerodynamic diameter less than or equal to 10   μ m (PM10) or 2.5   μ m (PM2.5) has been identified as a major risk factor for a broad range of adverse health outcomes, including respiratory disease, cardiovascular disease, and cancer [5]. The International Agency for Research on Cancer (IARC) classified outdoor air pollution and particulate matter as carcinogenic to humans and emphasized that outdoor air pollution is a leading environmental cause of cancer deaths worldwide [6]. These findings have reinforced the need for accurate statistical methods to characterize temporal patterns in particulate concentrations and to assess their health and policy implications.
Reflecting the accumulating epidemiological and toxicological evidence, the World Health Organization (WHO) issued updated Global Air Quality Guidelines in 2021, substantially tightening recommended exposure limits for PM2.5 and NO2 relative to the 2005 guidelines [7,8]. The new guidelines recommend, for example, that the annual mean concentration of PM2.5 should not exceed 5   μ g / m 3 , a target that is currently violated in many urban areas around the world [8]. The WHO and collaborating societies have emphasized that even relatively low concentrations of particulate matter are associated with increased mortality and morbidity, and that there appears to be no safe threshold below which adverse health effects can be ruled out [8]. In this context, statistical models that can capture the full temporal evolution of pollutant concentrations—including seasonal patterns, short-term peaks, and long-term trends—are essential to inform risk assessment and environmental regulation.
In Korea, public concern over particulate matter has increased markedly since the early 2010s, driven in part by frequent high-concentration episodes and extensive media coverage. Empirical analyses of national monitoring data have documented long-term trends and regional disparities in PM10 concentrations, as well as changes in the frequency and persistence of high-pollution events [9]. For instance, Yeo and Kim [9] showed that, while annual mean PM10 levels have generally decreased since the early 2000s, the occurrence of high-concentration cases remains non-negligible and exhibits distinct spatial patterns across provinces. At a more local scale, several studies have examined the current state of fine dust and its emission sources in specific regions, including Chungcheongbuk-do, highlighting the combined influence of industrial activities, traffic, and transboundary transport on regional air quality [10,11]. These findings underscore the importance of region-specific analyses that reflect local emission structures, meteorological conditions, and topographical features.
Chungcheongbuk-do (Chungbuk) is a landlocked province in the central region of South Korea, characterized by a mix of urban, industrial, and rural areas. Owing to its geographical location and industrial structure, Chungbuk experiences complex interactions between locally emitted pollutants and those transported from other regions or across national borders [10,11]. Previous work on Chungbuk has mainly focused on descriptive analyses of long-term averages, emission inventories, and source apportionment [10]. However, less attention has been paid to the detailed temporal dynamics of PM10 concentrations at daily or hourly scales, such as seasonal variation, intra-annual heterogeneity, and short-term peak behaviors. Given the increasing policy interest in region-specific mitigation strategies—for example, targeted emission controls and early-warning systems—there is a clear need for statistical methodologies that can summarize, compare, and model these complex temporal patterns in a systematic manner.
From a statistical perspective, time series of pollutant concentrations can naturally be viewed as functional data when measurements are collected on a fine and regular grid, such as daily or hourly monitoring records over multiple years. In this setting, each year (or season) can be regarded as a smooth curve representing the trajectory of PM10 concentrations, and FDA tools can be deployed to explore the main modes of temporal variation and to perform inference on their determinants [1,2]. Functional principal component analysis (FPCA), in particular, provides a flexible and interpretable way to decompose temporal variability into orthogonal components that capture, for example, overall level, seasonal amplitude, and timing of high-pollution episodes. This approach has been successfully applied in various environmental and epidemiological contexts, including analyses of mortality time series and air quality indicators [3,12]. By combining FPCA with penalized B-spline representations and appropriate roughness penalties, one can obtain smooth yet data-adaptive estimates of these components, even in the presence of irregular sampling or measurement noise [2,4].
The increasing visibility of FDA in both theory and applications is also reflected in recent developments in the literature, including functional time series analysis, functional regression, and functional principal component analysis under complex dependence structures [12]. These works demonstrate that FDA methods can be effectively integrated with modern dependence modeling, such as copula-based dynamic conditional correlation models, to analyze high-dimensional and temporally dependent functional data. Against this backdrop, the present study aims to contribute to the growing FDA literature by developing and applying a B-spline based functional principal component framework tailored to the analysis of PM10 concentration curves in Chungcheongbuk-do. Specifically, we focus on (i) constructing smooth functional representations of daily PM10 trajectories using penalized B-splines, (ii) extracting dominant modes of temporal variation via FPCA, and (iii) interpreting the resulting components in terms of seasonal patterns, high-concentration episodes, and long-term trends. The proposed approach not only yields a parsimonious description of complex temporal dynamics but also provides a basis for subsequent modeling and policy-relevant interpretation in the context of air quality management. Although the methods themselves are well-established, the primary contribution of this study lies in applying them to PM10 data in a novel manner, providing new insights into seasonal patterns, high-concentration episodes, and long-term trends.

3. Methodology

We briefly describe the methodological framework used in this study. The analysis consists of four main steps: (i) representation of daily PM10 curves as functional observations; (ii) penalized B-spline smoothing with total variation (TV) penalty; (iii) functional two-sample testing; and (iv) functional principal component analysis (FPCA). Throughout, the focus is on tools that are flexible, computationally stable, and well-suited to environmental time-series data.

3.1. Functional Representation

Let Y i ( t j ) denote the observed PM10 concentration at station i ( i = 1 , … , n ) on day t j ( j = 1 , … , T ). We assume that these discrete measurements arise from an underlying smooth function X i ( t ) observed with measurement error:
Y i ( t j ) = X i ( t j ) + ε i j , ε i j ∼ i . i . d . ( 0 , σ 2 ) ,
where X i ( t ) is defined on the continuous domain t ∈ [ 0 , 1 ] (after linear rescaling of the calendar year). Note that time is rescaled to [0, 1] for convenience, with t = 0 and t = 1 representing 1 January and 31 December, so seasonal interpretation is preserved.
The goal of smoothing is to obtain stable estimates X ^ i ( t ) that retain the important temporal features of the original trajectories.

3.2. B-Spline Basis Expansion

Each smooth curve X i ( t ) is represented using a finite number of cubic B-spline basis functions:
X i ( t ) ≈ ∑ k = 1 K β i k B k ( t ) = B ( t ) ⊤ β i ,
where B ( t ) = { B 1 ( t ) , … , B K ( t ) } ⊤ denotes the spline basis and β i = ( β i 1 , … , β i K ) ⊤ are the station-specific coefficients. Cubic B-splines ensure continuity of the first and second derivatives and offer local support and numerical stability [16].

3.3. Penalized Spline Smoothing with TV Penalty

Following [4], we estimate β i by solving the penalized least squares problem:
min β i ∑ j = 1 T Y i ( t j ) − B ( t j ) ⊤ β i 2 + λ TV d 3 d t 3 X i ( t ) ,
where λ ≥ 0 controls the degree of smoothness and TV ( · ) denotes the total variation functional.
For a function f defined on [ 0 , 1 ] , total variation is
TV ( f ) = sup P ∑ m = 0 | P | − 1 f ( x m + 1 ) − f ( x m ) ,
where the supremum is taken over all finite partitions P = { 0 = x 0 < ⋯ < x | P | = 1 } .
In the cubic B-spline setting, penalizing the total variation in the third derivative is closely related to automatic knot selection and yields smooth yet adaptive estimates [4].
We select λ by minimizing information criteria:
AIC ( λ ) = T log σ ^ 2 ( λ ) + 2 edf ( λ ) , BIC ( λ ) = T log σ ^ 2 ( λ ) + log ( T ) edf ( λ ) ,
where edf ( λ ) denotes the effective degrees of freedom. In practice, Akaike information criterion (AIC; [18]) tends to favor less oversmoothing than the Bayesian information criterion (BIC; [19]), which is beneficial for pollution data exhibiting episodic peaks. In the present setting with densely observed daily series, information-criterion-based selection provides a computationally efficient and stable choice, and allows direct comparison between AIC and BIC in terms of smoothness and peak preservation. Note that alternative criteria, such as cross-validation, could also be used to select the smoothing parameter.

3.4. Exploratory Functional Quantities

Given smoothed curves { X ^ i ( t ) } , the mean and variance functions are
X ¯ ( t ) = 1 n ∑ i = 1 n X ^ i ( t ) , Var ( t ) = 1 n − 1 ∑ i = 1 n X ^ i ( t ) − X ¯ ( t ) 2 .
The covariance surface is estimated as
Cov ( s , t ) = 1 n − 1 ∑ i = 1 n X ^ i ( s ) − X ¯ ( s ) X ^ i ( t ) − X ¯ ( t ) .

3.5. Functional Two-Sample t-Test

Suppose the stations are divided into two groups, Group 1 (urban) and Group 2 (rural), with sample sizes n 1 and n 2 . The pointwise test statistic proposed in [2] is
T ( t ) = X ¯ 1 ( t ) − X ¯ 2 ( t ) Var 1 ( t ) / n 1 + Var 2 ( t ) / n 2 ,
where X ¯ g ( t ) and Var g ( t ) denote the group-specific mean and variance functions.
To obtain global inference, we compute the maximum-type statistic max t T ( t ) and evaluate its null distribution using permutation-based resampling, thus preserving the functional dependence.

3.6. Functional Principal Component Analysis (FPCA)

FPCA decomposes temporal variability through the eigen-decomposition of the covariance operator. The eigenfunctions { ξ k ( t ) } solve
∫ 0 1 Cov ( s , t ) ξ k ( t ) d t = ρ k ξ k ( s ) ,
where ρ k are the eigenvalues ordered ρ 1 ≥ ρ 2 ≥ ⋯ ≥ 0 .
Each curve admits the Karhunen–Loève representation:
X i ( t ) = μ ( t ) + ∑ k = 1 ∞ α i k ξ k ( t ) , α i k = ∫ 0 1 X i ( t ) − μ ( t ) ξ k ( t ) d t ,
where α i k are uncorrelated principal component scores with Var ( α i k ) = ρ k . Note that Table 1 provides a brief summary of the key symbols used in the FPCA framework.
Table 1. Summary of key notation used in the FPCA framework.
In practice, only the first few components explaining the majority of total variance are retained, yielding a parsimonious but interpretable description of temporal behavior in PM10 concentrations [1,12].
Overall, the methodological framework presented in this section adopts an FDA perspective to treat PM10 concentrations as smooth temporal processes rather than discrete observations. In contrast to conventional time-series approaches, which typically focus on short-term dependence structures or pointwise temporal dynamics, the FDA framework represents each station’s record as a continuous trajectory defined over a common domain.
In contrast, the total variation penalty adopted in this study is naturally aligned with the methodological objective of capturing both structural commonality and localized heterogeneity across multiple functional trajectories. By allowing for roughness to be concentrated at a limited number of time points, the total variation penalty yields piecewise-smooth representations in which shared temporal patterns are preserved while station-specific abrupt changes are retained when supported by the data. This feature is particularly relevant for PM10 concentrations, where abrupt concentration increases associated with episodic environmental events may differ across regions and should not be uniformly penalized as roughness. Consequently, the total variation penalty provides a more appropriate roughness penalty than derivative-based alternatives for modeling collections of functional curves with both common and region-specific temporal features.
Compared with GAM-based analyses that emphasize regression relationships and conditional mean effects at individual time points, the proposed approach facilitates direct comparison of entire temporal profiles across monitoring stations. By combining penalized B-spline smoothing with FPCA, the framework provides a parsimonious representation of dominant modes of temporal variability and establishes a coherent basis for functional inference on seasonal patterns and regional differences in air pollution dynamics.

4. Data Analysis: PM10 Concentrations in Chungbuk

In this section, we apply the proposed functional data analysis framework to daily PM10 concentrations monitored across Chungcheongbuk-do (Chungbuk), Korea. The primary objective is to describe temporal variation, evaluate potential differences between urban and rural environments, and identify dominant modes of variability through functional principal component analysis (FPCA).
Air-quality measurements were obtained from AirKorea, the national monitoring network that records several major atmospheric pollutants, including SO2, NO2, O3, CO, PM10, and PM2.5. The PM10 values are determined using either gravimetric methods or the β -ray absorption method, and results are reported in μ g / m 3 . During 2022, a total of 30 monitoring stations operated within the province. To ensure reliability of inference, stations exhibiting more than 15% missing daily values (equivalent to at least 55 missing days) were removed. After this screening step, 28 stations remained and Figure 1 illustrates the administrative regions in which these stations are located. In Figure 1, thick boundary lines indicate city and county boundaries, while thin lines represent lower-level administrative units (eup, myeon, and dong). Gray-shaded regions indicate administrative areas where monitoring stations are located. Meanwhile, the total number of missing days across all stations was only 36. Because the proportion of missingness was small, these values were omitted rather than imputed.
Figure 1. Administrative regions of the 28 PM monitoring stations. Thick boundary lines indicate city and county boundaries, and thin lines represent lower-level administrative units (eup, myeon, and dong). Gray shading highlights regions with monitoring stations.
Korea typically experiences pronounced seasonal fluctuations in particulate concentrations. For this reason, the complete calendar year from January to December 2022 was analyzed, allowing us to capture recurring features such as winter peaks, spring transition periods, summer troughs, and late-autumn increases. Note that 2022 was primarily selected because it was the year in which data collection and analysis were initiated. Historical records from 2018 to 2022 indicate that annual mean PM10 concentrations and key meteorological variables (temperature, wind speed, and precipitation) were generally consistent with previous years, as shown in Supplemental Figure S1, suggesting that 2022 was not anomalous and is suitable for examining typical seasonal patterns.
Figure 2 displays the daily PM10 trajectories for all 28 stations. Most stations share a broadly similar temporal pattern: concentrations rise during winter and early spring, decline substantially during summer, and increase again toward the end of the year. Figure 3 summarizes these dynamics by showing the overall mean curve, in which several pronounced short-term surges are visible, particularly in January–March, May, and November.
Figure 2. Daily PM10 concentrations for 28 monitoring stations in Chungbuk over 2022. Each curve represents one station.
Figure 3. Overall mean daily PM10 concentration across the 28 stations.
To investigate spatial heterogeneity related to population density, monitoring sites were classified into urban (population ≥ 50,000) and rural (population < 50,000) groups. The urban group consists of stations located in Cheongju, Chungju, and Jecheon, whereas the remaining sites, such as those in Danyang, Goesan, Yeongdong, Eumseong, Jeungpyeong, Jincheon, and Okcheon, are categorized as rural. Supplemental Figure S2 presents daily curves for the urban stations with the overall mean superimposed. Several locations, including Songjeong in Cheongju and Jangrak in Jecheon, frequently record concentrations above the provincial average. By contrast, many rural stations remain slightly below the mean throughout most of the year; for instance, Danseong in Danyang consistently exhibits the lowest levels, as shown together with other rural stations in Supplemental Figure S3.
The raw daily values were then converted to smooth functional trajectories using cubic B-spline bases combined with penalized smoothing. The choice of smoothing parameter was determined by comparing the Akaike information criterion (AIC) and the Bayesian information criterion (BIC). Supplemental Figures S4 and S5 illustrate the effect of these two criteria. In general, AIC allows for greater local flexibility, thereby preserving episodic peaks that may be environmentally meaningful, whereas BIC imposes stronger regularization and occasionally suppresses such structure, particularly at stations with relatively low variability. Considering this behavior, subsequent analyses were conducted using AIC-based smoothed curves. Figure 4 displays the collection of smoothed curves for all stations, together with their corresponding mean functions.
Figure 4. Functional curves for all stations under AIC (top panel) and BIC (bottom panel) criteria, with mean curves shown in black.
Visual inspection of the sample mean functions in Figure 5 (left panel) suggests that urban stations tend to maintain higher levels than the provincial mean, while rural stations remain slightly below it. This pattern may reflect the influence of traffic and industrial emissions, which are generally higher in urban regions. However, to formally examine whether these apparent differences are statistically meaningful, a permutation-based functional two-sample t-test was implemented. The resulting test statistic is plotted in Figure 6. Although pointwise values occasionally exceed critical thresholds during certain time intervals, the global maximum statistic—used as the formal decision criterion—does not surpass the permutation-based critical value. This indicates that, despite the apparent differences in the mean curves shown in Figure 5 (left panel), the largely overlapping 95% confidence bands in Figure 5 (right panel) suggest that the patterns observed in the mean curves are not statistically significant.
Figure 5. Sample mean functions (left panel): overall (solid), urban (blue dashed), rural (red dashed). 95% confidence bands (right panel): urban (blue) and rural (red).
Figure 6. Functional two-sample t-test comparing urban and rural mean functions.
Further insight into temporal variability was obtained by examining the estimated variance function and covariance surface, shown in Figure 7. Variance is largest from late winter through spring and again in late autumn, indicating that inter-station variability is driven primarily by periods in which abrupt changes occur. Outside these intervals, temporal dependence is relatively weak, consistent with the smoother summer trajectories.
Figure 7. Variance function (left panel) and covariance surface (right panel) of PM10 functional curves.
Finally, FPCA was applied to the smoothed curves in order to extract the main modes of variation. Supplemental Figure S6 presents the first four principal components, along with the mean curve and corresponding perturbations. The first component, which explains 58.4% of the total variation, reflects overall elevation of pollution concentrations, particularly during winter and early spring. The second component, accounting for 10.6%, captures contrasts within the winter season, highlighting differing magnitudes of peak intensity. The third component (7.7%) is associated with fluctuations concentrated in May and June, while the fourth component (5.5%) represents relatively small deviations around the summer minimum. For comparison, FPCA based on BIC-smoothed curves (Supplemental Figure S7) yielded a similar set of principal components but with a reduced proportion of variance explained by the leading modes (58.4%, 10.6%, 7.7%, and 5.5% for PC1–PC4), reflecting the stronger regularization imposed by BIC. Because the primary goal of this study is to characterize temporal patterns and variability in daily PM10 concentrations, preserving such variability is essential; therefore, AIC-based smoothing was adopted for all subsequent functional analyses. Together, these components provide a concise but informative summary of the temporal structure underlying PM10 dynamics in Chungbuk and form the basis for interpretation in the subsequent discussion.
To further explore spatial heterogeneity, FPCA scores were compared between urban and rural stations (Supplemental Figure S8; Supplemental Table S1). For PC1 (overall pollution level), urban stations exhibit substantially higher median scores than rural stations, indicating generally elevated PM10 trajectories in urban areas. PC2 (seasonal deviations, particularly winter peaks) shows some differences in range and distribution between groups, with a few urban stations displaying low scores, reflecting localized variability, and lower median and mean and slightly larger standard deviation in urban stations. PC3 (late spring fluctuations) includes a few outliers in the urban group, but median scores are similar across groups, with mean and standard deviation comparable across groups, suggesting limited contribution to group-level differences. PC4 (minor summer deviations) shows largely overlapping distributions, indicating that this mode primarily reflects individual station-specific patterns rather than systematic group differences. Importantly, although PC3 and PC4 exhibit some variability across stations, formal statistical testing did not detect significant differences between urban and rural groups for these components. Together, these results indicate that the main regional differences are captured in PC1 and PC2, while PC3 and PC4 largely contextualize the null result from the formal functional two-sample test, confirming that overall temporal trajectories are broadly similar across urban and rural sites.

5. Discussion and Conclusions

This study applied a functional data analysis framework to daily PM10 concentrations collected from 28 atmospheric monitoring stations in Chungcheongbuk-do (Chungbuk), Korea. By representing each monitoring record as a smooth functional curve, we were able to examine seasonal behavior, assess potential differences between urban and rural regions, and summarize dominant temporal patterns through functional principal component analysis (FPCA). Viewing the pollution series as continuous trajectories rather than discrete observations provided a unified lens through which the temporal evolution of air quality could be interpreted in a more coherent and informative manner.
The empirical results revealed a pronounced seasonal cycle consistent with prior domestic and international findings. Concentrations were highest during late winter and early spring, declined markedly during the summer months, and increased again toward late autumn. The mean curve showed several short-lived but substantial surges, particularly in January–March, May, and November, indicating episodic events that likely reflect interactions among meteorological conditions, local emissions, and transported pollutants. When stations were partitioned into urban and rural groups, descriptive comparisons suggested somewhat elevated levels in urban locations throughout much of the year. Nevertheless, the permutation-based functional t-test indicated that these differences were not statistically significant once the full temporal profile was considered. This finding suggests that regional background influences may exert a strong homogenizing effect, reducing contrasts that might otherwise arise solely from local urbanization or traffic-related sources. Such an interpretation is broadly consistent with earlier studies highlighting the role of large-scale transport and synoptic weather patterns in shaping particulate dynamics.
FPCA offered an interpretable decomposition of temporal variability. The first principal component, explaining more than half of the total variation, primarily captured the overall level of pollution and its pronounced elevation during winter and early spring. Subsequent components isolated more localized phenomena, including contrasts within the winter period, variability concentrated in late spring, and subtle fluctuations around summer minima. Together, these components constitute a compact set of indices that summarize seasonal structure in a manner that is both statistically efficient and substantively meaningful. They also provide a natural foundation for future work that may seek to link temporal features of PM10 to meteorological covariates, emission sources, or health outcomes.
From a methodological perspective, the analysis illustrates how FDA can complement traditional time-series approaches. Representing the data as curves allows for the preservation of continuous-time information and avoids the need to focus exclusively on isolated time points or summary averages. The penalized B-spline estimator with total variation penalty proved particularly effective in balancing smoothness with the ability to capture episodic peaks, thereby reducing the likelihood of oversmoothing features that may hold environmental significance. Functional hypothesis testing enabled global comparison of temporal patterns across groups, while FPCA exposed low-dimensional structure that can be readily interpreted and incorporated into subsequent modeling efforts. Rather than replacing existing regression or spatio-temporal frameworks, FDA provides additional structure that can be combined with them to improve interpretability and, potentially, predictive accuracy.
The findings also carry several practical implications. The absence of a statistically significant distinction between urban and rural curves implies that mitigation strategies focused exclusively on urban centers may overlook broader regional drivers. Coordinated provincial policies, early-warning systems tailored to seasonal peaks, and public risk communication strategies may therefore be more effective than measures targeted only at specific cities. Moreover, FPCA-derived scores could be used as concise indicators of abnormal seasonal behavior, enabling earlier identification of periods in which unusually severe winter or spring pollution is developing. These results provide practical guidance for environmental monitoring agencies: the identification of dominant temporal patterns and seasonal deviations can help optimize monitoring station placement, support long-term trend analysis, and inform early warning and public communication strategies.
At the same time, several limitations warrant consideration. The analysis was restricted to a single year (2022), limiting the ability to evaluate long-term structural changes or policy effects across multiple years. Meteorological factors and co-pollutants, which are known to influence PM dynamics, were not explicitly included in the current modeling framework. As a result, the identified temporal patterns should be interpreted as descriptive characteristics of PM10 variability rather than as evidence of causal relationships, and any policy or environmental implications should be considered with appropriate caution. Importantly, these limitations may influence the FPCA and functional test results. Because the analysis is restricted to a single year, unusual day-to-day variability or rare events could disproportionately affect seasonal components (PC2–PC4), and the absence of meteorological covariates may lead to slight overestimation of short-term variability. Overall, the main temporal patterns appear robust, but specific components may be sensitive to single-year anomalies or unmeasured confounding factors. In addition, smoothing parameters were selected using information criteria; alternative procedures such as cross-validation or Bayesian methods may provide different perspectives on uncertainty and model selection. Addressing these issues in future work would enhance the robustness of conclusions and yield a more comprehensive picture of pollution dynamics.
Looking ahead, several natural extensions emerge. Functional regression models could be employed to relate PM10 curves to meteorological drivers or health indicators, offering a more explicit link between exposure trajectories and outcomes. Spatio-functional approaches would allow for joint modeling of temporal curves across monitoring locations, capturing both spatial correlation and temporal evolution. Extending the framework to multiple pollutants would facilitate analysis of joint exposure processes. Specially, extending FDA to multi-year data using a year-as-function approach could capture inter-annual variability and long-term trends, providing further insight into temporal stability and episodic events across years. These directions collectively underline the versatility of FDA as a tool for environmental monitoring and public-health assessment.
In conclusion, the study demonstrates that functional data analysis provides a powerful and insightful approach for investigating the temporal behavior of PM10 concentrations in Chungbuk. By combining penalized spline smoothing, functional hypothesis testing, and FPCA, we obtained a parsimonious yet detailed representation of seasonal variation, episodic peaks, and between-station heterogeneity. As air pollution monitoring networks continue to expand and data become more granular, the functional perspective is likely to assume an increasingly important role in statistical analyses that inform environmental policy and health-related decision-making.

Supplementary Materials

The following supporting information can be downloaded at https://www.mdpi.com/article/10.3390/axioms15030170/s1, Figure S1: Daily mean PM10 concentrations by year (2018–2022). The y-axis is fixed across panels to facilitate direct comparison of temporal patterns; Figure S2: Daily PM10 concentrations for urban stations (blue) with overall mean (gray) superimposed; Figure S3: Daily PM10 concentrations for rural stations (red) with overall mean (gray) superimposed; Figure S4: Functional transformation of PM10 data for urban stations. Red: AIC-based smoothing; blue: BIC-based smoothing; Figure S5: Functional transformation of PM10 data for rural stations. Red: AIC-based smoothing; blue: BIC-based smoothing; Figure S6: The first four functional principal components selected based on AIC, with mean ± principal component effects. The x-axis represents Date, and the y-axis represents the principal component function. The + and − symbols indicate the direction of each principal component, with + corresponding to the positive direction and − to the negative direction; Figure S7: The first four functional principal components selected based on BIC, with mean ± principal component effects. The x-axis represents Date, and the y-axis represents the principal component function. The + and − symbols indicate the direction of each principal component, with + corresponding to the positive direction and − to the negative direction; Figure S8: Boxplots of FPCA scores for the first four principal components, comparing urban (skyblue) and rural (pink) stations; Table S1: Summary statistics of FPCA component scores for the first four principal components (PC1–PC4).

Author Contributions

Conceptualization, E.-J.L. and J.H.L.; methodology, E.-J.L. and J.-H.J.; software, E.-J.L.; validation, E.-J.L. and H.-J.J.; formal analysis, E.-J.L. and J.H.L.; investigation, J.H.L.; resources, J.H.L.; data curation, E.-J.L. and J.H.L.; writing—original draft preparation, E.-J.L. and J.-H.J.; writing—review and editing, E.-J.L.; visualization, E.-J.L. and J.H.L.; supervision, J.-H.J.; project administration, H.-J.J.; funding acquisition, E.-J.L., H.-J.J. and J.-H.J. All authors have read and agreed to the published version of the manuscript.

Funding

This research was supported by Chungbuk National University Glocal30 project (2024). The research of Eun-Ji Lee was supported by the Basic Science Research Program through the National Research Foundation of Korea (NRF) funded by the Ministry of Education (RS-2024-00406939). The research of Hee-Jung Jee was supported by the National Research Foundation of Korea (NRF) grant funded by the Korea government (MSIT) (RS-2024-00354881). The research of Jae-Hwan Jhong was supported by the National Research Foundation of Korea (NRF) grant funded by the Korea government (MSIT) (RS-2024-00342014).

Institutional Review Board Statement

Not applicable.

Data Availability Statement

The data analyzed in this study are publicly available from the AirKorea website: https://www.airkorea.or.kr/web/ (accessed on 15 February 2023).

Conflicts of Interest

Author Ji Hyeon Lee was employed by the company DreamCIS Inc. All authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.

References

  1. Ramsay, J.O.; Silverman, B.W. Functional Data Analysis, 2nd ed.; Springer Series in Statistics; Springer: New York, NY, USA, 2005. [Google Scholar] [CrossRef] [Scilit]
  2. Ramsay, J.O.; Hooker, G.; Graves, S. Functional Data Analysis with R and MATLAB; Use R! Springer: New York, NY, USA, 2009. [Google Scholar] [CrossRef] [Scilit]
  3. Aneiros, G.; Horová, I.; Hušková, M.; Vieu, P. On functional data analysis and related topics. J. Multivar. Anal. 2022, 189, 104861. [Google Scholar] [CrossRef] [Scilit]
  4. Jhong, J.H.; Koo, J.Y.; Lee, S.W. Penalized B-spline estimator for regression functions using total variation penalty. J. Stat. Plan. Inference 2017, 184, 77–93. [Google Scholar] [CrossRef] [Scilit]
  5. Jang, A.S. Impact of particulate matter on health. J. Korean Med Assoc. 2014, 57, 763–768. [Google Scholar] [CrossRef] [Scilit]
  6. International Agency for Research on Cancer. Outdoor Air Pollution a Leading Environmental Cause of Cancer Deaths. Press Release No. 221. 2013. Available online: https://www.iarc.who.int/wp-content/uploads/2018/07/pr221_E.pdf (accessed on 5 February 2023).
  7. World Health Organization. WHO Global Air Quality Guidelines: Particulate Matter (PM2.5 and PM10), Ozone, Nitrogen Dioxide, Sulfur Dioxide and Carbon Monoxide; World Health Organization: Geneva, Switzerland, 2021; Available online: https://www.who.int/publications/i/item/9789240034228 (accessed on 5 February 2023).
  8. Hoffmann, B.; Boogaard, H.; de Nazelle, A.; Andersen, Z.J.; Abramson, M.; Brauer, M.; Brunekreef, B.; Forastiere, F.; Huang, W.; Kan, H.; et al. WHO Air Quality Guidelines 2021–Aiming for healthier air for all: A joint statement by medical, public health, scientific societies and patient representative organisations. Int. J. Public Health 2021, 66, 1604465. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  9. Yeo, M.J.; Kim, Y.P. Trends of the PM10 concentrations and high PM10 concentration cases in Korea. J. Korean Soc. Atmos. Environ. 2019, 35, 249–264. [Google Scholar] [CrossRef] [Scilit]
  10. Lee, J.; Lee, S. Analysis of the current state of fine dust and pollution sources in Chungcheongbuk-do. J. Inst. Constr. Technol. 2021, 40, 35–43. (In Korean) [Google Scholar]
  11. Allabakash, S.; Lim, S.; Chong, K.S.; Yamada, T.J. Particulate matter concentrations over South Korea: Impact of meteorology and other pollutants. Remote Sens. 2022, 14, 4849. [Google Scholar] [CrossRef] [Scilit]
  12. Kim, J.M. Copula dynamic conditional correlation and functional principal component analysis of COVID-19 mortality in the United States. Axioms 2022, 11, 619. [Google Scholar] [CrossRef] [Scilit]
  13. Dominici, F.; McDermott, A.; Zeger, S.L.; Samet, J.M. On the use of generalized additive models in time-series studies of air pollution and health. Am. J. Epidemiol. 2002, 156, 193–203. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  14. Cameletti, M.; Lindgren, F.; Simpson, D.; Rue, H. Spatio-temporal modeling of particulate matter concentration through the SPDE approach. AStA Adv. Stat. Anal. 2013, 97, 109–131. [Google Scholar] [CrossRef] [Scilit]
  15. Morris, J.S. Functional regression. Annu. Rev. Stat. Its Appl. 2015, 2, 321–359. [Google Scholar] [CrossRef] [Scilit]
  16. De Boor, C. A Practical Guide to Splines; Springer: New York, NY, USA, 1978. [Google Scholar]
  17. Liu, X.; Yang, L.; Chen, L. Functional time series modeling with dependence: A survey. Axioms 2023, 12, 433. [Google Scholar] [CrossRef] [Scilit]
  18. Bozdogan, H. Model selection and Akaike’s information criterion (AIC): The general theory and its analytical extensions. Psychometrika 1987, 52, 345–370. [Google Scholar] [CrossRef] [Scilit]
  19. Schwarz, G. Estimating the dimension of a model. Ann. Stat. 1978, 6, 461–464. [Google Scholar] [CrossRef] [Scilit]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Article Metrics

Citations

Article Access Statistics

Multiple requests from the same IP address are counted as one view.