1. Introduction
In recent decades, the rapid development of computing technologies and data acquisition systems has led to a substantial increase in the volume and dimensionality of observed data across a wide range of applied fields, including astronomy, physics, biomedicine, and telecommunications. Modern measurement systems operate at high sampling rates, providing fine temporal and spatial resolution. At the same time, the growth of observation density inevitably amplifies the impact of noise caused by external interference, hardware instability, and random environmental fluctuations. As a result, observed data typically represent a superposition of a weak informative signal and a dominant noise component, while truly informative observations constitute only a small fraction of the entire sample.
In many practical situations, the underlying signal exhibits a pronounced sparsity property: only a limited number of components carry significant information, whereas the majority of observations remain close to zero and can be treated as noise. Such sparse structures naturally arise in high-dimensional statistical problems, where informative events manifest themselves as rare and localized deviations. Sparse signal models have proven effective in a variety of applications, including:
Astronomical observations, where rare astrophysical events such as flares, gamma-ray bursts, or gravitational wave signatures appear against a dominant background noise;
Biomedical diagnostics, where abnormal physiological activity is reflected by infrequent pulses embedded in quasi-stationary biosignals such as ECG or EEG recordings;
Telecommunications, where occasional packet losses or excessive delays indicate network congestion or failures amid predominantly stable transmission conditions;
Industrial monitoring systems, where most sensor readings correspond to normal operation, while rare deviations signal faults or emergency situations.
Round-Trip Time Analysis as a Motivating Example. In modern distributed systems and telecommunication networks, monitoring packet delivery latency, commonly measured via round-trip time (RTT), plays a central role in performance assessment and anomaly detection. For each transmitted packet, RTT is defined as the time elapsed between sending the packet and receiving the corresponding acknowledgment. Under normal operating conditions, RTT values fluctuate within a narrow range around a baseline level, whereas rare, abnormally large delays often indicate transient congestion, queueing effects, or network malfunctions; see, for example, ref. [
1].
Accordingly, the sequence can be viewed as a noisy observation vector in which most entries are close to a baseline value, while a small number of observations correspond to anomalous events. After suitable centering, for instance by subtracting a robust location estimate such as the median, the data naturally conform to the sparse observation model studied in this paper. In this context, thresholding procedures serve as a simple and computationally efficient mechanism for suppressing noise while retaining rare, informative deviations associated with abnormal network delays.
RTT-based measurements and congestion detection mechanisms have been actively studied in the networking literature. Existing approaches focus primarily on protocol design, congestion control, and empirical performance evaluation; see, e.g., refs. [
2,
3,
4]. In contrast, the present paper addresses a complementary
theoretical problem: we study the asymptotic risk properties of thresholding procedures in sparse stochastic observation models, with RTT data serving as a motivating example rather than the only application domain.
Thresholding and Risk Analysis. Among various approaches to sparse signal recovery, thresholding methods remain particularly attractive due to their conceptual simplicity and strong theoretical guarantees. The central problem in thresholding consists in selecting an appropriate threshold value that balances noise suppression and signal preservation. A natural quantitative criterion for evaluating this trade-off is the mean-square risk of the thresholded estimator.
The asymptotic behavior of risk functionals and their empirical estimates has been extensively studied for a wide range of thresholding and shrinkage procedures; see, for instance, refs. [
5,
6,
7,
8,
9,
10,
11,
12,
13,
14,
15]. These works establish consistency, asymptotic normality, and minimax optimality properties under various sparsity and dependence assumptions.
The present paper contributes to this line of research by providing (i) an upper bound on the mean-square risk at the theoretically optimal threshold under an asymptotic sparsity regime, and (ii) asymptotic distributional results for the empirical risk estimate, including a central limit theorem and a strong law of large numbers.
The remainder of the paper is organized as follows.
Section 2 introduces the sparse stochastic observation model and the thresholding framework.
Section 3 studies the theoretical properties of the mean-square risk and its empirical estimate. In particular,
Section 3.1 derives an upper bound for the risk at the theoretically optimal threshold, while
Section 3.2 establishes asymptotic results for the empirical risk estimate, including a central limit theorem and a strong law of large numbers.
Section 4 presents a numerical illustration based on real RTT data. Concluding remarks are given in
Section 5.
2. Data Model
Consider the observed data vector
where
are the true (unknown) signal values, and
are independent Gaussian random variables modeling additive noise. Thus, the random variables
represent noisy measurements of the underlying signal. Throughout the paper, we adopt the standard additive white Gaussian noise (AWGN) model. In applied interpretations, external interference may be viewed as one of the physical sources contributing to this effective Gaussian perturbation, but no separate interference model is assumed.
To describe the signal structure, we assume that
where
are fixed (deterministic) amplitude coefficients, and
are random indicators independent of each other and of
, taking the value 1 with probability
and 0 with probability
. The parameter
may depend on the sample size
N and characterizes the proportion of nonzero signal components. Here, the random indicators
model the presence or absence of an informative signal component at position
i. The Bernoulli distribution reflects the assumption that such informative events are rare and occur sporadically, without a predetermined structure or fixed location within the observation vector. This modeling choice is natural in applications where anomalous events appear unpredictably in time or space.
The coefficients
represent the amplitudes of these rare signal components. They are treated as fixed but unknown quantities, allowing for heterogeneous magnitudes of anomalies. In the context of network delay analysis,
correspond to excess delays caused by transient congestion or queueing effects. The quantity
describes the number of active (non-zero) signal elements and has a binomial distribution
. We assume that the quantity
tends to zero as
N increases. This situation describes a case of asymptotic sparsity, when the number of informative components constitutes a very small fraction of the entire data vector. For the r.v.
, we introduce the corresponding notation
.
Thus, this model describes a noisy observation vector in which significant (non-zero) components are extremely rare and randomly distributed across positions. In other words, informative signal elements do not form a stable structure and can appear in any coordinate of the vector with a low probability . This assumption reflects the typical behavior of real-world sparse data encountered in signal recovery problems.
For further analysis, we introduce a quantitative condition for sparsity. Let there exist a parameter
such that
This condition formalizes an asymptotic sparsity regime in which the expected number of informative components grows sublinearly with respect to the sample size. The parameter quantifies the degree of sparsity: larger values of correspond to rarer signal occurrences. Such regimes naturally arise in high-frequency monitoring systems, where the observation horizon increases while anomalous events remain infrequent.
In other words, the mathematical expectation of the number of nonzero components grows slower than . This condition defines the degree of signal sparsity and will serve as a basic assumption in subsequent theoretical results.
Note that this stochastic structure directly corresponds to data on packet delivery delays in network systems. If we set , where is the observed transit time of the i-th packet, the centering is performed using a robust location estimate. Then typical delay values after centering are close to zero and are described by the noise component , while rare anomalous bursts correspond to nonzero values . Thus, the model can be interpreted as a scheme for detecting rare, anomalously high network delays against a background of Gaussian noise.
To reconstruct the true components
from the observed
, the threshold function is used
depending on the threshold parameter
. The corresponding estimate of the true signal has the form
This type of threshold function provides continuous truncation of small coefficients and is widely used, for example, in wavelet analysis of signals.
The function
corresponds to the classical soft-thresholding rule (
Figure 1). It continuously shrinks small-magnitude observations toward zero while preserving larger coefficients. Such thresholding functions naturally arise as solutions of quadratic risk minimization problems and are widely used in sparse signal recovery and wavelet shrinkage due to their stability and favorable theoretical properties [
16].
One of the popular methods for choosing a threshold in a thresholding problem is the universal threshold
. Intuitively, the universal threshold is chosen such that, with a high probability, all noise coefficients are less than this value in absolute value. Indeed, if
,
have a normal distribution
, then
Thus, the universal threshold ensures complete suppression of noise coefficients with an asymptotic probability of 1 and is the largest among the so-called useful thresholds (i.e., those that preserve significant signal coefficients but remove noise ones). In particular, any higher threshold will lead to excessive zeroing of informative coefficients.
Therefore, the universal threshold can be considered the upper limit of the range of rational (useful) thresholds. More details can be found in the monograph [
16].
To quantitatively describe the reconstruction quality, we introduce the mean-square risk
which serves as a natural measure of the thresholding procedure’s efficiency. Minimizing
with respect to the parameter
T leads to the definition of the optimal threshold
ensuring the smallest value of the mean-square deviation between the reconstructed and true signals. The value of
depends on the unknown components
and, therefore, cannot be calculated directly in practice.
To circumvent this difficulty, a risk estimate expressed in terms of observed data is used:
where the function
is defined by the rule
It is easy to verify that the estimate constructed in this way is unbiased, that is,
.
Similarly, we define an adaptive threshold that minimizes the empirical risk estimate:
which is known in the literature as the SURE threshold (Stein’s Unbiased Risk Estimate threshold). The threshold
is a random variable dependent on the observations
and serves as a practical replacement for the theoretically optimal
.
In the following discussion, the classical Bernstein inequality [
17] will be used to obtain probability estimates and upper bounds on risk:
where
a.s. This inequality allows one to obtain exponential estimates for the tails of distributions and is widely used in the analysis of sums of independent random variables.
5. Conclusions
This paper presents an asymptotic analysis of the properties of the mean-square risk during threshold processing of noisy signals in a sparse stochastic observation model. For the theoretically optimal threshold that minimizes the theoretical risk, an upper bound is obtained, showing that the risk does not exceed a value of the order .
Furthermore, a central limit theorem for the empirical risk estimate is established, as well as a strong law of large numbers, which guarantees the almost-everywhere convergence of the empirical risk to its theoretical counterpart. These results provide a rigorous theoretical basis for analyzing the asymptotic properties of thresholding procedures in problems of sparse signal recovery from observations with Gaussian noise.
The obtained estimates refine the classical results of Donoho and Johnstone (see [
20]) and confirm the effectiveness of thresholding methods under sparsity conditions.
At the same time, the proposed analysis has several limitations. In particular, the theoretical results are derived under the assumption of independent Gaussian noise and rely on an asymptotic framework with a large number of observations. Moreover, the sparsity structure is characterized by a specific rate of decay of the proportion of nonzero components, which may not fully capture more complex or heterogeneous sparse patterns encountered in practice.
A promising direction for future research is to extend the proposed approach to more general models incorporating dependent noise, inhomogeneous variances, and multidimensional data structures. Another important direction is to study the asymptotic behavior of the optimal threshold under different sparsity regimes and to analyze the minimax properties of the resulting procedures. Finally, it would be of interest to investigate models with a random sample size, where the number of observations N is itself a random variable. Such settings naturally arise in practical applications, including network monitoring problems with random traffic intensity or missing observations, and may require new asymptotic tools for analyzing thresholding procedures and their associated risk.