Abstract
Line transect sampling is widely used for estimating population abundance, but existing nonparametric estimators of detection density at the transect line often suffer from boundary bias and tuning sensitivity. In this paper, we propose two simple tuning-light estimators based on minimum order statistics of perpendicular distances, requiring measurement of only the judged-closest object within each set. Under mild regularity conditions, the proposed estimators are consistent and asymptotically normal, with low bias and variance demonstrated through simulation studies under exponential and half-normal detection models. An application to a wooden stakes transect survey illustrates the practical advantages of the proposed approach for low-effort ecological surveys.
1. Introduction
Estimating population abundance is a central problem in ecological and environmental statistics. Accurate measures of animal or plant density provide the foundation for understanding species distribution, monitoring biodiversity, and guiding conservation policies. Among the most widely used approaches, distance sampling—and, in particular, the line transect method—has attracted extensive attention due to its efficiency and practicality [1,2,3]. In this methodology, observers traverse predefined lines placed randomly in a study area with total length Lm and they record the perpendicular distances from each detected object to the transect line. These distances are then used to model the detection process and infer the abundance of the underlying population.
Let x denote the perpendicular distance from the line to an object of interest, and let be its probability density function (). At the transect line (i.e., ), the detection probability is assumed to equal one. In line transect surveys, abundance per unit area D is commonly expressed as
where denotes the expected number of detections [2]. Thus, the main statistical challenge lies in estimating the point value ; that is, the density of the observed perpendicular distances at the line.
Recent work has shown that spatial variation at multiple scales can challenge standard line transect modeling assumptions, motivating alternative designs and estimators that are less sensitive to model misspecification [4].
1.1. Background and Limitations of Existing Methods
A widely used estimator, under standard regularity assumptions, can simply be written as (see Buckland et al. [2])
such that n is the observed number of detected objects and is an estimate of f obtained from the measured sample of perpendicular distances.
Two broad classes of estimation techniques exist in the literature to produce . Parametric approaches assume a specific form for the detection function, such as the half-normal or exponential models, and they then estimate the associated parameters via the maximum likelihood or the method of moments [5,6]. While these methods are efficient when the assumed model is correct, they may lead to biased estimates when model assumptions are violated. Alternatively, nonparametric methods, such as histogram and kernel estimators, avoid strong distributional assumptions and offer flexibility in modeling the detection process that can reduce boundary bias or improve robustness when certain conditions are violated [7,8,9,10].
Nevertheless, despite their flexibility, nonparametric estimators of are often subject to boundary bias at the line and exhibit sensitivity to ad hoc tuning parameters, such as bandwidth in kernel estimators [2,9]. Related approaches, including pooling strategies, are frequently used to stabilize inference in distance sampling; however, their robustness properties require careful interpretation [3].
As a more principled alternative within the nonparametric paradigm, we have developed an approach that concentrates information near the line through order statistics. By exploiting the closest objects to the transect line and leveraging the efficiency gains afforded by order statistic sampling, we construct two tuning-light estimators based on independent, non-overlapping subsets of ranked distances. This design effectively replaces the arbitrary choice of bandwidth with a more interpretable decision on the number of subsets, thereby yielding smoother density estimates and improved computational performance. Recent theoretical developments have sharpened the basis for line transect inference, particularly by clarifying when design-based estimators remain valid under heterogeneity and what pooling robustness guarantees [3].
1.2. Motivation for and Contribution of This Study
In practical line transect surveys, observers are often able to reliably identify the object that is closest to the transect line using simple visual indicators such as apparent distance, viewing angle, or ease of detection, even when exact measurement of all distances is not feasible. This type of partial ranking information is inexpensive to obtain and is typically more reliable near the transect line, where detection probability is highest. Nevertheless, standard distance sampling methods do not explicitly exploit this information.
The existing parametric and nonparametric estimators of generally require either full distance measurements or careful tuning of smoothing parameters, both of which can be difficult to implement in resource-limited surveys. Recent advances in order statistic-based density estimation suggest that concentrating on extreme observations, such as minima, can yield stable and efficient estimators with reduced tuning requirements (Kung et al. [11]; Garg et al. [12]). Although these methods were not originally developed for line transect sampling, they provide important methodological insight for constructing estimators that emphasize near-boundary information.
The primary contribution of this study is the development of two simple, tuning-light estimators for the boundary density based on minimum order statistics of perpendicular distances. The proposed methods are constructed using independent, non-overlapping sets of detected objects and require measuring only the judged-closest object within each set. This design enforces independence by construction and replaces arbitrary smoothing choices with an interpretable set size parameter. Under mild regularity conditions, the proposed estimators are shown to be consistent and asymptotically normal. Their finite-sample performance was evaluated through simulation studies under common detection models, and their practical utility is illustrated using a wooden stakes transect survey. The practical relevance of this motivation is further supported by the application results in Section 3.3, which show that the proposed estimators achieve stable and competitive performance while requiring substantially fewer distance measurements than conventional approaches.
Theory has sharpened the basis for line transect inference, clarifying when design-based estimators stay valid under heterogeneity and what “pooling robustness” actually guarantees [3]. Simulations also show that pooling stabilizes totals but can mask group-level bias when detectability differs [13]. This motivates methods that emphasize well-identified near-line information, notably estimators targeting
Integrated distance sampling (IDS) is increasingly used to combine transect detections with auxiliary data, which can improve precision in large-scale analyses [14,15]. Such frameworks typically benefit from inputs that require minimal tuning, where robust estimators of may be useful.
In-field logistics matter where incomplete viewsheds reduce the detectable area with distance and bias inference if ignored; recent guidance shows how to account for this in design and analysis [16]. Comparisons of walked versus road transects likewise highlight coverage–detectability trade-offs that favor estimators with predictable behavior across modes [17].
Beyond the line transect sampling framework, several authors have proposed innovative strategies for general density estimation. For example, Kung et al. [11] introduced a block-based method that exploits localized information from order statistics to construct flexible and computationally efficient estimators of . Later, Garg et al. [12] extended this perspective by developing density estimators that combine subsampling with order statistic properties, thereby enhancing stability and reducing reliance on strong distributional assumptions. Although these approaches were not originally designed for line transect data, they highlight the potential of order statistic-driven techniques to achieve robustness and efficiency in density estimation, and they provided the methodological insights that motivated our development of subset-based estimators for .
2. Materials and Methods
The methods proposed in this paper are fundamentally design-driven: their accuracy depends not only on distributional assumptions, but also on how observations are selected, ranked, and recorded in the field (or in repeated sampling). Accordingly, we begin by specifying the study design and sampling framework that generates the minimum-based summaries on which our estimators rely, highlighting the key design parameters that control sampling effort and information content. We then introduce a general theoretical setup that abstracts these mechanisms into a probabilistic model, so that the resulting properties and comparisons of the proposed estimators can be stated and proved in a unified way.
2.1. Study Design and Sampling Framework
In what follows, we demonstrate the practical use of the proposed methodology for abundance estimation in line transect sampling. Within each set, observers employ a simple and efficient proxy to identify the unit judged to be closest to the transect line. At this stage, neither exact perpendicular distances nor a full ranking of all units are required. Instead, only the judged-closest unit is subsequently measured exactly, and these subset minima form the data used to estimate the target density at the line. The new procedure needs a relatively small set size, so ranking through distance comparisons can be performed or judged easily.
- Step 1.
- Choose design settings.Fix a reasonably small set size r, the number of sets m you will form per route or time window (in each “block”), and how many independent routes S you will run. This design yields detections encountered (not all measured) and exact distances measured (one per set).
- Step 2.
- Form disjoint sets in the field (block ). As observers traverse a line, group the next fixed number of successive detections into one set; start a new set for the next group of successive detections, and so on. Do not recycle detections across sets or routes (windows). Therefore, along a given route t, group r successive detections into a set; repeat m times to create m non-overlapping sets. Keep routes so that groups of detections do not overlap in space or time [2].
- Step 3.
- Rank efficiently using a cheap indicator. Within each set, use a rapid cheap proxy (e.g., sighting angle, apparent range, cue strength, visibility class, drone image cue) to select the single unit judged closest to the line. No exact perpendicular distances yet; no full ordering is required.
- Step 4.
- Measure only the one judged closest. In each set, measure the exact perpendicular distance for the judged-closest unit only; do not measure the rest. This yields one measured value from that set with minimal effort. For the purposes of theoretical development, we consider this minimum as the true minimum representing the perfect ranking. After m sets, a route contributes m measured distances; after S routes, the study collects measurements.
- Step 5.
- Repeat across the survey. Continue Step 2–Step 4 until forming m sets for the current route and recording one exact distance per set. Then move to the next route until you complete the planned number of sets for the route. Repeat on independent routes to maintain independence. For comparison with conventional line transect sampling, these steps can be repeated c times to obtain the desired sample size.
- Step 6.
- Extend the procedure for enhanced efficiency. To produce more efficient estimators, the exact measurement of the unit judged closest in Step 4 can be temporarily withheld. Instead, mark the identified minima without recording their distances, keeping them available for subsequent comparisons across different sets. To facilitate accurate judgment by the observer, flags or markers can be placed on these units. Then, perform a visual comparison of the marked units within each block to select the single unit judged closest to the line, and record its exact perpendicular distance. This procedure should be repeated at the block level to achieve the desired sample size.
- Step 7.
- Record keeping. For each set, store the route ID, set ID, quick ranks, which unit was measured, and the exact perpendicular distance. This ensures transparent auditing and reproducibility.
The following example illustrates the proposed set-based sampling construction and clarifies the notation used for the minimum order statistics.
Example 1
(Field construction with perpendicular distances). Throughout, let r denote the set size (number of consecutive detections per set), m the number of sets per block (per route or time window), and S the number of independent blocks. In this example, we take , , and .
Block . As the crew walks a transect, detections are grouped into non-overlapping sets of consecutive detections each. Let denote the perpendicular distance of the ith detection in set j on block t (labels only; no numbers are needed). The raw sets on block are
Within each set, observers use a rapid low-cost visual feature to identify the single unit judged closest to the line; only that unit’s perpendicular distance is then measured exactly. Denote the one measured value from set j on block by , .
Block . On an independent pass (different route or time window), the same construction yields three new sets and one exact perpendicular distance per set,
Across the two blocks we encounter detections (all counted), but we measure exactly one perpendicular distance per set, producing exact measurements:
These measurements are used to construct the first type of our estimators.
To enhance estimator efficiency, the exact measurement of the unit judged closest in Step 4 can be temporarily withheld. Denote the initially identified minima in set j on block t by , which are marked but not yet measured. These marked units are kept available for visual comparison across sets. To facilitate accurate judgment, flags or markers can be placed on them. Then, within each block the observer performs a visual comparison of the marked units to select the single unit judged closest to the line, whose exact perpendicular distance is then recorded and denoted . This procedure is repeated at the block level to achieve the desired number of exact measurements.
From a practical point of view, it is crucial that the subset size r remains identical across all sets within all blocks, while the number of blocks may be chosen as needed to achieve the desired sample size. Because sets within a block do not overlap and blocks are disjoint in space or time, the M measured perpendicular distances are obtained under a design that ensures independence both across sets and across blocks. This structure yields i.i.d. minimum order statistics, which form the basis of the estimation procedure proposed in this paper.
Relative to the above procedure of element selection, two estimation methods are proposed in this paper for estimating and, subsequently, the abundance D. The first method applies estimation based on the set of minimums We refer to this as the first stage of selection. The second method requires an additional step at the block level by selecting the minimum of these minimums. The produced elements are , where each element is defined as for blocks. We refer to this as the second stage of selection.
2.2. General Theoretical Setup
Let be i.i.d. perpendicular distances with right-continuous density f on with absolutely continuous cumulative distribution function (cdf) F, where denotes the truncation distance (i.e., the strip half-width). Then, the minimum order statistics are i.i.d. with cdf and pdf , such that
where . Accordingly, the minima of these minimum order statistics are also i.i.d. with cdf and pdf , such that
where .
The following lemma, which is analogous to Garg et al. [12], can be applied respectively on and , the first- and second-stage minimum perpendicular distances, as follows. Suppose that the minimum perpendicular distances, denoted by , are obtained from either the first- or second-stage sampling procedures described in Section 1.
Lemma 1.
Let denote i.i.d. copies of the minimum order statistic arising from samples of size η, whose parent distribution has continuous pdf and cdf . Then, as , for any fixed block t:
- 1.
- .
- 2.
- If f is twice differentiable in a neighborhood of 0 and is continuous at 0 then
- 3.
Proof.
Define , so that are i.i.d. standard uniform random variables. Let ; then, with (see David and Nagaraja [18]) (p. 14)). Moreover, assume is the quantile function; thus, by the monotonicity of .
The inverse-function identity yields
Equivalently, there exists a function with as , such that
Using and (2),
Select an arbitrary . Since , choose , such that for all . Therefore,
where by continuity of Q on . Now , so exponentially fast. Therefore
Letting gives . Thus,
Hint: If a higher-order bias expression is required, such as the second-order bias, under the conditions (which follows from the condition commonly referred to as the shoulder condition in line transect sampling) and if exists then
This completes the proofs of parts (1) and (2). For part (3), Poisson approximation is employed. Hence, , and it follows that limiting variance is . □
Lemma 1 is derived under the assumption of perfect identification of the minimum within each set, which represents an idealized sampling scenario. In practical field applications, particularly those relying on rapid visual assessment, occasional misclassification of the closest unit may occur. As discussed in the preceding Section 2.3, the probability of correctly identifying the minimum remains high for moderate separation margins, and the difficulty of minimum-only ranking is substantially lower than that of full ranking. Moreover, the robustness study in Section 3.2.3 demonstrates that moderate departures from perfect ranking primarily result in increased variability rather than severe bias inflation. These findings indicate that the theoretical results provide a reliable approximation under realistic field conditions, even when perfect ranking cannot be guaranteed.
Proposition 1.
1. The first-stage minimum order statistics satisfy Lemma 1 by setting with .
- 2.
- The second-stage minimum order statistics satisfy Lemma 1, where , and the following hold:
- (i)
- which consequently yields
- (ii)
- If is twice differentiable in a neighborhood of 0 and is continuous at 0 then , which leads to
- (iii)
- which in turn implies
2.3. Challenge of Ranking
The accuracy of our forthcoming estimators for the boundary density and, hence, for abundance D, depends on the feasibility of performing accurate and inexpensive ranking of the minima. In this section, we show that minimum perfect ranking can, in some sense, be achieved with relative ease in the first stage, and that it is even simpler in the second stage. This property is essential for the practical usefulness of our new estimation method.
To formalize the ranking difficulty, we adopt a margin-based criterion. The two units and can be correctly ordered if for some preassigned positive [19]. Extending this idea, Al-Saleh and Al Kadiri [20] defined a set as perfectly rankable if so that all pairwise differences exceed the margin.
For random variables , the associated probability of full perfect ranking is given by
where corresponds to the first and second sampling stages.
Table 1 illustrates these probabilities for the standard exponential distribution, as an example, showing that for and that the values decrease as grows.
Table 1.
Illustrative values of full perfectly ranking and under the standard exponential distribution. (Source: [20]).
In the new procedure developed in this paper, the practically relevant ranking task is to distinguish only the closest unit from the rest by a preassigned margin . This requirement is strictly easier than “full” perfect ranking, which demands that all pairwise separations exceed , and, thus, it yields to higher probabilities.
Definition 1
(Minimum-only Separability). Let be random variables with order statistics . We say the set is minimum-only perfectly rankable at margin ε if For stage we define the corresponding probability by
Since full perfect ranking implies minimum-only separability, we have for each k and m.
For minimum-only identification, it suffices to ensure that the smallest element is at least below all the others. Since is the closest competitor to , we have
Consequently, the following statements are equivalent:
Hence, the condition is sufficient for minimum-only separability at margin , and no additional pairwise comparisons are required. When the cdf F is continuous, ties occur with probability zero, so strict inequalities are well-defined.
In the following example, we illustrate Definition 1 in the context of the two sampling stages introduced above.
Example 2.
- 1.
- Ranking minimum in the first stage: If are i.i.d. random variables with pdf f and cdf F then the probability of first stage minimum-only separability at margin ε is
- 2.
- Ranking minimum in the second stage: Let denote the jth order statistic from an i.i.d. size-m sample with cdf F. In addition, assume are independent across j (e.g., they arise from different sets or blocks). Then, the probability of second stage separability isFor Exp(1) density with sample size , we have
For , , while typically does not admit a simple closed form and must be evaluated numerically. Another example is the standard Half-Normal distribution, for which the corresponding minimum-only separability probabilities are also obtained numerically.
Table 2 presents the numerical values of and for the and Half-Normal distributions with and . The results indicate that consistently exceeds for both sample sizes, and that the probabilities vary as a function of . Furthermore, in comparison with Table 1, which reports the probabilities of full perfect ranking under the same distribution, it is evident that the probability of correctly ranking the minimum observation is uniformly greater than its full perfect ranking counterpart. This indicates that ranking the minimum elements is easier than performing a full ranking of all elements. Moreover, ranking the minima in the second-stage sampling yields a more favorable outcome than ranking the minima in the first-stage sampling.
Table 2.
Values of minimum-only probability of correct identification and for the standard Exponential and Half-Normal distributions.
In practical field surveys, identification of the judged-closest unit may be subject to observer error due to imperfect visual cues, occlusion, or time constraints. Such judgment error may lead to occasional misranking within a set, where the selected unit is not the true minimum. However, the proposed methodology relies only on minimum-only ranking rather than full ranking, which substantially reduces the difficulty of correct identification, as discussed above. As shown by the probabilities in this section, Section 2.3, the likelihood of correctly identifying the minimum remains high even under moderate separation margins. From a statistical perspective, imperfect ranking primarily increases variability rather than introducing systematic bias, provided that misranking occurs symmetrically and with bounded probability. This robustness motivates the use of low-cost visual proxies in rapid field implementations.
3. Results
This section reports the main outcomes of the study. We first present the proposed minimum-based estimators and summarize their implementation in a unified procedure, emphasizing how the design parameters control the information extracted from repeated sampling. We then provide the associated theoretical and empirical results—covering analytic properties, simulation-based performance, and the applied study—to quantify the accuracy, efficiency, and practical robustness across the considered settings.
3.1. Theoretical Results
3.1.1. Proposed Estimators
Using the sampling framework of Section 1 as a foundation, this section proposes two complementary methodologies to estimate the boundary density and, hence, the population abundance D. We examine their key properties, including bias, variance, and asymptotic behavior, extending classical order statistic approaches through the distributional structure of minimum samples.
Stage I constructs minimum-based summaries from repeated samples by selecting and retaining the relevant minima under the prescribed sampling design, thereby producing the core statistics that drive the subsequent estimation.
Let denote the minimum observation within the set of block t, for , . The collection consists of i.i.d. copies of the minimum order statistics from samples of size r, with pdf and cdf . Here, r denotes the set size, m the number of subsets within each block and S the number of independent blocks, as defined in Section 1. By setting , we define the empirical mean of these units as
Motivated by Proposition 1, our first-stage estimators of the boundary density for and are, respectively:
Hint: under the mild regularity conditions satisfied in our setting, the reciprocal of a consistent and asymptotically normal estimator remains consistent and asymptotically normal [11]. Since our primary estimator targets and possesses these properties, the inverted estimator inherits consistency and asymptotic normality (see Proposition 2).
The corresponding abundance estimator with observed count M and transect length L is
We now proceed to Stage II, where these minimum-based summaries are combined to yield the final estimator. For each block t, we define the block representative—that is, the block minimum-of-minima—as
which are i.i.d. with pdf and cdf . We define
Using Proposition 1, hence we define the second-stage estimator
Thus, the abundance estimate is
Relative to the first-stage approach, this second-stage method compresses information within each block by considering only the block-minimum statistics. This results in a smaller effective sample size (S instead of ) but yields a sharper functional dependence on . Intuitively, the “min-of-minima” amplifies information about the boundary density by pushing samples closer to zero, at the cost of increased variability. Thus, and highlight a trade-off between efficiency and variance reduction, depending on how r and m are tuned.
3.1.2. Improving Density Estimation via Nearest-Neighbor
From a general density-estimation perspective, minimum order statistics can be naturally interpreted as nearest-neighbor distances to a target point. In particular, estimating the boundary density in line transect sampling can be viewed as a pointwise density-estimation problem, where the most informative observations are those closest to the transect line. Nearest-neighbor methods are well known for their stability and low tuning requirements in such settings, especially near boundaries. In this context, an analogous viewpoint arises in nearest-neighbor methods evaluated at a fixed point. Kung et al. [11] and Garg et al. [12] showed that, under mild regularity conditions, splitting the sample into subsets and averaging absolute distances yields a consistent estimator of the density . These results directly motivated the following estimation method, which adapts the minimum-distance principle to the line transect setting, targeting as the key quantity and, consequently, the abundance density D.
Suppose are i.i.d. perpendicular distances with pdf f, where we seek the value of f at a particular location . A convenient way to recast this pointwise task is to work with distances to the target: The variables are non-negative with density so In other words, estimating the mass of the folded distribution at zero, , immediately yields the desired . A bonus of this reduction is that g is typically smoother at zero than f is at , which stabilizes estimation at the point of interest [12].
From a practical point of view, the analyst chooses a target point at which the density is to be estimated. Two practically motivated types dominate near the transect: (i) blind-zone anchors and (ii) anti-heaping anchors. A blind-zone anchor sets , where it is the smallest distance at which detections are reliably made (e.g., beyond a vehicle hood, boat rail, or dense verge). In practice, is defined from a short pilot or calibration pass as the minimal distance with stable detection and measurement—often the maximum of the empirically observed blind width and twice the instrument’s rounding resolution. An anti-heaping anchor sets to avoid known rounding bins that inflate ; here, is chosen from small, non-heaped offsets (e.g., , , m), typically guided by a fine near-zero histogram or rug plot. In both cases the minimum local distance mechanics are identical—only the chosen changes.
To produce new density estimators, return to the raw data ; , from density f before forming . Randomly partition the data into m subsets of size r each, and repeat the entire procedure S times (block level). At the block level and for each m subsets, let be x’s nearest-neighbor distance computed within the subset. For each block t, define
The quantity labeled in (3) is simply the mean of these over j values; that is:
Similarly, in (6) is also the mean of these across the t blocks.
Thus, two additional estimators can be constructed from
that are
3.1.3. Large Sample Behavior and Convergence of Minima-Based Estimators
Let and ; let and . Under the standing assumptions the density f is continuous at 0 with and finite derivative , and it is twice continuously differentiable. For an interior target , we assume and that is finite. Blocks and sets are independent across the design, and within each set of size r, the observations are i.i.d. Thus, we establish the following propositions.
Proposition 2.
Assume that the sizes, and S are as defined in Section 1 and set . Then:
- (i)
- If and with thenandConsequently, for .
- (ii)
- If r- and m for the second stage-are fixed while both and then
Both the first- and second-stage estimators are asymptotically unbiased for their limits and attain variances of order and , respectively. When f is differentiable at zero, the bias relative to decays at rate and rate for the two estimators, respectively; see Lemma 1.
Proposition 3
(Standard error estimation). Let be the sample variance of the M values of and let be the sample variance of . Then
Accordingly,
To compare the SE of estimators [ and ] when and for some leads to the SE optimal rate for any non-negative . In smoothing the histogram, the rate is ; with the Kernel density method, the rate is .
Proposition 4
(Rates for minima-based estimators). For a design comprising S blocks and m sets per block, the estimators and , defined in (4) and (7), respectively, satisfy
where for and for . Similarly, the folded estimators at an interior point , and , as defined in (12), exhibit the same convergence rates with the corresponding .
The proposition can be established straightforwardly once we note that and . Apply the delta method to with :
Hence, and similarly for ; the second order bias is , so . For , replace X by ; the same argument applies and the extra factor only changes constants.
Remark 1
(Rates for histogram and KDE at the boundary for fair comparators). Let be the size of the reduced sample used by the comparator ( for the first-stage level; S for the second-stage level). With bin width h and kernel bandwidth h chosen at standard optimal orders,
3.2. Simulation Studies
To investigate the finite-sample performance of the developed estimators, we conducted a series of Monte Carlo simulations under both exponential and half-normal detection models. The design followed the structure of Section 1, with set size r, number of sets per block m, and number of independent blocks S varied across scenarios. Unless otherwise noted, each scenario was replicated 5000 times. All our simulations and data analyses were implemented in R (version 4.4.3). Implementation details and the corresponding R code are provided in Appendix A.
3.2.1. Boundary Estimation Under Perfect Ranking
We compared four estimators derived from subset minima as illustrated in Equations (4), (7) and (11): , corresponded to the boundary case with perfect detection at the line and folded estimators evaluated at a small interior point , denoted and . Two standard nonparametric benchmarks were also included for comparison: the histogram estimator (Hist) and a boundary-corrected kernel density estimator (KDE with reflection). For fairness, all the comparators were applied to the same reduced samples ( or S units) as the proposed methods. Thus, the effective sample sizes were for and and for and , ensuring a like-for-like comparison that did not advantage the benchmarks by giving them access to the full raw sample of perpendicular distances.
This section focuses on boundary targets, examining the bias–variance patterns of the proposed estimators for the boundary density under several detection-distance models representing both non-shoulder and shoulder behavior. Performance was evaluated in terms of Monte Carlo bias and mean squared error (MSE). Table 3 and Table 4 summarize the results. Table 3 focuses on boundary estimation , while Table 4 reports estimation at an interior target motivated by blind-zone or anti-heaping adjustments.
Table 3.
Estimation of using first- and second-stage methods. Results for random samples from Exponential, Half-Normal, and Hazard-Rate Distributions. True values: , , and for Hazard-Rate. Boldface indicates the smallest MSE among the compared methods for the same setting within each distribution.
Table 4.
Estimation of using first- and second-stage methods (folded at ). Results for random samples from Exponential, Half-Normal, and Hazard-Rate distributions with . True values: , , and for Hazard-Rate. Boldface indicates the smallest MSE among the compared methods for the same setting within each distribution.
In line transect sampling, commonly studied detection functions include the Exponential model, representing a non-shoulder case, the Half-Normal model with a moderate shoulder, and the Hazard-Rate model, which allows for a pronounced shoulder. Across these generating models, the pooled-minima estimator exhibited the anticipated bias decay as r increased: for moderate , the averages were around – under Exponential with negative bias that shrank toward zero as r grew, consistent with an bias term. In contrast, the block min-of-minima often attained smaller bias at a fixed configuration but showed larger variance because its effective sample size was only S; the MSE reflected this trade-off and decreased when S grew or when m was large enough to stabilize block-level minima. These patterns are consistent with the theoretical orders and .
Histogram and KDE Benchmarks
On the reduced samples, the boundary-corrected KDE (reflection) reliably dominated the histogram in the MSE, in line with the canonical rates versus . For very small (e.g., large r but small m), both nonparametric benchmarks could be unstable at the boundary, with the histogram particularly sensitive to the first-bin mass at . As increased, the KDE improved steadily; however, typically matched or surpassed the KDE once was moderately large, reflecting the variance rate (up to constants and second-order bias in ).
3.2.2. Interior-Point Estimation (Folded Estimators)
For interior evaluation at , the folded estimator stabilized rapidly as r increased and was frequently near-unbiased in both models, mirroring the boundary behavior of but benefiting from smoother local curvature at . The block-folded estimator again showed the expected bias reduction at the cost of higher variance due to its reliance on S block minima. In most of the scenarios, the KDE remained a strong benchmark at the interior points; nevertheless, was often competitive or superior in the MSE once was moderate, while the histogram remained sensitive to bin choice.
As seen in Table 5 and Table 6, across values, the folded pooled set estimator was essentially unbiased with small variance, yielding the lowest MSE in all three values. The MSE declined mildly as increased, consistent with greater stability of the folded minima away from the boundary. The block min-of-minima exhibited a persistent positive bias and substantially larger variance, reflecting its smaller effective sample size S and therefore larger MSE.
Table 5.
Estimation of interior density for samples from distribution with fixed sample sizes design. Boldface indicates the smallest MSE among a set values of .
Table 6.
Estimation of interior density for samples from Half-Normal distribution with fixed sample sizes design. Boldface indicates the smallest MSE among the compared values of .
3.2.3. Robustness Against Closest-Unit Misclassification
To assess robustness against imperfect closest-unit identification, we conducted a small Monte Carlo study incorporating judgment error in the set-based selection. Within each set of size r, the true minimum was selected with probability , while with probability p the second-smallest distance was selected, following the standard ranking-error models used in the ranked set sampling literature. We considered misclassification probabilities . The results show that moderate judgment error led to gradual increases in variability without severe bias inflation, supporting the robustness of the proposed methodology under realistic field conditions.
Bias and MSE were evaluated for the first-stage and second-stage estimators under the Exponential and Half-Normal detection models using the same design settings as in Section 3.2. Table 7 summarizes the results. As expected, increasing misclassification led to gradual efficiency loss, reflected in larger MSE, while the estimators remained well-behaved. In the extreme case of severe misclassification, performance approached that of simple random sampling using the same number of measured distances. Related results in the ranked set sampling literature show that efficient estimation remains feasible under ranking or measurement errors, with performance degrading smoothly rather than failing [21].
Table 7.
Robustness of and to closest-unit misclassification for random samples with configuration . True values: for standard Exponential and for standard Half-Normal distributions. Boldface indicates the smallest MSE among the compared values of p.
In conclusion, all of the above results confirm that minima-based estimators provide a computationally light, tuning-robust alternative to standard nonparametric methods, with especially strong gains in scenarios with moderate-to-large set sizes and block counts.
3.2.4. Practical Guidance for Choosing r, m, and S
The performance of the proposed estimators depends on the design parameters r (set size), m (sets per block), and S (independent blocks), which determine the bias–variance trade-off and effective sample size. We summarize our practical recommendations below.
The set size r primarily controls bias: larger r yields minima closer to the transect line and lower bias but increases judgment difficulty and sensitivity to ranking error. The parameters m and S control variance through the effective sample size, with for the first-stage estimator and for the second-stage estimator. Increasing S is particularly important for stabilizing the second-stage estimator.
Table 8 provides concise guidance for selecting under common field constraints, consistent with the theoretical results in Section 3.1.1and the simulation findings in Section 3.2.
Table 8.
Practical guidance for selecting r, m, and S.
3.2.5. Choice of the Folding Target
The folded estimators in Section 3.2.2 require selecting a target point , motivated by practical issues near the transect line, such as blind zones and distance heaping. In blind-zone settings, may be chosen as the smallest distance beyond which detections are reliably recorded, as identified from a brief pilot or calibration pass. In anti-heaping settings, can be selected as a small offset from zero that avoids rounded or heaped values.
In practice, simple empirical diagnostics (e.g., near-zero histograms or measurement resolution) are sufficient for choosing . Typical values are small relative to the truncation distance, and the simulation results in Section 3.2 indicate that the folded estimators are not overly sensitive to modest changes in within this near-line region.
3.2.6. When Should Minima-Based Estimators Be Used?
Minima-based estimators are particularly well suited for surveys where precise measurement of all perpendicular distances is costly, but reliable identification of the closest unit within small sets is feasible. They are advantageous when robustness to tuning parameters is desired, when interest centers on near-line behavior, or when field effort must be minimized.
In contrast, KDE and parametric distance-sampling methods may be preferable when full distance measurements are readily available, detection functions are well specified, or likelihood-based inference is required. In practice, the proposed estimators are best viewed as complementary to the existing methods, offering a low-effort alternative rather than a replacement for standard approaches.
3.3. Real Data Study: Wooden Stakes Transect Survey
3.3.1. Study Design and Data Collection
To illustrate the applicability of the proposed estimators in practice, we designed a field-style study in which artificial wooden stakes were used as a proxy for objects encountered along a line transect, closely resembling the classical Wooden Stakes experiment of Burnham et al. [5]. A straight transect of length m was laid out randomly across a uniform open field. Wooden stakes were positioned randomly within a strip extending m on either side of the transect line. The stakes were placed flush with the ground to mimic small animal or plant features and to reduce visual bias from above-ground height. Our field team consisted of the principal investigator and a graduate student assistant. The investigator carried a laser rangefinder and a measuring tape to obtain exact perpendicular distances when required. The assistant was responsible for making the initial “closest unit” judgments within each set.
This experiment represented an intentionally idealized setting in which all objects were detectable with certainty. Its primary role in this study was to serve as a proof-of-concept, allowing the proposed estimators to be evaluated under controlled conditions without confounding effects from distance-dependent detection or imperfect observability. By design, this experiment isolated the statistical behavior of the minimum-based estimators and facilitated direct comparison with existing methods. The results should therefore be interpreted as illustrating methodological performance under ideal detection, rather than as a fully realistic ecological survey.
Data Collection Procedure
Following the sampling design outlined in Section 1, we adopted a set-based sampling approach. The transect was partitioned into segments (routes) per cycle, and within each segment sets of candidate objects were defined. Each set contained candidate stakes, positioned naturally within the field of view. The assistant visually judged the nearest stake to the transect line within each set, using simple visual signs (apparent proximity, relative angle, or sighting ease). This was an efficient approach done quickly without ordering the remaining units and with low-cost judgment. Then, the perpendicular distance of the judged-closest stake was measured exactly, using laser rangefinder (or measuring tape). Thus, from each set exactly one distance was measured and recorded.
After completing sets per route and routes, a total of distances were obtained per cycle. To increase replication, we conducted independent cycles of the procedure, resulting in a final dataset of observed set minima.
Extensions. To accommodate block-minima and folded estimators, and , we also considered (i) block minima across each route (yielding block minima), and (ii) folded distances around a chosen interior reference , reflecting practical cases where interior density evaluation is relevant. In our study we adopted m. This choice was justified by two practical considerations:
- Blind zone at the line: stakes very close to the line 25 cm) were hard to measure precisely due to grass or ground cover.
- Heaping and measurement error: Measurements extremely close to zero are subject to “rounding” errors, as field assistants tend to record them as exact zeros. Folding at avoided this artifact and produced a stable estimation point.
These matched directly to the estimators , , , and defined in Section 3.1.1.
3.3.2. Observed Perpendicular Distances
The following table presents the observed perpendicular distances (in meters) collected using the judged-closest protocol described in Section 3.3.1. The distances are truncated to three decimal places.
3.3.3. Estimation Results
Let denote the mean of the 60 set minima and the mean of the 15 block minima; similarly, and denote the means of the folded set and block minima, as defined in Section 3.1.1. From the data in Table 9, we obtained Accordingly, the four estimators were , , , and , with m.
Table 9.
Observed perpendicular distances (judged-closest, one per set) for all cycles from Wooden Stakes Transect Survey. Design: , , .
Our designed protocol, where stakes, m and m, corresponding to a covered strip area , produced the true population density and (i.e., 3 cycles with 60 observed distances each). This yielded the design-based reference which equaled
As seen in Table 10, the four estimators all produced values close to this design-based reference, with and aligning most closely. The folded estimator was more sensitive to the choice of , and here provided a conservative lower value. This exercise illustrates how the proposed minima-based estimators can be applied to real field-style data, offering valid alternatives to classical full-data methods in situations where only partial observations are feasible.
Table 10.
Estimated boundary and interior densities from the wooden stakes dataset. Design: , , per cycle, cycles ( set minima, 15 block minima). Interior folding at m. ESW = effective strip width .
A simple significant note is that the relatively large ESW implied by at m reflects the smoothing effect induced by evaluating the folded estimator away from the boundary, rather than model misspecification; this mild inflation is consistent with the bias–variance trade-off inherent in avoiding near-zero detection irregularities and remains compatible with the known truncation width m.
Complete Distance Data
In addition to the judgment-minimum protocol, we measured all perpendicular distances for the full population of wooden stakes. The complete dataset lets us exploit the survey’s full information content, directly compare the minima-based estimators with conventional line transect estimators that use all detections, and transparently assess the efficiency trade-off between collecting a few targeted distances and undertaking exhaustive measurement. While our proposed estimators are designed to achieve efficiency with minimal effort, the complete dataset provides a benchmark for validation and ensures that comparisons to standard methods are transparent and scientifically rigorous.
4. Discussion
The proposed minima-based estimators are developed under the standard assumptions of line transect sampling, namely perfect detection at the transect line, independence across sets and blocks, and correct identification of the closest unit within each set. These assumptions are shared by most design-based and nonparametric distance-sampling estimators and are therefore not specific to the present approach.
Violations of perfect detection at the line affect the scale of the boundary density and consequently induce proportional bias in abundance estimates, a well-known issue in line transect inference. As with existing methods, this can be mitigated through blind-zone anchoring, auxiliary detection modeling, or integrated distance-sampling frameworks, within which the proposed estimators can be readily embedded.
Independence is primarily required at the block level and is enforced by construction through the use of non-overlapping sets and spatially or temporally separated blocks. Because only minimum observations are retained from each set, the estimators are less sensitive to local dependence among detections than methods that rely on full within-set measurements.
Accurate identification of the closest unit corresponds to ranking by judgment in the ranked set sampling literature. As established in that literature, moderate ranking errors primarily reduce efficiency rather than consistency. In the extreme case of complete misranking—where the selected unit within each set is effectively random—the proposed procedure reduces to a standard estimator based on simple random sampling using the same number of measured distances. Thus, ranking errors lead to a gradual loss of efficiency rather than structural bias, and the method remains well-behaved under realistic field conditions.
The results of both our simulation experiments and our field-style wooden stakes survey highlight the practical strengths and limitations of the proposed minima-based estimators. At the boundary, and converge rapidly and are competitive—often superior-to histogram and boundary-corrected KDE. For interior targets, the folded estimators and perform well and offer a flexible extension. A key advantage is operational, where only one exact distance per set is required after a quick closest judgment, substantially reducing measurement burden.
In the full-data stakes study (), minima-based estimates of and ESW aligned closely with conventional analyses. Choice of can be guided by blind-zone or anti-heaping considerations, providing stable inference near the line. These findings directly support the motivation of the proposed methodology by demonstrating that reliable abundance estimation can be achieved using minimal measurement effort when only the closest objects to the transect line are recorded.
Overall, the study demonstrates that minima-based estimators can provide accurate, efficient, and operationally simple density estimates in field surveys. While further work is needed to extend the framework to more complex detection functions and clustered populations, the current results provide a clear foundation for practical adoption. A possible refinement would also be to add further selection stages, which may increase the method’s efficiency.
Author Contributions
Conceptualization, M.A.A.K.; methodology, M.A.A.K.; formal analysis, M.A.A.K. and M.H.A.-H.; writing—original draft, M.A.A.K.; writing—review and editing, M.A.A.K. and M.H.A.-H. All authors have read and agreed to the published version of the manuscript.
Funding
This research received no external funding.
Data Availability Statement
The dataset supporting the findings of this study is openly available in Zenodo at DOI: 10.5281/zenodo.17692827.
Acknowledgments
The authors thank the anonymous referees for their crucial comments on an early version of this paper. All computations and simulations were carried out in R (version 4.4.3).
Conflicts of Interest
The authors declare no conflicts of interest.
Appendix A. R Code for Simulation Studies
The script below implements the proposed estimators (boundary) and (interior, folded at ), along with the histogram and KDE comparators computed on the same reduced samples for a fair comparison. Alternative generating models (Exponential and Hazard-Rate) can be evaluated by replacing only the data-generation function, while keeping the estimation and summarization routines unchanged.
| Listing A1. R script for simulation study (Half-Normal baseline). | |
![]() ![]() ![]() ![]() | |
References
- Buckland, S.T.; Anderson, D.R.; Burnham, K.P.; Laake, J.L. Distance Sampling: Estimating Abundance of Biological Populations; Chapman and Hall: London, UK, 1993. [Google Scholar]
- Buckland, S.T.; Anderson, D.R.; Burnham, K.P.; Laake, J.L.; Borchers, D.L.; Thomas, L. Introduction to Distance Sampling: Estimating Abundance of Biological Populations; Oxford University Press: Oxford, UK, 2001. [Google Scholar]
- Baer, B.R.; Thomas, L.; Buckland, S.T. A foundation for the distance sampling methodology. arXiv 2025, arXiv:2504.12439. [Google Scholar] [CrossRef] [Scilit]
- Olav, N.B.; Hans, J.; Martin, J.; Martin, B. Spatial Variation on Multiple Scales in Line Transect Data; the Case of Antarctic Fin Whales. J. Am. Stat. Assoc. 2025, 1–13. [Google Scholar] [CrossRef] [Scilit]
- Burnham, K.P.; Anderson, D.R.; Laake, J.L. Estimation of Density from Line Transect Sampling of Biological Populations. Wildl. Monogr. 1980, 72, 3–202. [Google Scholar]
- Ababneh, F.; Eidous, O. Weighted Exponential Model for Line Transect Sampling. Stat. Probab. Lett. 2012, 82, 2000–2006. [Google Scholar] [CrossRef] [Scilit]
- Chen, S.X. Empirical Likelihood Confidence Intervals for Nonparametric Density Estimation. Biometrika 1996, 83, 329–341. [Google Scholar] [CrossRef] [Scilit]
- Eidous, O.M. Bias Correction for Histogram Estimator Using Line Transect Sampling. Environmetrics 2005, 16, 61–69. [Google Scholar] [CrossRef] [Scilit]
- Silverman, B.W. Density Estimation for Statistics and Data Analysis; Chapman & Hall: London, UK, 1986. [Google Scholar]
- Eidous, O.M. A New Kernel Estimator for Abundance Using Line Transect Sampling Without the Shoulder Condition. J. Korean Stat. Soc. 2012, 41, 267–275. [Google Scholar] [CrossRef] [Scilit]
- Kung, Y.H.; Lin, P.S.; Kao, C.H. An Optimal k-Nearest Neighbor for Density Estimation. Stat. Probab. Lett. 2012, 82, 1786–1791. [Google Scholar] [CrossRef] [Scilit]
- Garg, V.V.; Tenorio, L.; Willcox, K. Minimum Local Distance Density Estimation. Commun. Stat.-Theory Methods 2017, 46, 148–164. [Google Scholar] [CrossRef] [Scilit]
- Rexstad, E.; Buckland, S.T.; Marshall, L.; Borchers, D.L. Pooling robustness in distance sampling: Avoiding bias when there is unmodelled heterogeneity. Ecol. Evol. 2023, 13, e9684. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Nabias, J.; Lorrillière, R.; Dupuy, J.; Couzi, L.; Barbaro, L. Improving national-scale breeding bird surveys with integrated distance sampling. Sci. Rep. 2025, 15, 18312. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Nilsen, E.; Nater, C. Open integrated distance sampling for modelling age-structured population dynamics. Peer Community J. 2025, 5, 587. [Google Scholar] [CrossRef] [Scilit]
- Delaney, D.M.; Harms, T.M.; Harris, J.; Kaminski, D.J. Methods to account for incomplete viewsheds in distance sampling. Methods Ecol. Evol. 2025, 16, 1201–1214. [Google Scholar] [CrossRef] [Scilit]
- Ferrarini, M.; Nielsen, Ó.K. Distance sampling: Comparing walked transects and road transects for rock ptarmigan densities and population trends. Wildl. Biol. 2025, e01350. [Google Scholar] [CrossRef] [Scilit]
- David, H.A.; Nagaraja, H.N. Order Statistics; Wiley: Hoboken, NJ, USA, 2003. [Google Scholar]
- Takahasi, K. Practical note on estimation of population mean based on samples stratified by means of ordering. Ann. Inst. Stat. Math. 1970, 22, 421–428. [Google Scholar] [CrossRef] [Scilit]
- Al-Saleh, M.; Al Kadiri, M. Double-ranked set sampling. Stat. Probab. Lett. 2000, 48, 205–212. [Google Scholar] [CrossRef] [Scilit]
- Ahmadini, H.; Singh, R.; Raghav, Y.S.; Kumari, A. Estimation of Population Mean Using Ranked Set Sampling in the Presence of Measurement Errors. Kuwait J. Sci. 2024, 51, 2307–4108. [Google Scholar] [CrossRef] [Scilit]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.



