Next Article in Journal
Two-Stage Tensor Linear Mixed Model for Multi-Scale Longitudinal Data Analysis
Previous Article in Journal
Relaxed Local Tracking Conditions for Markovian Jump Systems with Application to DC-DC Synchronous Buck Converters Under Load Resistance Switching Variations
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Emergence of Gamma-Type Upward-Phase Statistics in the Collatz Map: An Effective Poisson Process Mechanism

1
Department of Physics, Tianshui Normal University, Tianshui 741001, China
2
Key Laboratory of Atomic and Molecular Physics & Functional Material of Gansu Province, College of Physics and Electronic Engineering, Northwest Normal University, Lanzhou 730070, China
3
Lanzhou Center for Theoretical Physics, Lanzhou University, Lanzhou 730000, China
4
Key Laboratory for Quantum Theory and Applications of the Ministry of Education, Lanzhou University, Lanzhou 730000, China
*
Authors to whom correspondence should be addressed.
Mathematics 2026, 14(15), 2739; https://doi.org/10.3390/math14152739
Submission received: 25 June 2026 / Revised: 28 July 2026 / Accepted: 30 July 2026 / Published: 2 August 2026
(This article belongs to the Section D1: Probability and Statistics)

Abstract

The Collatz map is a simple deterministic transformation whose orbit structure remains highly nontrivial. A recent direction-phase decomposition partitions each orbit into upward and downward steps, and numerical observations indicate that the number of upward phases, N , follows an approximate Gamma distribution. In this work, we provide a mechanistic explanation for this statistical regularity by modeling the occurrence of upward phases in the odd-compressed, or Syracuse, version of the Collatz map as a homogeneous Poisson process. From the mean-field logarithmic balance and the geometric distribution of 2-adic valuations, we derive closed-form expressions for the Gamma parameters: the scale parameter θ = 2 / ( 2 log 2 3 ) 2 11.61 is constant, whereas the shape parameter K grows logarithmically with the maximal initial value X 0 = 2 L + 1 . We also analyze the closure conditions for periodic orbits, showing that nontrivial cycles are severely constrained, which supports the plausibility of the statistical framework. Numerical validation for L ranging from 10 5 to 10 15 confirms the theory with relative errors below 3 % , and a bias-corrected mean estimate reduces the error to 10 3 10 2 % . These results establish a quantitative link between the arithmetic properties of the Collatz map and Gamma-type statistics, and suggest possible extensions to generalized Collatz-type problems.

1. Introduction

The Collatz map is a classical problem at the intersection of discrete dynamical systems [1,2,3,4,5,6], elementary number theory [7,8,9], and probabilistic modeling [10,11]. It is defined by
X = 3 X + 1 , X is odd ; X / 2 , X is even .
The Collatz conjecture asserts that, for every positive integer X, repeated iteration eventually enters the period-three cycle 4 2 1 4 . The problem was proposed by L. Collatz in the 1930s [8]. Despite its elementary formulation, the associated orbit structure is highly intricate, and a complete proof remains out of reach. A substantial body of work has investigated the problem from the perspectives of stopping times, total stopping times, modular structures, random models, 2-adic dynamics, and probabilistic methods [12,13,14,15].
In recent years, several noteworthy developments have further advanced the study of the Collatz problem. One important direction is understanding the typical behavior of Collatz orbits through probability theory and random dynamical models [14,15]. Tao proved that almost all Collatz orbits attain almost bounded values. Although this result does not show that every orbit reaches 1, it demonstrates that, in the sense of logarithmic density, the overwhelming majority of orbits exhibit a strong tendency toward descent [14]. This work refines the long-standing random-walk heuristic for the Collatz problem into a more precise probabilistic framework, suggesting that the deterministic Collatz map may display effective randomness at appropriate scales. Closely related to this viewpoint, random models of Collatz orbits have long served as an important tool for understanding the problem. Kontorovich and Lagarias systematically studied random models for the 3 X + 1 and 5 X + 1 problems, showing that forward iteration can be approximated by additive random walks, biased random walks, or Markov processes [8].
Another active direction concerns large-scale computational verification. Although computational evidence cannot replace a rigorous proof, it provides important support for understanding the statistical structure of orbits and for excluding counterexamples within large finite ranges [16,17]. Barina recently extended the verification of the Collatz conjecture to X < 2 71 , using improved algorithms and parallel computation, and discussed the role of GPUs and distributed computing in accelerating this task [17]. Such computations show that no nontrivial cycle or divergent orbit appears within an extremely large range of initial values.
Beyond stopping-time analysis and computational verification, recent work has also attempted to characterize the structure of Collatz orbits from the perspective of nonlinear dynamical systems. In Ref. [18], Collatz orbits were decomposed into upward and downward phases, leading to a direction-phase decomposition and a family of recursive functions parameterized by the number of upward phases N . A key numerical observation in that work is that the statistics of odd integers classified by the number of upward phases are well described by a Gamma-type distribution. This suggests that the number of local growth events in Collatz orbits is not merely an irregular count, but may instead exhibit a stable statistical law.
The present work is motivated by the following question: what mechanism gives rise to Gamma-type statistics in the number of upward phases of Collatz orbits? To answer this question, we model the occurrence of upward phases in the odd-compressed, or Syracuse, version of the Collatz map as a homogeneous Poisson process. Combining the mean-field logarithmic balance with the geometric distribution of the 2-adic valuations, we obtain closed-form estimates for the Gamma parameters and explain why the scale parameter is approximately constant whereas the shape parameter grows logarithmically with the maximal initial value. Large-scale numerical experiments are then performed to validate the proposed Poisson process mechanism and the resulting Gamma distribution approximation. Beyond its implications for the statistical analysis of Collatz dynamics, this framework also provides a pedagogically useful example for undergraduate and graduate courses in nonlinear dynamics, number theory, computational physics, and computational mathematics, illustrating how deterministic arithmetic rules can be effectively described through probabilistic and statistical models.
The remainder of this paper is organized as follows. Section 2 introduces the Collatz map and defines the direction-phase decomposition, including the numbers of upward and downward phases. Section 3 develops a probabilistic approximation for the odd-compressed Collatz dynamics and derives the Gamma distribution approximation for the statistics of N , together with theoretical estimates for the fitting parameters and the mean value. Section 4 discusses exact closure conditions for possible periodic orbits and clarifies the role of finite-size corrections in the double-logarithmic asymptotic picture. Section 5 presents numerical results that test the theoretical predictions over a wide range of L. Finally, Section 6 summarizes the main conclusions and discusses the scope and limitations of the proposed approximation.

2. The Collatz Map and Definition of Direction Phases

For brevity, the map (1) is abbreviated as
X n = F ( X n 1 ) = F n ( X 0 ) ,
where n , X n Z + . The Collatz conjecture asserts
lim n F n ( X 0 ) = { 4 , 2 , 1 } for X 0 Z + .
To formalize the dynamics, we introduce “direction phases” [19]:
P ( n ) = 1 , if X n + 1 > X n ( up phase ) ; P ( n ) = 1 , if X n + 1 < X n ( down phase ) .
Notably, N and N correspond to the counts of odd and even terms in the sequence. The total iterations satisfy
N = N + N , with F N ( X 0 ) = 1 ,
where N is parameterized as a function of N as follows:
N = N log 2 3 + 1 X 0 + log 2 ( X 0 ) ,
where x denotes rounding x to the nearest integer [18]. Clearly, for a given X 0 , the total number of iterations is completely determined under the condition that N is known. Moreover, Ref. [18] numerically observed that, over finite ensembles of odd initial values, the upward-phase count N is well described by a Gamma distribution. In the following, we develop a theoretical framework to explain this empirical regularity and derive analytical estimates for the corresponding distribution parameters. Providing a mechanistic account of the numerically observed statistics constitutes both the principal motivation and the central contribution of the present work.
The direction-phase decomposition used here could be understood as a specific modeling convention rather than as a canonical decomposition universally adopted in the Collatz literature. Its application to Collatz trajectories, together with the associated upward-phase count N , was introduced in Ref. [18] from a nonlinear-dynamics perspective. Although the terminology was motivated by the more general concept of phase ordering in nonlinear maps [19], the present upward–downward partition is tailored to the arithmetic structure of the Collatz map.
The advantage of this convention is its direct correspondence with the two elementary operations of the map. An upward phase occurs when an odd integer is transformed according to X 3 X + 1 , whereas a downward phase occurs when an even integer is transformed according to X X / 2 . Consequently, N and N coincide with the numbers of odd and even terms, respectively, along the trajectory before it enters the known cycle. Moreover, under the odd-compressed, or Syracuse [14], representation, each compressed iteration consists of exactly one upward operation followed by one or more divisions by two. Hence, N is exactly the number of iterations of the odd-compressed map. This correspondence makes the decomposition analytically convenient for describing the trajectory, deriving the logarithmic balance, and constructing the effective statistical model developed below. Note that this decomposition is not unique or universally preferred. Other trajectory observables or alternative partitions may lead to different statistical descriptions. The results presented in this work are specifically concerned with the direction-phase convention of Ref. [18] and with the distribution of the resulting observable N . Other decompositions may lead to different statistical observables and possibly richer behavior; exploring such alternatives remains an interesting subject for future research.

3. Poisson-Process Mechanism for Gamma-Type Upward-Phase Statistics

Before introducing the probabilistic approximation, we specify the statistical ensemble considered in this work. For a fixed positive integer L, let
Ω L = { 3 , 5 , , 2 L + 1 }
denote the finite population of odd initial values. For each X 0 Ω L considered numerically, the Collatz trajectory is iterated until it enters the known cycle, and the direction-phase decomposition assigns to it a deterministic upward-phase count N ( X 0 ) . The resulting population of upward-phase counts is therefore
N L = N ( 3 ) , N ( 5 ) , , N ( 2 L + 1 ) .
Both the initial-value population and the map are deterministic; no intrinsic randomness or pseudo-random generation of the initial values is assumed.
A probability distribution may nevertheless be associated with this finite ensemble by assigning the uniform counting measure to Ω L . Equivalently, one may regard X 0 as being selected uniformly from Ω L , in which case N ( X 0 ) becomes an induced random variable with probability mass function
p L ( n ) = 1 L # X 0 Ω L : N ( X 0 ) = n .
The Gamma distribution introduced below is used as a continuous effective approximation to this discrete finite-population distribution. It does not represent intrinsic stochasticity of the Collatz dynamics, nor is it introduced to estimate an unspecified external population parameter. Instead, its shape and scale parameters characterize the location, dispersion, and shape of the distribution of N across the prescribed ensemble of initial values.
The standard Collatz map consists of an odd step, 3 X + 1 , and an even step, X / 2 . Since 3 X + 1 is always even for any odd integer X, an odd step is necessarily followed by one or more successive divisions by 2. It is therefore natural to combine these consecutive even steps and consider the odd-only accelerated map acting on the set of positive odd integers:
X n + 1 = 3 X n + 1 2 h n , h n = ν 2 ( 3 X n + 1 ) ,
where ν 2 ( m ) denotes the 2-adic valuation of m, i.e., the exponent of the highest power of 2 dividing m. This map is also referred to as the reduced Collatz function, the accelerated Collatz function, the odd-only Collatz map, or the Syracuse function [14]. In this compressed representation, each step from X n to X n + 1 consists of one upward operation, X n 3 X n + 1 , followed by h n downward divisions by 2. Hence, for an orbit segment containing M compressed steps, the number of upward phases is N = M , while the total number of downward phases is
N = n = 0 M 1 h n .
Thus, the number of iterations of the odd-compressed map is naturally identified with the number of upward phases.
Figure 1a,b compare the iteration trajectories of the odd-compressed map, indicated by the red lines, with the cobweb plots of the original Collatz map for the initial values X 0 = 19 and X 0 = 31 , respectively. For these two initial values, the corresponding numbers of upward and downward phases are ( N , N ) = ( 6 , 14 ) and ( 39 , 67 ) , respectively. The trajectory generated by the odd-compressed map consists entirely of odd integers, as illustrated by the red dots in Figure 1. In the log 2 log 2 coordinate system, the grid lines correspond to powers of two. Therefore, except for the fixed point X = 1 = 2 0 , no point along the odd-compressed trajectory can coincide with these power-of-two grid lines.
We now seek an effective probabilistic model for the finite-population distribution defined in Equation (9). In the odd-compressed Collatz dynamics, each compressed iteration corresponds to one upward phase. Motivated by the approximately geometric statistics of the associated 2-adic valuations [14], we adopt an effective homogeneous Poisson-process approximation for the occurrence of these upward phases. Under this interpretation, the Gamma family provides a continuous approximation to the distribution of N over the initial-value ensemble, i.e.,
ρ ( N ) = 1 Γ ( K ) θ K N K 1 exp N θ ,
where K and θ are the shape and scale parameters, respectively. Note that the Gamma family is only employed as an effective continuous model, not as an exact distribution derived from the deterministic Collatz dynamics.
The remaining task is to estimate K and θ from the arithmetic structure of the odd-compressed Collatz map. Taking logarithms of the compressed map yields
log 2 X n + 1 = log 2 ( 3 X n + 1 ) h n .
For a sufficiently large X n , we use the approximation
log 2 ( 3 X n + 1 ) log 2 ( 3 X n ) = log 2 X n + log 2 3 .
Substituting this approximation into Equation (13) gives
log 2 X n + 1 log 2 X n + log 2 3 h n .
Equivalently, the logarithmic evolution can be written as
log 2 X n + 1 log 2 X n ξ n ,
where
ξ n = h n log 2 3
denotes the single-step logarithmic decrease. Since h n has mean value larger than log 2 3 , the average value of ξ n is positive, corresponding to a net logarithmic contraction on average.
For odd integers that are approximately uniformly distributed among residue classes modulo powers of 2, the 2-adic valuation h n = ν 2 ( 3 X n + 1 ) may be approximated by a geometric distribution [14], i.e.,
P ( h n = h ) 2 h , h = 1 , 2 , 3 , .
The geometric law in Equation (18) specifies the marginal distribution of each h n , but does not by itself imply independence of successive valuations. A stronger finite-dimensional justification was given by Tao [14], who proved that, when a random odd initial value is approximately uniformly distributed modulo a sufficiently large power of 2, its finite Syracuse valuation vector is exponentially close in total variation to a vector of independent Geom ( 2 ) random variables. Motivated by this result, we treat the successive logarithmic increments as approximately independent. This approximation gives
E [ h n ] = 2 , Var ( h n ) = 2 .
Since ξ n = h n log 2 3 , it follows that
μ E [ ξ n ] = 2 log 2 3 0.4150375 ,
and
σ 2 Var ( ξ n ) = Var ( h n ) = 2 .
For an initial odd integer X 0 , reaching the small attracting cycle requires the accumulated logarithmic decrease to be of the order of log 2 X 0 . At the mean-field level, this logarithmic balance gives
μ N log 2 X 0 .
Taking X 0 = 2 L + 1 as the representative upper boundary of the sampling interval, we obtain
N log 2 X 0 μ = log 2 ( 2 L + 1 ) 2 log 2 3 1 + log 2 L 2 log 2 3 2.4094 + 8.0039 log 10 L .
This estimate is in good agreement with the numerical results reported in Ref. [18].
We next estimate the variance of N by error propagation. After N compressed steps, the accumulated logarithmic decrease is
S N = n = 1 N ξ n .
Motivated by the finite-dimensional independence result discussed above, and within the independent-increment approximation, the variables ξ n are treated as having negligible serial dependence. Hence, for a fixed value of N ,
Var ( S N ) σ 2 N .
At the mean-field level, the accumulated logarithmic decrease is related to the number of upward phases by
S N μ N .
Thus, a small fluctuation in S N induces a corresponding fluctuation in N according to
Δ N Δ S N μ .
It follows that
Var ( N ) Var ( S N ) μ 2 σ 2 μ 2 N .
For a Gamma distribution with shape parameter K and scale parameter θ , one has
E [ N ] = K θ , Var ( N ) = K θ 2 .
Therefore,
θ = Var ( N ) E [ N ] σ 2 μ 2 = 2 ( 2 log 2 3 ) 2 11.6106 ,
and
K = E [ N ] 2 Var ( N ) N θ 2 log 2 3 2 ( 1 + log 2 L ) 0.2075 + 0.6894 log 10 L .
Thus, the theoretical estimate predicts that θ is independent of L, while K grows logarithmically with L. In the effective Poisson-process interpretation, the approximately constant scale parameter is consistent with a constant inverse rate. This numerical and analytical consistency supports the usefulness of the homogeneous approximation, but does not constitute a proof of an exact Poisson point process.

4. Closure Conditions for Periodic Orbits

The statistical description developed above implicitly concerns orbits that eventually enter the known cycle. If other nontrivial periodic orbits existed, their long-time behavior would not be described by the same convergent-orbit statistics. It is therefore useful to examine the exact closure conditions for periodic orbits of the accelerated Collatz map.
Suppose that there exists a periodic orbit of the accelerated map on positive odd integers with period M, namely,
X 0 X 1 X M 1 X 0 ,
where X M = X 0 . Iterating the accelerated map gives the exact balance condition
2 N = n = 0 M 1 3 + 1 X n ,
where N = n = 0 M 1 h n is the total number of downward divisions by 2 along the cycle. For the known cycle 1 4 2 1 , the accelerated map has the fixed point X 0 = 1 , corresponding to M = 1 and N = 2 . In this case, Equation (33) reduces to 2 2 = 3 + 1 .
Equation (33) may be rewritten as
2 N 3 M = n = 0 M 1 1 + 1 3 X n .
For any nontrivial positive cycle, all odd elements satisfy X n 3 . Hence,
1 < 2 N 3 M < 10 9 M .
Taking logarithms, we obtain the necessary condition
log 2 3 < N M < log 2 10 3 .
Moreover, if m = min { X 0 , X 1 , , X M 1 } , then
0 < N M log 2 3 < log 2 1 + 1 3 m < 1 3 m ln 2 .
Thus, any large nontrivial cycle would require the rational number N / M to approximate the irrational number log 2 3 from above with extremely high accuracy.
For example, the numerical verification of the Collatz conjecture up to X 0 < 2 71 implies that any possible nontrivial positive cycle must have m > 2 71 > 2.361183 × 10 21 . Substituting this lower bound into Equation (37) gives
0 < N M log 2 3 < log 2 1 + 1 3 × 2 71 < 1 3 × 2 71 ln 2 < 2.0368 × 10 22 .
Therefore, if one considers a hypothetical sequence of nontrivial cycles with m , then N / M would have to converge to log 2 3 from above with an error tending to zero.
In the asymptotic approximation where 3 X n + 1 is replaced by 3 X n , the closure condition would reduce to
N = M log 2 3 .
This equation cannot hold exactly, since N and M are integers whereas log 2 3 is irrational. Hence, in the double-logarithmic asymptotic picture, nontrivial periodic closure is excluded in the limit m . The constraint becomes increasingly stringent as the minimum element m of the cycle increases.
However, this asymptotic obstruction is not a proof of the nonexistence of all nontrivial cycles. The finite correction 1 / X n in 3 X n + 1 cannot be discarded rigorously in the integer dynamics, because it determines the 2-adic valuation and hence the parity structure of the subsequent trajectory. Indeed, this correction is precisely what produces the known cycle 1 4 2 1 . Therefore, a complete exclusion of all finite-size corrections would amount to proving the uniqueness of the known Collatz cycle. This illustrates the subtle difference between taking asymptotic limits in the real-valued logarithmic approximation and preserving the exact arithmetic structure of the dynamics on positive integers. For example, in the real-valued asymptotic sense, the approximation 3 X n + 1 3 X n is valid as X n . In the integer dynamics, however, this approximation is not innocuous: for odd X n , the quantity 3 X n + 1 is even, whereas 3 X n remains odd. Thus, replacing 3 X n + 1 by 3 X n changes the parity structure and, consequently, the 2-adic valuation that determines the subsequent divisions by 2.

5. Numerical Results

To show how well the proposed probabilistic model captures the statistical behavior of the accelerated odd-compressed Collatz dynamics, this section displays a series of goodness-of-fit analyses comparing theoretical predictions with numerical experiments. The objective is to evaluate the agreement between the empirical distributions of the upward-phase count N and the Gamma-based density ρ ( N ) , examine the scaling of the fitted parameters K and θ with respect to L, and quantify deviations in the mean behavior relative to the theoretical expectation.
Figure 2 shows the statistical counts of N for X 0 { 3 , 5 , , 2 L + 1 } and L = 10 5 , , 10 15 on a semilogarithmic scale. The orange dots denote the numerical counts, while the blue stars show the corresponding results obtained by restricting X 0 to odd values in the range X 0 { 0.2 L + 1 , 0.2 L + 3 , , 2 L + 1 } . The two sets of data nearly overlap, indicating that, for sufficiently large X 0 , the statistical distribution of N is insensitive to the lower cutoff of the sampling interval. The green curves represent the theoretical prediction obtained from the probability density function in Equation (12), namely N ( N ) = L ρ ( N ) , which agrees well with the numerical results.
Figure 3a shows the frequency distribution of N as a function of N for X 0 { 3 , 5 , , 2 L + 1 } at fixed L = 10 15 . The red solid curve denotes the Gamma distribution fit. To better resolve the discrepancy between the numerical data and the fitted curve at small frequencies, the same data are replotted on a double-logarithmic scale in the inset. The fitted curve shows excellent overall agreement with the numerical results. Figure 3b presents the empirical cumulative distribution function constructed from the frequency statistics. Since the cumulative curve is smooth, we fit it using the cumulative distribution function of the Gamma distribution to determine the fitting parameters. The fitted curve in Figure 3a is then drawn using these parameters. The red solid curve in Figure 3b denotes the fitted cumulative distribution function. To further illustrate the fitting accuracy, the inset shows the residual Δ between the empirical and fitted values. The residual satisfies | Δ | < 4 × 10 3 , indicating that the Gamma distribution provides an accurate approximation to the statistics of N .
To quantitatively characterize the dependence of the Gamma approximation on the sampling range of X 0 , Figure 4a shows the fitted Gamma parameters K and θ as functions of L. As L increases, the fitted value of θ gradually approaches an approximately constant value, θ 11.245 , which is slightly smaller than the theoretical prediction θ T 11.61 . The fitted value of K changes from slightly below to slightly above the theoretical prediction, while its overall logarithmic growth remains consistent with the theory. Quantities without subscripts in the legend are obtained from the statistics for X 0 { 3 , 5 , , 2 L + 1 } , whereas quantities with the subscript 2 are obtained from the restricted range X 0 { 0.2 L + 1 , , 2 L + 1 } . The difference between the two data sets is small, indicating that the lower cutoff of X 0 does not affect the qualitative behavior of the distribution.
Figure 4b further examines this dependence by plotting the local slopes obtained from linear fits between adjacent data points for K and θ . The slope of θ approaches zero as L increases, consistent with an asymptotically constant scale parameter. In contrast, the slope of K converges toward its theoretical value, as indicated by the horizontal reference line.
Figure 4c shows the mean value of N as a function of L. The mean estimated from the Gamma distribution fit agrees closely with the directly computed statistical mean, whereas the theoretical estimate is slightly larger. The inset shows the corresponding local slopes obtained from adjacent-point linear fits. The slope of the statistical mean is nearly identical to the theoretical prediction. Although the slope obtained from the fitted mean is slightly smaller, it gradually approaches the theoretical value 8.0039 as L increases.
Figure 4d shows the relative errors of the statistical mean and the fitted mean with respect to the theoretical prediction, i.e., N = K T θ T . The relative error remains below 3 % over the full range of L. When X 0 is restricted to the range from 0.2 L + 1 to 2 L + 1 , the relative error is further reduced to approximately 2 % . Overall, the error decreases with increasing L, consistent with the large- X 0 approximation used in Equation (14). Thus, larger values of X 0 lead to smaller approximation errors in the theoretical derivation.
Figure 4c,d show that the original theoretical prediction slightly overestimates the mean value of N , while accurately capturing its scaling slope. When X 0 is restricted to the interval X 0 { 0.2 L + 1 , 0.2 L + 3 , , 2 L + 1 } , the numerical results agree better with the theoretical prediction. This indicates that using the upper boundary X 0 = 2 L + 1 in Equation (23) as the representative logarithmic scale of the whole sampling interval leads to a systematic overestimation for smaller values of X 0 . This bias decreases as X 0 approaches 2 L + 1 . As a compact empirical finite-range correction for the restricted sampling interval considered here, we replace log 2 ( 2 L ) by the midpoint 1 / 2 + log 2 L of the logarithmic interval [ log 2 L , log 2 ( 2 L ) ] . The empirically corrected estimate is therefore
N 1 / 2 + log 2 L 2 log 2 3 1.2047 + 8.0039 log 10 L .
Figure 5a compares this corrected prediction with the numerical results, showing good agreement. Figure 5b further displays the relative error of the directly computed statistical mean as a function of L. With the corrected theory, the relative error is reduced to the order of 10 3 % 10 2 % . In addition, the relative error of the mean estimated from the fitted Gamma parameters decreases and falls below 0.1 % at L = 10 15 . This correction is sampling dependent and is not claimed to represent a universal finite-size law.

6. Summary and Discussion

In summary, we have provided a mechanistic explanation for the emergence of Gamma-type statistics in the direction-phase structure of Collatz orbits. By decomposing the Collatz dynamics into upward and downward phases and using the odd-compressed representation, we modeled the occurrence of upward phases through an effective homogeneous Poisson process mechanism. Together with the geometric approximation for the 2-adic valuations h n , this framework leads naturally to a Gamma-family approximation for the statistics of N . The corresponding Gamma parameters can be estimated analytically from the arithmetic structure of the map. The scale parameter is predicted to be θ T = σ 2 μ 2 = 2 ( 2 log 2 3 ) 2 11.61 . In contrast, the shape parameter K grows logarithmically with L, consistent with the mean-field logarithmic balance between the accumulated decrease and log 2 X 0 . A corrected estimate for the mean value of N , obtained by replacing the upper-boundary approximation with the average logarithmic scale of the sampling interval, further reduces the relative error to the order of 10 3 % 10 2 % .
We also examined the closure conditions for possible periodic orbits of the accelerated Collatz map. The exact balance condition 2 N = n = 0 M 1 3 + 1 / X n implies that any nontrivial large cycle would require N / M to approximate log 2 3 from above with an accuracy better than 1 / ( 3 m ln 2 ) , where m is the minimum element of the cycle. As m , this condition becomes arbitrarily stringent. In the asymptotic approximation where 3 X n + 1 is replaced by 3 X n , the closure condition reduces to N = M log 2 3 , which is impossible because log 2 3 is irrational. This provides an asymptotic obstruction to large nontrivial cycles, although it does not constitute a rigorous exclusion of all finite-size corrections.
Several questions remain open. Although the effective Poisson-process approximation is supported by the numerical results, a rigorous derivation from the deterministic Collatz dynamics is still lacking. In particular, it would be important to quantify correlations among successive values of h n and to determine whether a suitable limit theorem can be established for the odd-compressed map. Such a correlation analysis would not only provide a more reliable basis for the homogeneous Poisson-process approximation, but could also help explain and correct the slight systematic overestimation of the value of θ relative to the numerical measurement. A careful investigation of these effects is therefore left for future work. Another direction is to extend the direction-phase framework to generalized Collatz-type maps, such as a X + b problems, in order to clarify whether the observed Gamma-type statistics are specific to the classical 3 X + 1 map or reflect a broader feature of parity-driven integer dynamics.
Beyond its implications for the statistical analysis of Collatz dynamics, the present framework also provides a useful pedagogical example. For undergraduate and graduate teaching, it provides a compact example showing how an elementary deterministic rule can generate highly nontrivial behavior and how number-theoretic structure, discrete dynamics, probabilistic modeling, and numerical evidence can interact. In this sense, the Collatz problem is not only a classical topic in elementary number theory, but also an instructive model for illustrating how deterministic arithmetic dynamics may admit effective stochastic descriptions, while numerical evidence can motivate but not replace rigorous mathematical analysis.

Author Contributions

Conceptualization, W.F.; methodology, W.F., X.L. and Y.W.; software, X.L.; validation, W.F., X.L. and Y.W.; formal analysis, W.F., X.L. and Y.W.; investigation, W.F., X.L. and Y.W.; resources, W.F., X.L. and Y.W.; writing—original draft preparation, W.F.; writing—review and editing, W.F., X.L. and Y.W.; visualization, W.F.; project administration, W.F.; funding acquisition, W.F. and Y.W. All authors have read and agreed to the published version of the manuscript.

Funding

This research was funded by the National Natural Science Foundation of China (Grants No. 12465010, No. 12247106, and No. 12575040). W. Fu was also supported by the Gansu Province Long-yuan Youth Talent Project, the Fei-tian Scholars Project of Gansu Province, the Leading Talent Project of Tianshui City, the Innovation Fund from the Department of Education of Gansu Province (Grant No. 2023A-106), the Open Project Program of the Key Laboratory of Atomic and Molecular Physics & Functional Material of Gansu Province (Grant No. 6016-202404), and the Young Faculty Achievement Award (Grant No. PYJ-01252291) & Postgraduate Curriculum Construction Project (Grant No. TKXM2601) of Tianshui Normal University.

Data Availability Statement

The raw data supporting the conclusions of this article will be made available by the authors on request.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Crandall, R.E. On the “3x +1 ” Problem. Math. Comput. 1978, 32, 1281–1292. [Google Scholar] [CrossRef]
  2. Kozak, J.J.; Musho, M.K.; Hatlee, M.D. Chaos, Periodic Chaos, and the Random-Walk Problem. Phys. Rev. Lett. 1982, 49, 1801–1804. [Google Scholar] [CrossRef]
  3. Krasikov, I. How Many Numbers Satisfy The 3x + 1 Conjecture? Int. J. Math. Math. Sci. 1989, 12, 791–796. [Google Scholar] [CrossRef]
  4. Abraham, R.H.; Gardini, L.; Mira, C. Chaos in Discrete Dynamical Systems, 1st ed.; Springer: New York, NY, USA, 1997. [Google Scholar] [CrossRef]
  5. Wirsching, G.J. The Dynamical System Generated by the 3n+1 Function, 1st ed.; Lecture Notes in Mathematics; Springer: Berlin/Heidelberg, Germany, 1998; Volume 1681. [Google Scholar] [CrossRef]
  6. Siegel, M.C. The Collatz Conjecture & Non-Archimedean Spectral Theory—Part I—Arithmetic Dynamical Systems and Non-Archimedean Value Distribution Theory. p-Adic Numbers Ultrametric Anal. Appl. 2024, 16, 143–199. [Google Scholar] [CrossRef]
  7. Akin, E. Why is the 3x+ 1 problem hard? Contemp. Math. 2004, 356, 1–20. [Google Scholar] [CrossRef]
  8. Lagarias, J.C. (Ed.) The Ultimate Challenge: The 3x+1 Problem; American Mathematical Society: Providence, RI, USA, 2010. [Google Scholar]
  9. Wang, M.; Yang, Y.; He, Z.; Wang, M. The Proof of the 3X + 1 Conjecture. Adv. Pure Math. 2022, 12, 10–28. [Google Scholar] [CrossRef]
  10. Borovkov, K.A.; Pfeifer, D. Estimates for the Syracuse Problem via a Probabilistic Model. Theory Probab. Appl. 2001, 45, 300–310. [Google Scholar] [CrossRef]
  11. Sinai, Y.G. Statistical (3x + 1) problem. Commun. Pure Appl. Math. 2003, 56, 1016–1028. [Google Scholar] [CrossRef]
  12. Terras, R. A stopping time problem on the positive integers. Acta Arith. 1976, 30, 241–252. [Google Scholar] [CrossRef]
  13. Lagarias, J.C. The 3x + 1 Problem and Its Generalizations. Am. Math. Mon. 1985, 92, 3–23. [Google Scholar] [CrossRef]
  14. Tao, T. Almost all orbits of the Collatz map attain almost bounded values. Forum Math. Pi 2022, 10, e12. [Google Scholar] [CrossRef]
  15. Polli, J.G.; Raposo, E.P.; Viswanathan, G.M.; da Luz, M.G.E. Stochastic-like characteristics of arithmetic dynamical systems: The Collatz hailstone sequences. J. Phys. Complex. 2024, 5, 015011. [Google Scholar] [CrossRef]
  16. Barina, D. Convergence verification of the Collatz problem. J. Supercomput. 2021, 77, 2681–2688. [Google Scholar] [CrossRef]
  17. Barina, D. Improved verification limit for the convergence of the Collatz conjecture. J. Supercomput. 2025, 81, 810. [Google Scholar] [CrossRef]
  18. Fu, W.; Wang, Y. The Structure of the Route to the Period-Three Orbit in the Collatz Map. Math. Comput. Appl. 2026, 31, 23. [Google Scholar] [CrossRef]
  19. Wang, W.; Liu, Z.; Hu, B. Phase Order in Chaotic Maps and in Coupled Map Lattices. Phys. Rev. Lett. 2000, 84, 2610–2613. [Google Scholar] [CrossRef] [PubMed]
Figure 1. Cobweb plot of the orbits with X 0 = 19 (a) and X 0 = 31 (b), respectively. The red lines are the iterative trajectories of the odd-compressed Collatz map.
Figure 1. Cobweb plot of the orbits with X 0 = 19 (a) and X 0 = 31 (b), respectively. The red lines are the iterative trajectories of the odd-compressed Collatz map.
Mathematics 14 02739 g001
Figure 2. The orange dots denote the counts N ( N ) as a function of N ( X 0 ) for X 0 { 3 , 5 , , 2 L + 1 } , plotted on a semilogarithmic scale. The dots from bottom to top correspond to L = 10 5 , , 10 15 . The blue stars are obtained in the same way, but with X 0 restricted to X 0 { 2 L × 10 1 + 1 , , 2 L + 1 } . The green solid curves are the theoretical estimate, i.e., N ( N ) = L ρ ( N ) , see Equation (12).
Figure 2. The orange dots denote the counts N ( N ) as a function of N ( X 0 ) for X 0 { 3 , 5 , , 2 L + 1 } , plotted on a semilogarithmic scale. The dots from bottom to top correspond to L = 10 5 , , 10 15 . The blue stars are obtained in the same way, but with X 0 restricted to X 0 { 2 L × 10 1 + 1 , , 2 L + 1 } . The green solid curves are the theoretical estimate, i.e., N ( N ) = L ρ ( N ) , see Equation (12).
Mathematics 14 02739 g002
Figure 3. (a) Frequency distribution of N as a function of N for X 0 { 3 , 5 , , 2 L + 1 } at fixed L = 10 15 . The red solid curve denotes the Gamma distribution fit. Inset: the same data plotted on a log-log scale. (b) Empirical cumulative distribution function of the frequency statistics. The solid curve represents the cumulative distribution function of the fitted Gamma distribution. The fitting parameters are K = 10.70924 ± 0.00621 and θ = 11.24522 ± 0.00661 , with a reduced chi-square of 4.3560117 × 10 7 and an adjusted R 2 of 0.9999961 . Inset: residuals between the empirical data and the fitted Gamma distribution, defined as the difference between the empirical and fitted values.
Figure 3. (a) Frequency distribution of N as a function of N for X 0 { 3 , 5 , , 2 L + 1 } at fixed L = 10 15 . The red solid curve denotes the Gamma distribution fit. Inset: the same data plotted on a log-log scale. (b) Empirical cumulative distribution function of the frequency statistics. The solid curve represents the cumulative distribution function of the fitted Gamma distribution. The fitting parameters are K = 10.70924 ± 0.00621 and θ = 11.24522 ± 0.00661 , with a reduced chi-square of 4.3560117 × 10 7 and an adjusted R 2 of 0.9999961 . Inset: residuals between the empirical data and the fitted Gamma distribution, defined as the difference between the empirical and fitted values.
Mathematics 14 02739 g003
Figure 4. (a) Fitting parameters θ (upward triangles) and K (circles) of the Gamma distribution as functions of L, for X 0 { 3 , 5 , , 2 L + 1 } . The solid and dash-dotted lines denote the theoretical predictions θ T [Equation (30)] and K T [Equation (31)], respectively. (b) Local slopes obtained from linear fits between adjacent data points in panel (a) as functions of L. The horizontal reference lines indicate the corresponding theoretical values. (c) Statistical mean of N (circles) and the mean estimated from the Gamma distribution fit, K θ (upward triangles), as functions of L. The solid line represents the theoretical estimate K T θ T [Equation (23)]. Inset: local slopes obtained from linear fits between adjacent data points in the main panel. The horizontal reference line indicates the theoretical value. (d) Relative error between the mean of N shown in panel (c) and the corresponding theoretical prediction. Quantities labeled with the subscript 2 are obtained in the same way, but with X 0 restricted to X 0 { 0.2 L + 1 , , 2 L + 1 } .
Figure 4. (a) Fitting parameters θ (upward triangles) and K (circles) of the Gamma distribution as functions of L, for X 0 { 3 , 5 , , 2 L + 1 } . The solid and dash-dotted lines denote the theoretical predictions θ T [Equation (30)] and K T [Equation (31)], respectively. (b) Local slopes obtained from linear fits between adjacent data points in panel (a) as functions of L. The horizontal reference lines indicate the corresponding theoretical values. (c) Statistical mean of N (circles) and the mean estimated from the Gamma distribution fit, K θ (upward triangles), as functions of L. The solid line represents the theoretical estimate K T θ T [Equation (23)]. Inset: local slopes obtained from linear fits between adjacent data points in the main panel. The horizontal reference line indicates the theoretical value. (d) Relative error between the mean of N shown in panel (c) and the corresponding theoretical prediction. Quantities labeled with the subscript 2 are obtained in the same way, but with X 0 restricted to X 0 { 0.2 L + 1 , , 2 L + 1 } .
Mathematics 14 02739 g004
Figure 5. (a) Statistical mean of N (circles) and the mean estimated from the Gamma distribution fit, K 2 θ 2 (stars), as functions of L, where X 0 restricted to X 0 { 0.2 L + 1 , , 2 L + 1 } . The solid line represents the empirically corrected theoretical estimate K T θ T [Equation (40)]. (b) Relative error between the mean of N shown in panel and the corresponding corrected theoretical prediction (a).
Figure 5. (a) Statistical mean of N (circles) and the mean estimated from the Gamma distribution fit, K 2 θ 2 (stars), as functions of L, where X 0 restricted to X 0 { 0.2 L + 1 , , 2 L + 1 } . The solid line represents the empirically corrected theoretical estimate K T θ T [Equation (40)]. (b) Relative error between the mean of N shown in panel and the corresponding corrected theoretical prediction (a).
Mathematics 14 02739 g005
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Fu, W.; Liu, X.; Wang, Y. Emergence of Gamma-Type Upward-Phase Statistics in the Collatz Map: An Effective Poisson Process Mechanism. Mathematics 2026, 14, 2739. https://doi.org/10.3390/math14152739

AMA Style

Fu W, Liu X, Wang Y. Emergence of Gamma-Type Upward-Phase Statistics in the Collatz Map: An Effective Poisson Process Mechanism. Mathematics. 2026; 14(15):2739. https://doi.org/10.3390/math14152739

Chicago/Turabian Style

Fu, Weicheng, Xiaobin Liu, and Yisen Wang. 2026. "Emergence of Gamma-Type Upward-Phase Statistics in the Collatz Map: An Effective Poisson Process Mechanism" Mathematics 14, no. 15: 2739. https://doi.org/10.3390/math14152739

APA Style

Fu, W., Liu, X., & Wang, Y. (2026). Emergence of Gamma-Type Upward-Phase Statistics in the Collatz Map: An Effective Poisson Process Mechanism. Mathematics, 14(15), 2739. https://doi.org/10.3390/math14152739

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop