1. Introduction
The Collatz conjecture, also known as the
problem, is one of the most well-known unsolved problems in mathematics. It concerns the iteration of the map
defined by
The Collatz problem has been studied from several viewpoints, including analytic estimates, probabilistic arguments, computational verification, and structural reformulations. A comprehensive overview of the problem and its major developments is given by Lagarias [
1], who discusses analytic estimates, probabilistic models, computational verification, and structural interpretations of Collatz trajectories. Krasikov and Lagarias [
2] developed analytic bounds using difference inequalities, showing how probabilistic and inequality-based methods can yield nontrivial information about the distribution of Collatz trajectories. Tao [
3] proved that almost all orbits of the Collatz map attain almost bounded values, providing a major probabilistic advance in understanding typical orbit behavior. Another important line of research focuses on the structure of trajectories through cycle lengths, directed graphs, and finite state representations. Eliahou [
4] obtained lower bounds on nontrivial cycle lengths, contributing to the study of possible periodic behavior in Collatz dynamics. From a graph-theoretic viewpoint, Andaloro and, later, van Laarhoven and de Weger investigated graph-based descriptions of Collatz trajectories, including directed graph models and connections with De Bruijn graphs, making it possible to study their iterative behavior through combinatorial and dynamical patterns [
5,
6]. Conway’s (see [
7]) foundational work highlighted the complex and potentially unpredictable nature of generalized Collatz iterations, showing how simple recursive rules can generate highly nontrivial computational behavior. A clustering perspective on the Collatz conjecture has also been proposed to study [
8] structural regularities among Collatz trajectories and to highlight pattern-based relationships in their iterative behavior. Motivated by these structural approaches, we study the level sets
through the parity patterns of the corresponding trajectories.
Starting from a positive integer
n, the associated Collatz sequence is
where
denotes the
k-th iterate of
f. The Collatz conjecture asserts that for every
, there exists a finite integer
k such that
The smallest such
k is called the Collatz length of
n, and we denote it by
For each integer
, we define the level set
which consists of all positive integers whose trajectories reach 1 after exactly
x iterations.
Level sets of Collatz lengths have been investigated from various perspectives, including probabilistic models, analytic bounds, and large-scale computational verification [
9,
10,
11,
12]. In this paper, we study the structure of the sets
through the parity patterns of the corresponding Collatz trajectories.
For each integer
n with Collatz length
, we associate a parity vector
where
Thus, each element of
is encoded by a binary sequence of length
x that records the parity structure of its trajectory.
A fundamental quantity derived from this representation is the number of odd steps,
which counts how many times the odd transformation
occurs in the trajectory of
n. For a fixed length
x, this leads naturally to the subsets
which partition the level set
according to odd-step count. We denote the mean of each such subset by
The main idea of this paper is that the parity representation provides a natural framework for understanding the numerical distribution of integers inside each level set
. Heuristically, if a trajectory of length
x contains
k odd steps, then its starting value is expected to scale approximately like
where
C depends on the specific parity pattern of the trajectory. This suggests that the means
should exhibit an approximate geometric behavior as functions of
k.
Motivated by this observation, we develop a statistical analysis of the parity classes . We prove a basic mean bound under a natural scaling hypothesis, derive bounds for the successive ratios , and present numerical results showing that the logarithms of these means are very well approximated by linear functions of k. These results indicate that the odd-step count is a useful structural parameter for describing the distribution of integers within the level sets .
For
, let
denote the positions of the odd steps in the parity code of
n. It is convenient to introduce the trajectory-dependent quantity
which depends only on the parity pattern
p of the trajectory. The following lemma gives an exact expression for
n and hence for
.
Lemma 1. Let , and let be the positions of the odd steps in its parity code. Then,Equivalently,where Proof. Let
be the Collatz trajectory of
n.
At each step, the update has the form
Hence, every odd step contributes a multiplicative factor 3 and an additive term 1, while every even step contributes a factor
.
If the parity code contains exactly
k odd steps, then the total multiplicative contribution from
to
is
Now consider the additive contribution coming from the
r-th odd step, which occurs at position
. After this step, there remain exactly
odd steps and
even steps. Therefore, the additive term 1 generated at position
contributes
to the final value
.
Since
, summing the multiplicative and additive contributions gives
Multiplying both sides by
, we obtain
Simplifying the exponents yields
Finally, since
we may factor out
and write
where
Substituting the explicit formula for
n gives
□
The rest of the paper is organized as follows. In
Section 2, we introduce the parity classes
and establish the main theoretical bounds for their means. In
Section 3, we present examples and numerical results illustrating the stability of these mean ratios and the corresponding regression behavior across a range of Collatz lengths.
2. Parity Representation and Statistical Clustering
In addition to studying the numerical distribution of integers within each level set,
we examine the internal structure of Collatz trajectories using a binary feature representation.
Let
be the Collatz trajectory of an integer
n with length
. We define the
parity vector
by
Thus, each integer can be associated with a binary sequence describing the parity pattern of its Collatz trajectory. The length of this sequence equals the Collatz length x.
Example 1. For Collatz length , the following integers and their parity codes arise: | L(n) | Parity Code |
| 17 | 12 | 100100010000 |
| 96 | 12 | 000001010000 |
| 104 | 12 | 000100010000 |
| 106 | 12 | 010000010000 |
| 113 | 12 | 100100000000 |
| 640 | 12 | 000000010000 |
| 672 | 12 | 000001000000 |
| 680 | 12 | 000100000000 |
| 682 | 12 | 010000000000 |
| 4096 | 12 | 000000000000 |
A natural feature derived from the parity vector is the number of odd steps,
which counts the total number of occurrences of the odd transformation
in the trajectory of
n.
In the above example, the elements of
are grouped according to the value of
as follows:
Thus, the number of odd steps provides a natural way to organize the integers in
.
This can be understood heuristically from the multiplicative behavior of the Collatz map. If a trajectory of length
x contains
k odd steps and
even steps, then each odd step contributes approximately a factor of 3, while each even step contributes a factor of
. Ignoring the additive
terms in the odd steps, the overall effect on the initial value is therefore approximately
Since the trajectory eventually reaches 1, this suggests the approximation
and hence,
where
C depends on the precise parity pattern of the trajectory. The constant
C accounts for the cumulative effect of the additive
terms in the odd steps as well as the ordering of odd and even operations.
Taking logarithms gives
which shows that, for fixed length
x, the logarithm of
n should vary approximately linearly with the number of odd steps
k.
This provides a statistical explanation for the geometric behavior of the means . Since integers with similar values of k tend to lie in comparable numerical ranges, the means of the classes are expected to exhibit approximately constant ratios. Thus, the observed geometric behavior may be viewed as a consequence of the underlying parity structure of Collatz trajectories.
Fix
and consider the level set
For each , let denote the number of odd steps in its Collatz trajectory.
For each admissible value
k, define the subset
Thus, the level set decomposes as
For each
k, we define the mean value of the set
by
Lemma 2. Assume there exist constants such that for every admissible k and every ,Then, the mean value satisfies Proof. Fix
k. By assumption, for every
,
Summing over all
yields
Dividing by
, we obtain
□
Remark 1. Recall thatwhere depends on the parity pattern p of the trajectory. If one wishes to determine the sharpest constants and satisfying the hypothesis of Lemma 2, then they are given byand Indeed, the inequalitiesare equivalent to Thus, the optimal lower and upper bounds are obtained by taking the minimum and maximum of over all admissible pairs .
Example 2. Grouping by odd-step count gives If we use the scaling law , then the optimal constants are Hence, for every admissible k and every , For instance, when , Corollary 1. Suppose the hypothesis of Lemma 2 holds. Then, for every k, Proof. Dividing the second inequality by the first gives
which simplifies to
□
Theorem 1. Suppose the hypothesis of Lemma 2 holds. Then, for every k, Proof. By Corollary 1, both ratios
belong to the interval
Hence, the absolute difference between them is bounded by the length of this interval; then, we have
□
Statistical Model for Parity-Based Clusters
The parity representation introduced in the previous section suggests that the integers in the set
can be naturally grouped according to the number of odd steps in their Collatz trajectories.
For a fixed Collatz length
x, define
where
denotes the parity indicator defined earlier. Thus,
counts the total number of odd steps in the trajectory of
n.
The empirical analysis indicates that these means follow an approximate exponential relationship with respect to
k. Specifically, the data suggest a model of the form
where
and
are constants depending on the Collatz length
x.
Taking natural logarithms yields the linear relation
Thus, the parameters
A and
b can be estimated by fitting a linear regression model to the points
3. Numerical Results
To investigate the statistical structure of the level sets
, we computed all integers
and determined their Collatz lengths. For each length
, the set
was constructed. Integers in
were then grouped according to the number of odd steps
in their parity representation. The regression analysis is used to test whether the relationship between the cluster means
and the number of odd steps
k follows an exponential pattern. Accordingly, for each cluster indexed by
k, the mean value
was computed, and a linear regression of
against
k was performed. The resulting regression model has the form
To summarize these numerical findings,
Table 1 and
Figure 1 and
Figure 2 present the behavior of the regression parameters across Collatz lengths.
Table 1 shows that, across the tested lengths, the parity-based clustering is consistently pure (purity
, meaning that each cluster is perfectly separated by the odd-step count in the computed data. Across all tested values of
L, the linear fits were extremely strong, with coefficients of determination
values that are extremely close to 1, indicating that the logarithm of the cluster means is very well approximated by a linear function of the cluster index. Moreover, the fitted slopes
b are highly stable, (
). This supports the claim that the cluster means of the parity classes follow an approximate exponential trend as the odd-step count increases.
Figure 1 shows the estimated slope
b as a function of the Collatz length
L. The graph demonstrates that the slope remains highly stable across the entire range
.
Figure 2 illustrates the growth of the level-set size
as
L increases, while exhibiting fluctuations likely tied to the irregular structure of Collatz trajectories. This provides context for the regression analysis by showing how many integers contribute to each computation.