Next Article in Journal
Fuzzy Iterative Learning Contouring Control
Previous Article in Journal
Relational Nonlinear Almost Contractions of Mukherjea Type and Applications to Boundary Value Problems
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Geometric Structure of Genomes Across the Tree of Life: Toward a Geometric Theory of Sequence Structure

by
Valentin E. Brimkov
1,* and
Reneta P. Barneva
2
1
Mathematics Department, SUNY Buffalo State University, 1300 Elmwood Ave., Buffalo, NY 14222, USA
2
School of Business, SUNY Fredonia, 280 Central Ave., Fredonia, NY 14063, USA
*
Author to whom correspondence should be addressed.
Mathematics 2026, 14(10), 1760; https://doi.org/10.3390/math14101760
Submission received: 29 November 2025 / Revised: 1 May 2026 / Accepted: 15 May 2026 / Published: 20 May 2026
(This article belongs to the Section E3: Mathematical Biology)

Abstract

This work develops a geometric and statistical framework for analyzing the structure of biological sequences and explores its implications for understanding the emergence and evolution of life. Motivated by questions concerning the transition from prebiotic chemistry to living systems, the quantification of negentropy in organic matter, and the distinction between random and biologically viable sequences, we introduce mathematical descriptors that measure deviation from linearity and related geometric irregularities of self-replicating macromolecules. These descriptors reveal a pronounced geometric separation between biological DNA and random sequences, underscoring the non-random structural organization characteristic of living systems. Using these descriptors, we compare a broad range of species across the Tree of Life and examine how geometric complexity varies between primitive and more advanced organisms. We further investigate whether these measures provide a natural way to compare organismal complexity, characterize the structure of viable sequence space, and identify potential constraints on evolutionary trajectories. The framework also offers an initial perspective on how natural selection and stochastic mutations may jointly influence genomic organization. Finally, we outline speculative connections between increasing geometric irregularity and the emergence of biological complexity, suggesting that such geometric transitions may offer insight into the origins of life and the theoretical limits of evolutionary development.

1. Introduction

The origin of life and the mechanisms that shaped early evolution remain among the most challenging and conceptually profound questions in science (for classic overviews, see [1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22]). Despite extensive research across chemistry, biology, geology, and physics, many aspects of life’s emergence are inaccessible to direct observation, and competing theories often coexist without the possibility of decisive empirical resolution. This is not surprising: the relevant events occurred billions of years ago under environmental conditions that can only be partially reconstructed, and the available evidence is fragmentary, indirect, and shaped by subsequent geological processes. Even seemingly simple questions, such as the precise timing of life’s appearance on Earth, admit uncertainties of hundreds of millions of years, reflecting both the incompleteness of the fossil record and the difficulty of interpreting ancient geochemical signatures.
These limitations do not imply that progress is impossible but, rather, that fully deterministic or experimentally reproducible theories of life’s origin are unlikely to be achieved. Instead, conceptual frameworks that identify structural, informational, or geometric constraints on viable biological organization may offer a productive way forward. Such approaches aim not to reconstruct specific historical events but to characterize the quantitative properties that distinguish living systems from random chemistry and that shape the space of possible evolutionary trajectories. Within this perspective, mathematical structures and quantitative descriptors can play a central role in clarifying what kinds of molecular organization are compatible with replication, heredity, and functional diversification.
Given the inherent difficulties in formulating a complete and experimentally verifiable theory of the origin of life and early evolution, any meaningful conceptual framework must satisfy several broad criteria.
  • Proposed scenarios should be consistent with physical and chemical plausibility, avoiding assumptions that contradict established principles;
  • They must agree with the available empirical evidence, however fragmentary, including geological, biochemical, and genomic data;
  • Finally, the underlying hypotheses should be logically coherent and grounded in general principles rather than relying exclusively on isolated or incomplete observations.
These criteria reflect the methodological constraints of origins-of-life research, where direct experimentation under ancient Earth conditions is impossible and the surviving evidence is sparse. At the same time, they highlight the increasing role of mathematics and computation in biological inquiry. Over the past decades, fields such as mathematical biology, theoretical biology, computational biology, and bioinformatics have emerged precisely to address questions that cannot be resolved through empirical observation alone. Mathematical models and computational analyses have provided insight into diverse biological phenomena, from molecular organization to evolutionary dynamics, by identifying structural patterns and constraints that are not accessible through traditional experimental approaches. Within this conceptual landscape, several fundamental questions arise concerning the quantitative structure of biological sequences and the constraints that govern their organization and evolution. Among them are the following:
  • What quantitative characteristics must be satisfied for a transition from prebiotic chemistry to self-replicating, living systems?
  • How can one measure the negentropy of individual organisms, groups of organisms, or living systems as a whole?
  • To what extent do biological sequences differ from random sequences, and what structural features account for these differences?
  • Can quantitative parameters distinguish primitive from more complex organisms, and do they provide a meaningful basis for comparing biological complexity?
  • Is there a mathematical structure that naturally represents the DNA sequence of each organism and unifies all such sequences within a common framework?
  • What are the properties of this structure; how does it evolve; and what theoretical limits, if any, constrain further evolutionary development?
Some of these questions have been explored in earlier work [23,24,25], where mathematical constructs and computational experiments were introduced to analyze the geometry of DNA sequences. Traditional approaches to biosequence analysis often rely on combinatorial methods such as pattern matching and combinatorics on words. In contrast, the present work adopts a geometric perspective, applying the idea of string geometrization to genomic sequences. A central contribution of this work is the introduction of a new geometric technique for the analysis of genomic sequences based on the representation of DNA sequences as monotone paths and monotone graphs. This framework provides quantitative descriptors, such as geometric deviation and path-based negentropy, that are not available through existing combinatorial or statistical methods. Because these constructions are general and reusable, they offer a methodological foundation that can support future studies of genomic organization, complexity, and the structure of viable sequence space.
Although modern molecular biology routinely analyzes DNA, RNA, and protein sequences, it remains a fundamental biological fact that all cellular organisms—bacteria, archaea, protists, fungi, plants, and animals—possess DNA-based genomes. RNA genomes occur exclusively in viruses, which are non-cellular entities with extremely small genomes and replication mechanisms that differ profoundly from those of cellular life. Because the present study aims to compare genomic organization across species in the biological sense, our analysis is restricted to cellular organisms, for which DNA is the universal hereditary substrate. Including RNA viruses would conflate cellular and non-cellular systems, introduce orders-of-magnitude differences in genome size, and require distinct methodological assumptions and scaling frameworks. Our focus on DNA genomes is therefore a deliberate and biologically coherent choice that ensures comparability across the Tree of Life. While RNA and protein analyses are essential in many modern contexts, such as transcriptomics, regulatory biology, and structural studies, they address questions different from those considered here. The geometric properties investigated in this work pertain specifically to large-scale organization of genomic DNA; extending these methods to viral RNA genomes would constitute a separate line of inquiry with its own biological and methodological constraints.
In developing this framework, we conducted a series of computational experiments on diverse genomic sequences, examining their geometric, structural, and statistical properties. The patterns observed in these analyses motivate several hypotheses and suggest possible interpretations of the underlying mechanisms that shape genomic organization. These interpretations are not presented as definitive conclusions but as informed conjectures grounded in computational evidence and guided by theoretical considerations. Their purpose is to stimulate further investigation and to highlight potential connections between geometric descriptors of sequences and broader questions concerning the emergence, structure, and evolution of living systems.
The remainder of this work is organized as follows. Section 2 introduces the basic notions and notation used throughout the paper. Section 3 recalls key definitions related to string geometrization and presents several new concepts, including the discrete irregular helix, the monotone graph, and measures of negentropy for monotone paths and monotone graphs; we also examine fundamental properties of these structures. Section 4 introduces the DNA Graph of Life and discusses how the proposed mathematical framework can be applied to its analysis. Section 5 defines the negentropy of individual species, groups of species, and the DNA Graph of Life as a whole. Section 6 describes, in detail, the computational procedures used in our experiments, and Section 7 presents the principal results obtained from these computations. In particular, we explore distinctions between biological and random sequences, potential boundaries of life expressed through the introduced geometric parameters, and gradients between primitive and complex organisms. Section 8 offers hypotheses and biological interpretations derived from our findings, while Section 9 outlines a speculative outlook concerning transitional states and the emergence of biological complexity. Section 10 concludes with final remarks and directions for future research.
Throughout the paper, selected key conclusions are highlighted for clarity. Several theoretical results related to string geometrization were previously stated without proof; the corresponding proofs are provided in Appendix A. Appendix B contains tables summarizing selected outcomes of our computational experiments.

2. Notions and Notation

2.1. General

| X | denotes the cardinality of a set X, and x y ¯ denotes a straight line segment with endpoints x and y. d ( x , y ) = | | x y | | denotes the Euclidean distance between points x and y, and  d ( x , Y ) = inf y Y { d ( x , y ) } denotes the distance between point x and set Y.
Given a list T of non-negative real numbers t 1 , t 2 , , t k , not all of which equal 0, a normalization of T is obtained by multiplying each value in T by 100 t max , where t max = max 1 i k { t i } .
Given an approximation algorithm A for a minimization problem Π with a set of instances D Π , let A ( I ) be the value of an approximate solution on instance I D Π found by A. The approximation ratio of A on I is R A ( I ) = A ( I ) O p t ( I ) , where O p t ( I ) is the optimal solution for I; the worst-case performance ratio of A is R A = sup { R A ( I ) : I D Π } . We will say that an algorithm with performance ratio r finds an r-approximation to the optimal solution. For more details, the reader is referred to [26,27].

2.2. Notions of Theory of Words

The theory of words studies the structural properties of strings composed of letters of a given alphabet and provides algorithms for solving diverse problems defined on strings. Among the most important motivations for the discipline is its relevance to computational biology and, more precisely, to the automated analysis of biosequences. This includes a great variety of problems whose portrayal is beyond the purposes of the present paper. Some avenues of the ongoing research are surveyed in [28,29,30,31,32,33,34].
In the literature, the terms word, sequence, and string are often used interchangeably. A sequence is often defined in mathematics as a function whose domain consists of a set of consecutive integers, and a string over X, where X is a finite set, is often defined as a finite sequence s of elements from X (X is also sometimes called the alphabet). The term word is frequently used as an abstraction of the other two terms. In biology, the prevalent term is biosequence; the DNA sequences that are the subject of our interest in the present work are built from the four letters A, T, C, and G and have finite lengths.
Below, we recall a few basic notions and fix some notations to be used in this paper.
In the string s = s 1 s m over set X, s i is the i th term of s ( 1 i m ), which is some element of X. The number of elements in s is called the length of s and denoted as | s | . If  | s | = 0 , we say that s is the empty string, denoted by λ . By x k we denote k 1 consecutive repetitions of term x in string s.
A s u b s t r i n g of a string s is obtained by selecting some or all consecutive elements of s. More formally, a string v is a substring of a string s if there are strings u and w such that s = u v w (where we may have u = λ or v = λ ).

2.3. Notions of Graph Theory

Let G ( V , E ) be a simple graph, i.e., without multiple edges between any pair of vertices and with no loops (edges that connect a vertex to itself). A simple path in G is a sequence of vertices of G such that no vertex is repeated. A shortest path between two vertices of G is a simple path of a minimum length (number of edges) connecting the vertices. A simple cycle is a closed simple path. G is triangle-free if it contains no triangle, i.e., a cycle of length 3. G is connected if there is a path of edges of G between any two of its vertices. G is sparse if | E | = O ( | V | ) ; otherwise, it is dense. G is bipartite if its vertices can be divided into two parts such that the graph edges can connect vertices only from different parts but not from the same part. A graph is bipartite if and only if it is two-colorable (i.e., its vertices can be colored by only two colors so that any two adjacent vertices are colored by different colors).
Graph G is geometric if its vertices are points of R n and its edges are segments containing graph vertices. A geometric graph G is rectilinear if each edge of G is a straight line segment that is parallel to a coordinate axis.
A graph is called directed if its edges have directions. A directed path in a directed graph is a sequence of vertices such that for each vertex v in the sequence, there is a directed edge pointing to a successor of v in the sequence. A directed graph is connected if there is a directed path between any two vertices. A directed graph is strongly connected if, for any pair of vertices (u, v), there is a directed path from u to v and a directed path from v to u.

3. Theoretical Background: String Geometrization, Monotone Paths, Monotone Graphs, and Negentropy

3.1. Monotone Paths

In order to make the paper self-contained, in this section, we first recall some basic definitions and related properties from [23,24,25]. Then, we extend those considerations by introducing new structures and concepts to be used in DNA sequence interpretation.
Let θ be the origin of the Cartesian coordinate system. A monotone path in Z + n is a sequence of points (nodes) a 0 = θ , a 1 , a 2 , , a m Z + n , connected by segments that are unit edges of the rectangular grid, where the coordinates of any point a i , i 2 , are pairwise greater than or equal to the corresponding coordinates of any preceding point.
Let L = q 0 q m , q 0 = θ be a monotone path of length m. The segment q 0 q m ¯ is determined by the total composition of the string. It is the “ideal” (or “average”) direction of the path determined by base composition and stands for linearity. The actual path L fluctuates around that line. d ( q i , q 0 q m ¯ ) measures how much the partial composition at position i deviates from the final composition “direction” determined by the segment q 0 q m ¯ , i.e., how much it deviates from linearity.
Let L = q 0 q m be a monotone path of length m. If, for every node q of L, the voxel centered at q intersects the line segment θ p ¯ , then we say that L is a linear path or, equivalently, that L exhibits linearity. In the degenerate case where a monotone path is aligned with one of the coordinate axes, L is called an inline path.
It is easy to see that the following facts hold:
Fact 1.
Given a point p Z + n , there is at least one linear path from θ to p.
Fact 2.
If L is a linear path from θ to p, then d ( q , θ p ¯ ) n 2 q L .
Remark 1.
Note that when n = 4 (which is the case for DNA sequences), the  a d v and m d v of a linear string are, at most, 1.
We define the maximum deviation from linearity of L as
m d v ( L ) = max i = 0 m { d ( q i , q 0 q m ¯ ) } ,
and the average deviation from linearity of L as
a d v ( s ) = i = 0 m d ( q i , q 0 q m ¯ ) m + 1 .
These two characteristics will play a crucial role in our experiments and in interpreting the results we obtain.
The third characteristic of a string s will be called the number of local maxima of s and is denoted as n l m ( s ) . Commonly, by a local maximum one means a point q i L for which d ( q i , q 0 q m ¯ ) is greater than d ( q i 1 , q 0 q m ¯ ) and d ( q i + 1 , q 0 q m ¯ ) . However, regarding the usual applications of extrema of discrete functions—in particular, in view of our own purposes—counting all such maxima does not seem to be very relevant. Instead, local maxima can be counted only if they “stand out" compared to other, “indistinguishable” local extrema, which differ very little from neighboring points. Thus, we adopt the notion of the number of local maxima as “method-dependent”. Specifically, our choice of method is MATLAB’s peakfinder function provided in [35]. This is a computational method for identifying local maxima in a numerical sequence by evaluating changes in slope or curvature under specified thresholds. It defines a peak as a local extremum that exceeds its neighbors by at least a threshold (named sel), which is user-specified or designated by default. The algorithm isolates local maxima whose values exceed their neighbors according to that criterion. More specifically, the algorithm maintains two running values: m x , which is the highest value found since the last minimum, and  m n , which is the lowest value found since the last maximum. These values are updated constantly as the function rises and falls. A minimum is recorded whenever the signal switches from falling to rising, and a peak is accepted only if m x m n sel .
Let p = ( p 1 , , p n ) be a point in Z + n . Let H p be the set of all monotone discrete paths between θ and p. It is easy to see that the following holds.
Fact 3.
| H p | = ( p 1 + + p n ) ! p 1 ! p n ! .
Each path L H p consists of 1 + i = 1 m p i points, with the initial point denoted as θ and the terminal point denotes as p. Now, let H ( m ) be the set of all monotone paths of length m. Then, we have
Fact 4.
H ( m ) = p H p , where p 1 + + p n = m .

3.2. String Geometrization

Let s = s 1 s m be a string on an alphabet X = { x 1 , x 2 , , x n } . We inductively construct an ordered set L ( s ) of points q 0 = θ , q 1 , , q m corresponding to string s as follows.
Let q i = ( q i , 1 , , q i , n ) be the i th element of L ( s ) for 0 i < m . If s i + 1 = x j for some j, 1 j n , then we set q i + 1 = ( q i , 1 , , q i , j + 1 , , q i , n ) . Thus, we obtain a monotone discrete path L ( s ) with | L ( s ) | = | s | = m + 1 associated with string s, in which the coordinates of a point are pairwise greater than or equal to the corresponding coordinates of any preceding point. Figure 1 (left) gives an example of a string over the alphabet { x , y } and its corresponding monotone discrete path.
With reference to Formulas (1) and (2), we can define the maximum deviation from linearity and average deviation from linearity of s as m d v ( s ) = m d v ( L ( s ) ) and a d v ( s ) = a d v ( L ( s ) ) . We call s a linear string if its corresponding monotone path L ( s ) is linear. If L ( s ) is inline, then s is an inline string.
Denote by m a x m d v ( s ) the greatest m d v attained by a permutation of string s, i.e., 
m a x m d v ( s ) = max { m d v ( s ) : s is a permutation of s } .
The following theorem (stated in [23] without proof) will be used in Section 4 in the deliberation on “borders of life”. The proof of the theorem is given in Appendix A.
Theorem 1.
Given a string s over an alphabet X = { x 1 , , x n } in which letter x i appears a i times, 1 i n ,
max m d v ( s ) = max P ( i N a i 2 ) ( j Z a j 2 ) k = 1 n a k 2 ,
where P is the set of all partitions of { 1 , 2 , , n } into two disjoint, nonempty sets N and Z.
If P = { N * , Z * } P is a partition for which max m d v ( s ) is attained, then for every permutation s of s, the corresponding discrete monotone path contains the point expressed as p = ( p 1 , p 2 , , p n ) , where p i = a i if i N * and p i = 0 if i Z * , m d v ( s ) = max m d v ( s ) .
Remark 2.
Given a string s of length m, the numbers a 1 , , a n of appearances of the letters x 1 , , x n can be counted with O ( m ) operations. Once this is done, max m d v ( s ) can be computed in O ( 2 n ) time, i.e., the overall solution of the considered problem takes O ( m + 2 n ) time. While this is exponential for an unbounded n, for a fixed n—as in the case of DNA sequences—the computation time is linear in the string length.
Having defined our geometric descriptors, let us also take notice of a property that underlies their use as tools for analyzing the geometric organization of genomic sequences. Our geometric descriptors are intentionally invariant under permutations of the nucleotide alphabet. This reflects the fact that the framework is designed to quantify large-scale compositional structure rather than biochemical identity. Arbitrary global relabelings of nucleotides do not occur in nature, but they provide a useful mathematical perspective: they show that the descriptors capture structural organization, such as long-range heterogeneity, domain architecture, and compositional drift, rather than local functional semantics. Many widely used genomic measures (e.g., GC-content, k-mer spectra, and codon-usage indices) share this invariance and remain biologically informative because they characterize statistical structure rather than molecular meaning. Our goal is not to encode functional annotation but to analyze the geometric organization of genomes across species, for which alphabet-invariant descriptors are both appropriate and revealing.

3.3. Relation to Minimum Enclosing Cylinder

Our definition of a linear string refers to a discrete monotone path whose voxels intersect the line segment between the initial and terminal points. Correspondingly, deviation from linearity refers to the distances from the points of the monotone path to that line segment.
Another possible approach is to consider a straight line that minimizes the maximal distance over all points of L ( s ) , which is the axis of the minimum enclosing cylinder for set L ( s ) . Note that, while in two dimensions, the problem of finding the minimum enclosing cylinder can efficiently be solved in linear time, it is not so in higher dimensions. Even in three-dimensional space, the available exact algorithms take super-cubic time (see, e.g., [36,37,38]). This makes the problem practically intractable for strings of considerable size, e.g., the DNA sequences we investigate in the following sections. The following theorem demonstrates that the deviation from linearity that we adopt is no more than twice as great as the one defined by the minimum enclosing cylinder. Moreover, the computation of the former requires only a linear number of operations for a fixed dimension n (that is 4 in the case of DNA sequences) and is therefore, without a doubt, advantageous from a computational complexity perspective. Note also that in the course of our experiments, we measured a significant difference between the deviation from linearity of DNA sequences and random sequences; thus, a minimum enclosing cylinder approach—provided that one could afford to wait for the solution—would provide no advantage in distinguishing DNA sequences from random sequences aside from changing the magnitude of distinction by, at most, a factor of 2.
Theorem 2.
Let L = q 0 q m Z n be a monotone discrete path with a length of at least 3. A 2-approximation to a minimum enclosing cylinder for L can be found with O ( m n ) operations.
The proof is given in Appendix A.

3.4. Monotone Path and Irregular Discrete Helix

The monotone paths introduced and discussed in the previous sections can be regarded as irregular discrete helices. Figure 2 gives examples of two monotone paths in Z + 3 . The one on the left resembles a discrete analog of a standard cylindrical helix, which is known as a triangular helix or regular skew-apeirogon. The one on the right features a similar (although irregular) helical structure, which we call an irregular discrete helix. Clearly, in dimensions higher than three, relevant illustrations are not possible. In dimension four, one can obtain certain evidence by viewing the projections of the monotone path on the four coordinate planes.

3.5. Negentropy of a Monotone Path

Broadly speaking, the term “negentropy” (short form of “negative entropy”) is interpreted as a trend towards creation, maintenance, and increasing order within a system. For this, “free energy” is utilized (which, according to the theory of thermodynamics, is energy available to perform work). In biology, negentropy hints at the capability of living organisms to create, maintain, and increase their internal order and complexity by utilizing free energy from the environment through metabolism and adaptation (thereby resisting the second law of thermodynamics).
For the purpose of adding the probabilities of independent events, the entropy is customarily defined as a logarithmic function. Likewise, we can define the negentropy of a monotone path L as
N ( L ) = log 2 ( 1 + m d v ( L ) ) ,
or, alternatively, as 
N ( L ) = log 2 ( 1 + a d v ( L ) ) .
We have N ( L ) 0 , as  N ( L ) = 0 if and only if L is an inline path.
In turn, the negentropy of a string s is
N ( L ( s ) ) = log 2 ( 1 + m d v ( L ( s ) ) ( resp .   N ( L ( s ) ) = log 2 ( 1 + a d v ( L ( s ) ) ) .
Let S be a set of strings and L ( S ) be the corresponding set of monotone paths. Then, the total negentropy of L ( S ) is
N ( L ( S ) ) = s S log 2 m d v ( L ( s ) ) ( resp .   N ( L ( S ) ) = s S log 2 a d v ( L ( s ) ) ) .
Note that the logarithmic function is strictly monotone; therefore, both N ( L ) and m d v ( L ) (as well as a d v ( L ) ) are fairly good for comparative analysis of strings. Thus, as an alternative, one can measure the negentropy of a monotone path L according to its maximum (or average) deviation from linearity. As in the definition using a logarithmic function, m d v ( L ) = 0 (resp. a d v ( L ) = 0 ) if and only if L is an inline path.

3.6. Monotone Graphs

Let L be a set of monotone paths in Z n , n 2 , and let G ( L ) be their union. G ( L ) can be regarded as a geometric graph, called monotone graph with vertices–the nodes of the paths, and edges–the grid edges of the paths in L .
The monotone paths of L are subgraphs of G ( L ) . If we assign directions on the edges of a path L L , starting from the origin, we obtain a directed path L with the first vertex at the origin. Taking the directed paths L for all L L , we obtain a directed graph G ( L ) .
  • Properties of  G ( L )   and  G ( L )
Property 1.
G ( L ) is connected. G ( L ) is connected but not strongly connected and has no directed cycles.
Property 2.
In G ( L ) and G ( L ) , all simple monotone paths between any two vertices have the same length.
Property 3.
G ( L ) and G ( L ) are rectilinear. According to Property 2, all cycles of G ( L ) are of even length. Hence, G ( L ) is bipartite (and, thus, 2-chromatic).
Property 4.
G ( L ) and G ( L ) are sparse. (The proof is given in Appendix A).
Property 5.
G ( L ) is triangle-free (follows from Property 3).
Note: It is known [39] that a triangle-free graph G ( V , E ) with sufficiently many edges (e.g., of the order of Ω ( n 2 ) ) is bipartite; otherwise, it may be non-bipartite. Monotone graphs are always bipartite, regardless their low density.
Property 6.
Both in G ( L ) and G ( L ) , one can find a shortest path from the origin to any vertex v in O ( p ) time, where p is the length of the shortest path (searching backward from v to the origin). Clearly, all of them have the same length.
Property 7.
Since any subgraph of a bipartite graph is bipartite, for any vertex v in G ( L ) , the subgraph reachable from v within G ( L ) is a monotone graph.

4. Graph of Life

Graph-based methods are central in evolutionary biology for the study of interactions among proteins, genes, and DNA sequences (e.g., [40]). Previous monographs [41,42,43,44] provide a broad conceptual and methodological foundation for the use of graph theory in the biosciences. They collectively introduce the mathematical and computational principles underlying biological network analysis, ranging from metabolic and regulatory systems [42] to evolutionary and ecological interactions [44]. Foundational overviews of cellular organization through network structure [41] and general network theory with strong biological applications [43] further situate graph-based methods as central tools in modern biological research. A substantial body of research papers—too extensive to report in full here—has been developed on the subject, elucidating various facets of biological research. For instance, refs. [45,46,47,48] illustrate how graph-theoretic approaches have shaped contemporary systems biology. Scale-free network models [45] and centrality-based analyses of essential genes [46] demonstrate how structural properties of graphs reveal functional and evolutionary constraints. Methods for detecting overlapping communities [47] and for predicting protein function through network propagation [48] show how graph algorithms can extract biologically meaningful patterns from large-scale interaction data. Graph-based and graph grammar approaches to the modeling of structural determination, development, and evolution are investigated in [49].
Perhaps most widely used in the modeling of biological structures are bipartite graphs. For example, refs.  [50,51,52] highlight the particular relevance of bipartite graphs for the representation of heterogeneous biological relationships. Bipartite protein-complex networks [51] reveal modularity and organizational principles of the proteome, while bipartite gene-function mappings [50] support functional annotation and evolutionary inference. Bipartite transcription-factor target graphs [52] provide insight into regulatory architecture and its evolutionary constraints. Bipartite species-interaction networks [53] demonstrate how ecological and evolutionary structure emerges from mutualistic relationships, illustrating the versatility of bipartite models across biological scales. For applications of bipartite graphs in evolution and gene flow, as well as recent work on best-match graphs in phylogenetics, the reader is referred to [54].
In what follows, we define a graph in four-dimensional Euclidean space, which can be viewed as a geometric map of the biosphere’s viability region in nucleotide composition space.
DNA molecules consist of the adenine (A), guanine (G), cytosine (C), and thymine (T) nucleotide bases and can be represented as strings over the alphabet Σ = { A , G , C , T } . Through string geometrization (Section 3.2), each sequence yields a monotone path L ( s ) Z + 4 whose coordinates correspond to the four bases. Every point of L ( s ) represents a nucleotide position and defines a state in a state space embedded in Z + 4 .
We define Graph of Life as GL = L ( s ) , the union taken over all living organisms. Similar networks can be formed for any biological group. GL is a monotone graph satisfying properties 1–7 and is extraordinarily large. With an estimated 8.7 million species and genomes ranging from millions to billions of bases, its size is of the order of 10 12 to 10 15 . Individual variation further enlarges the network: about 20 quintillion animals existed in 2020, on Earth contains roughly 5 × 10 37 DNA base pairs, giving the scale of GL ’s nodes.
Beyond its size, GL is only partially known and constantly changing. Most species that have ever lived—over 99%—are extinct; new organisms appear and disappear, and mutations continually alter existing sequences. However, the network’s graph-theoretic properties remain stable. Conceptually, GL occupies a five-dimensional hyperspace, with time as the fifth dimension. Although not physically present in 3D space, it models all living organisms from their emergence approximately 4.1 billion years ago until the eventual end of life on Earth.
Note that the same construction applies verbatim to RNA-based organisms: their genomes also define monotone cumulative-count paths and an analogous graph representation with the same graph-theoretic properties as the DNA-based Graph of Life. Geometric descriptors m d v , a d v , and  n l m can likewise be computed for RNA genomes, although their biological interpretation may differ because RNA organisms exhibit distinct compositional biases and evolutionary constraints.

4.1. Evolutionary Interpretation of the Graph of Life

The “Graph of Life” can be viewed as a compressed portrayal of genomic diversity. If every genome contributes a monotone path in Z + 4 , then the union of all these paths forms a gigantic geometric object encoding the global structure of genomic variation. This object is sparse (Property 4), which means that genomes occupy only a micro fraction of the theoretically possible space. In biological terms, this reflects strong restrictions: only a tiny subset of all possible DNA sequences is compatible with life. Sparsity is a mathematical way of demonstrating the idea that life traverses a thin, structured region of the entire sequence space.
Concurrently, the Graph of Life is bipartite (Property 3). Biologically, this suggests that the transitions between genomic states (as represented by monotone paths) feature a sort of hierarchical or directional constraint. This seems to be harmonious with evolution: genomes do not roam arbitrarily. Instead, they follow constrained trajectories shaped by selection, mutation biases, and biochemical feasibility.
The computational inferences of this construction are consequential. Sparse and bipartite graphs admit efficient algorithms for various problems that are NP-hard to solve on general graphs. Examples of such tasks are clustering (e.g., detecting groups of species), computation of distances between genomes (e.g., evolutionary divergence), ancestral reconstruction, and detection of structural “bottlenecks” of a biological system. These properties of the Graph of Life imply that wide-ranging comparative genomics may be more tractable than the underlying combinatorial complexity of the sequence space would suggest. In this sense, the graph provides not only a descriptive framework but also a computationally advantageous representation of genomic relationships. One could also suppose that the evolution may be navigating a space of DNA sequences whose structure makes adaptation feasible rather than combinatorially impossible.
Note also that in a dense, unconstrained sequence space, almost all random mutations would most probably be catastrophic. However, in a sparse and structured (bipartite) Graph of Life, neighboring genomes would tend to remain viable, with evolution proceeding through a network made of “safe” paths; as a result, catastrophic gamboling would naturally be avoided because they would lie outside the sparse region of DNA sequences. In other words, the geometry of the Graph of Life may reflect the fact that life evolves in a region of sequence space where survival is statistically stable.
Conceptually, the Graph of Life can be viewed as a geometric map of the biosphere’s viability region in sequence space. Its structure reflects the existence of “safe corridors” through which genomes can evolve while maintaining functionality. This perspective aligns with the broader view that evolution operates within a constrained environment shaped by biochemical feasibility and selective pressures. Thus, the graph offers a compact and informative representation of the deep structural constraints that govern genomic evolution, providing a foundation for future work on the geometry of sequence space, the dynamics of evolutionary change, and the global organization of genomic diversity.
The Graph of Life could be seen as a global viability manifold, which is a low-dimensional and structured subset of the astronomically large space of all DNA and RNA sequences where life can actually exist. The “good properties” of the graph can be viewed as a mathematical reflection of the fact that life has evolved within a constrained, navigable region of sequence space. Thus, life is possible because the underlying genomic space has a structure that the Graph of Life captures: sparse, constrained, and algorithmically passable rather than chaotic. See Figure 3 for a conceptual illustration.

4.2. Helical and Non-Helical Structures and Analysis Using Projections

DNA is not structurally uniform. The familiar Watson–Crick B-DNA double helix is the dominant form, but modern research has shown that DNA can adopt many alternative, experimentally confirmed structures under physiological conditions depending on the sequence, stress, and environment. These “non-canonical” forms are not rare curiosities. They appear in real genomes and often play regulatory or structural roles. Some alternatives remain helical (A-DNA, Z-DNA, and some triplexes), but many are not helices at all. For example, A-DNA is a compressed, right-handed helix often found in DNA–RNA hybrids. Z-DNA is a left-handed helix forming in GC-rich, torsionally stressed regions. G-quadruplexes are four-stranded stacks of planar G-tetrads—geometrically columns, not helices. i-Motifs are interlocked C-rich folds with no helical axis. Hairpins, cruciforms, and slipped structures form loops, branches, and misaligned segments rather than spirals. These structures appear in specific genomic contexts, such as telomeres, promoters, repetitive regions, and replication origins. Evolution preserves these non-helical shapes because they serve functional roles: regulating transcription, relieving torsional stress, protecting chromosome ends, controlling replication, shaping 3D genome architecture, and generating adaptive variation. They act as molecular switches, stress absorbers, binding platforms, and structural organizers. For more details, see [55,56].
For all the above alternative possibilities, the 4D monotonic path encodes sequence composition, not physical molecular geometry, i.e., it provides a mathematical encoding of the sequence, not a physical model of the molecule. The real 3D structure of DNA depends on various factors, such as electrostatic interactions, the hydration shell, ionic environment, protein binding, supercoiling, torsional stress, chromatin packing, temperature, pH, and mechanical forces inside the cell. These are not contained in the sequence alone, and none of them is represented in a monotonic 4D walk.
Note, however, that the 3D projections of the 4D monotone path on the four 3D coordinate planes can reveal informational geometry that correlates with structural and functional DNA behavior, including biologically meaningful features that are in tune with structure, function, and genomic behavior. Such features include the order of local nucleotide composition, long-range patterns, mirror-like local symmetry signals, and palindromes, which can form hairpins or cruciforms, periodicities, biases (GC-rich vs. AT-rich), motifs, and repeats. For example, strong drift in one coordinate highlights compositional bias (e.g., G-rich, implying potential G-quadruplex regions). Straight or nearly straight segments indicate tandem repeats or microsatellites, often linked to slipped-strand structures. Sharp bends or slope changes mark transitions between GC-rich and AT-rich domains, often aligning with regulatory boundaries. Dense, compact clusters in mixed C/G steps correspond to CpG islands and promoter regions. Periodic oscillations reflect structured repeats with known instability profiles. Large-scale curvature patterns can reflect isochores or replication domains. Thus, while projections cannot show the physical 3D shape of DNA, they reveal the sequence-encoded architecture that influences real DNA folding, stability, and function.
Projection from a 4D monotonic path into 3D inevitably removes one coordinate, but this loss does not undermine the biological conclusions one could draw from the geometry. The structural signals we care about, such as G-richness, palindromes, repeats, and domain boundaries, are encoded in patterns of cumulative drift rather than in the exact four-dimensional point. These patterns are geometric and remain visible, even when one axis is dropped. Features such as steep slopes, mirrored segments, straight lines, sharp bends, and compact clusters survive projection because they depend on relative changes among the remaining coordinates. Different projections can be used to recover different aspects of the sequence, much like viewing an object from multiple angles. The lost coordinate rarely contains unique structural information, since many biologically meaningful patterns arise from correlations among bases rather than from a single base count. The method is not intended to reconstruct the sequence but to detect large-scale compositional and structural tendencies. Because these tendencies produce strong geometric signatures, the conclusions remain reliable despite the dimensional reduction.
The 3D projections of a 4D monotonic path do not add new biological information, but they reorganize the sequence into a geometric form where important patterns become far easier to detect. Many meaningful features in DNA, such as G-rich stretches, palindromes, repeats, and GC/AT domain shifts, are present in the raw sequence but are visually subtle, scattered, or overlapping. When the sequence is converted into cumulative geometric drift, these same features produce strong shapes like steep slopes, mirror-symmetric curves, straight or periodic segments, sharp directional bends, or compact clusters. This transformation suppresses local noise and highlights long-range structure that is nearly impossible to see by reading the letters alone. It also allows different sequences with a similar functional architecture to produce similar geometric signatures, making comparisons intuitive. Palindromes and inverted repeats become literal reflections, and compositional transitions appear as abrupt changes in trajectory. Therefore, the projections serve as a kind of visual microscope for sequence architecture, revealing structural tendencies that remain hidden in symbolic form.

5. Negentropy of Species and Groups of Species

Most theories of the origin of life—RNA World, Oparin-Haldane [57,58], Metabolism First, and related scenarios—begin with some form of “primordial soup” from which the first organic molecules emerged. In this shared framework, simple inorganic compounds (water vapor and carbon-, hydrogen-, and nitrogen-containing species such as ammonia or methane) reacted to form the earliest organic monomers. Among these were canonical nucleobases adenine, guanine, cytosine, thymine, and uracil, which later assembled into RNA and DNA.
At this initial stage, these molecules did not exhibit any organized arrangement. They were not yet connected into macromolecules but existed as disordered collections of monomers, with no structural constraints and no long-range correlations. Their spatial and compositional distribution was effectively random. In this sense, on the eve of life’s emergence, organic matter existed as chaotic, unstructured aggregates of nucleobases. Such a state can be modeled, at the level of sequence statistics, by a random string of nucleotides. Its entropy is therefore maximal (and its negentropy minimal) compared with the structured genomes of later living organisms.
It is worth noting that low entropy alone does not imply biological relevance. A perfectly ordered system, such as a crystal at absolute zero, has zero entropy, but such a configuration is incompatible with life. In the language of sequence entropy introduced earlier, this would correspond to a string over a single-letter alphabet whose monotone path is perfectly linear and exhibits no deviation. Life requires structure but not perfect order. Chemical reactions in the prebiotic environment generally increase the entropy of the surrounding inorganic matter. However, these same processes can decrease the entropy (i.e., increase the negentropy) of the organic subsystem by producing more complex, more structured molecules. This reflects a fundamental thermodynamic trade-off:
An increase in the entropy of some inorganic matter was exchanged for a decrease in the entropy (increase in the negentropy) of certain organic matter.
This trade-off enabled the emergence of increasingly organized molecular structures, ultimately setting the stage for the appearance of life.
The definitions of Section 3.5 apply directly to the negentropy of DNA sequences and assemblies of such sequences. The negentropy of a DNA sequence s is
N ( L ( s ) ) = log 2 ( 1 + m d v ( L ( s ) ) ) ( a l t e r n a t e l y ,   N ( L ( s ) ) = log 2 ( 1 + a d v ( L ( s ) ) ) ) ,
where L ( s ) is the monotone path corresponding to s. Given a set S of a DNA sequences (in particular, S may be the entire Graph of Life GL ) and the corresponding set L ( S ) of monotone paths, the total negentropy of L ( S ) is
N ( L ( S ) ) = s S log 2 m d v ( L ( s ) ) ( a l t e r n a t e l y ,   N ( L ( s ) ) = s S log 2 m d v ( L ( s ) ) ) .
Let us reiterate that because of the monotonicity of the logarithmic function, both N ( L ( s ) ) and m d v ( L ( s ) ) (resp. a d v ( L ( s ) ) can be used for any comparative studies concerning DNA sequences of species.

6. Experimental Study of DNA Sequences Grounded in Deviation from Linearity

The concept of string linearity provides a straightforward and practical method for comparing the DNA sequences of different organisms. It is natural to hypothesize that organisms with greater biological complexity possess more structured DNA whose associated monotone path deviates significantly from a straight line. In contrast, simpler organisms likely have less structured DNA, producing monotone paths that remain closer to linear. A fully random sequence over the alphabet {A,T,C,G} contains no inherent structure, so its monotone path should lie even nearer to a straight line.
To examine this hypothesis, experiments were conducted in two stages: first on an initial baseline panel of 25 species plus a random sequence, then on an expanded set of 50 species after adding 25 additional genomes and a random sequence. Reporting both stages allows us to show that the observed patterns are not artifacts of a small dataset but persist and even sharpen when the taxonomic breadth is doubled. This presentation highlights that the emerging geometric trends are robust, reproducible, and stable under substantial expansion of phylogenetic diversity.
In Stage 1 (25 species; see Table A1), all DNA sequences were obtained from Ensembl Genomes, a genome-scale repository and browser maintained by the European Bioinformatics Institute, using the most recent assemblies available at the time the analyses were performed. The specific analyzed strains were: Chlamydia (Nigg), Tuberculosis (CCDC5180), Gingivalis (W83), and Streptococcus (ND03). All DNA samples were taken from chromosome 1 of each organism, with the exception of fruit fly and yeast, for which chromosomes 2L and 4 were used, respectively. For the additional 25 species included in Stage 2 (see Table A3), sequences were obtained from a range of publicly available repositories (including Ensembl and other FASTA sources), selecting the most clearly annotated chromosomal or scaffold sequence available for each species. Together, these 50 genomes provide a representative and sufficiently diverse set for assessing the geometric properties of cumulative paths across species.
For each organism, we extracted relatively short substrings and subsequences in FASTA format, chosen at random from a genomic excerpt of approximately two million characters. We assumed that excerpts of this size are sufficiently large to mitigate the effects of genomic non-stationarity, though future work may involve sampling from complete genomes.
Our first step was to analyze how string length influences the proposed linearity measures and to determine an appropriate length for more extensive experiments. The results presented in the next section show that comparing samples of a reasonably chosen size yields the same relative ordering of organisms as comparisons based on much larger samples. Fixing the sample length also simplifies comparisons, since the absolute deviation from linearity depends on string length. Because genomes vary widely in size—and some remain incomplete or only partially studied—direct comparisons of full genomes are not feasible.
We also distinguish between substrings and subsequences of a string s: substrings consist of consecutive characters, whereas subsequences are ordered selections of characters that need not be consecutive. Most experimental work on biosequences focuses on substrings, but Apostolico and Cunial have recently explored the structure and randomness of polypeptides using subsequences satisfying specific constraints [29,30]. Their promising results motivated us to conduct all experiments using both substrings and subsequences. The similarities and differences between these two approaches are examined in the following sections.

6.1. Effects of Substring and Subsequence Length

Before detailing our research procedure, we explain our choice of data. We investigated how the length of a DNA sequence affects deviation from linearity and ascertained that deviation increases with the size of the string. However, the rate of increase appears independent of the organism, meaning that an organism’s deviation from linearity relative to others remains stable as long as all samples have the same sufficiently large length. Initial fluctuations disappear around lengths of approximately 50,000 bases, after which m d v and a d v grow smoothly and predictably. For this reason, all subsequent analyses use substrings with a length of 50,000, which is long enough to suppress small-scale noise while remaining applicable to species with incomplete or unevenly assembled genomes. Using entire genomes would introduce confounding effects due to large differences in genome size and completeness, and focusing on specific genomic regions would shift the analysis toward local functional or structural features. Because our goal is to compare species as a whole, we average m d v and a d v over 1000 uniformly chosen substrings per species, providing a stable estimate of each organism’s global compositional geometry rather than region-specific effects (e.g., taking care of regions that span centromeres, telomeres, microsatelites and other low-complexity regions) that would belong to other types of studies.
Figure 4 supports this observation by showing the absolute and normalized maximum and average deviations from linearity for substrings and subsequences of increasing length for our controlled baseline set of 25 species used in Stage 1. The top four graphs display the a d v and m d v values for substrings, while the bottom four show the same measures for subsequences. The left column presents absolute values, and the right column shows the corresponding normalized values. The normalized graphs essentially “lift” the left sides of the absolute-value graphs, making the curves more directly comparable.
All substrings and subsequences were selected randomly from the 26 sources listed in Table A1. For the substring graphs, linearity measures were computed for lengths of 10 , 000 × k for 1 k 20 . For each length, 200 random substrings were drawn from each source, and the average of these trials was plotted. For the subsequence graphs, lengths were 10 , 000 × ( 2 k 1 ) for 1 k 10 , with 100 randomly chosen subsequences per source; the averages of these trials were plotted. (Fewer lengths and trials were used for subsequences because additional resolution would not provide meaningful new information, given the near-linearity of the resulting graphs).
Figure 4 shows that a d v and m d v derived from subsequences depend even less on length than those derived from substrings, as indicated by the reduced crossing of lines in the lower graphs. Nonetheless, even for substrings, the measures remain largely length-independent: across all tested lengths, organisms maintain roughly the same relative ordering.
As substring and subsequence lengths approach zero, deviation from linearity naturally approaches zero as well, and the measures become less stable. The normalized graphs indicate that initial fluctuations disappear around lengths of 50,000. For this reason, all subsequent experiments (which involve more trials) use substrings and subsequences with lengths of 50,000.

6.2. Description of Computational Procedure

The computations were performed using MATLAB R2011a for Stage 1 and MATLAB R2025b for Stage 2. We used the built-in randi and randseq functions to generate random selections and random sequences with independent, uniformly distributed symbols. We also used the peakfinder function from [35] to identify local extrema. The computational workflow consisted of three main components:
  • Selection of samples;
  • Computation of linearity measures;
  • Compilation and normalization of data.

6.2.1. Selection of Samples

In both stages, to extract a random substring b 1 , , b m from a larger string a 1 a n , we selected a random integer i [ 1 , n m + 1 ] using randi and set b j = a i + j 1 for 1 l e q j m . Substrings of length m from each organism were stored in an array, along with a random sequence of length m over {A,T,C,G} generated using randseq. To obtain a random subsequence b 1 , , b m , we generated a list C = c 1 c m of m distinct random integers between 1 and n, sorted them to obtain c 1 < < c m , and set b i = a · c i for 1 i m . These subsequences were stored in an array alongside a random sequence of the same length.
Because earlier analysis showed that fluctuations in linearity measures stabilize at a length of 50,000, all substrings and subsequences used in later experiments were chosen to have this length.
Because our geometric statistics are computed on long substrings (50 kb) and averaged over 1000 randomly sampled segments per species, they are expected to be robust to minor differences between successive assembly versions or small local corrections in updated releases. The focus of this work is on large-scale geometric organization of genomes rather than on fine-grained annotation; therefore, the precise assembly version is not expected to materially affect the qualitative patterns reported here. Using updated assemblies with explicit accession numbers would further enhance reproducibility, but this falls outside the scope of the present study.

6.2.2. Computation of Linearity Measures

For a string S of length m, we first count the occurrences x l of each letter l { A , T , C , G } , so that x l = m . Let S t denote the prefix of S of length t, and let t l be the number of occurrences of letter l in S t , with  t l = t . For each 1 t m , we compute the distance D ( S t ) from the point ( t A , t T , t C , t G ) to the line through the origin and ( x A , x T , x C , x G ) with the following formula:
D ( S t ) = ( x A k t A ) 2 + ( x T k t T ) 2 + ( x C k t C ) 2 + ( x G k t G ) 2 where k = x A t A + x T t T + x C t C + x G t G x A 2 + x T 2 + x C 2 + x G 2 .
The D ( S t ) values are stored in a list, from which we compute the average and maximum deviations, as well as the number of local maxima, using peakfinder.

6.2.3. Compilation and Normalization of Data

The two stages follow the same procedure. In Stage 1, for each set of 26 samples (25 biological sequences and one random sequence), we computed the linearity measures a d v , m d v , and  n l m , producing a 26 × 3 table. This entire process was repeated 1000 times with different random samples. We then calculated the mean and standard deviation for each measure across the 1000 trials and normalized the mean values to highlight relationships among organisms. We also computed two measures of statistical dispersion: the coefficient of variation (CV), defined as the standard deviation divided by the mean, and the quartile coefficient of variation (QCV), defined as ( Q 3 Q 1 ) / ( Q 3 + Q 1 ) , where Q 1 and Q 3 are the first and third quartiles. This procedure yields a table of normalized means that enables straightforward comparison of organisms across the three linearity criteria, along with a table summarizing variability across trials. The entire analysis was performed twice—once using substrings and once using subsequences. The results appear in Table A1 and Table A2.
Stage 2 repeats the same procedure on an expanded set of 50 species and a random sequence. The results appear in Table A3 and Table A4.

7. Basic Results

This section expands on our experimental procedures and examines the results in greater depth. We highlight patterns that are readily apparent while acknowledging that additional insights may emerge for readers with more extensive expertise in the biological sciences.

7.1. General Observations and Comments

The first two columns of Table A1 (resp., Table A3) list the scientific and common names of the organisms included in our study. Columns 3–5 present the values of a d v , m d v , and  n l m computed from substrings of the DNA sequences and one random sequence. The final three columns report the same measures computed from subsequences rather than substrings. Each measure was calculated on samples of length 50,000, averaged over 1000 trials.
  • Stage 1. For most species, the normalized average values of a d v and m d v derived from substrings differ by less than 1%. Only four organisms show differences greater than 2%, and none exceeds 4.2%. When subsequences are used, the discrepancy between a d v and m d v is more variable: although typically under 2%, six organisms exceed 4%, and a few show differences as high as 10–20%. Table A2 further shows that the CV and QCV of a d v and m d v differ by no more than 0.02 for most organisms and never by more than 0.06, regardless of whether substrings or subsequences are used. Figure 2 reinforces these similarities, as the a d v and m d v curves are nearly indistinguishable, though slight differences appear more clearly in the subsequence-based plots.
Distributions with coefficients of variation below one are generally considered low-variance. As Table A2 indicates, the variance across the 1000 trials is, indeed, low, with subsequence-based trials showing substantially lower variance than substring-based trials.
We hypothesize that analyzing subsequences rather than substrings removes some structural information from the DNA, reducing precision. However, subsequences are less sensitive to sample length and exhibit lower variance. Substrings, by preserving the full local structure of the DNA, appear to yield linearity measures that more accurately reflect organismal complexity.
As shown in the fifth and eighth columns of Table A1, the  n l m measure effectively distinguishes random sequences from biological ones, but it is less informative for ranking organisms by biological complexity.
  • Stage 2. Turning to Stage 2, we find that the expanded set of species exhibits qualitative patterns similar to those in Stage 1, thereby strengthening the overall conclusions.
The normalized differences between a d v and m d v remain small for most organisms (at most, 4 and 1.47, on average), and subsequence-based values show somewhat larger discrepancies (at most 19.3 and 2.42 on average), just as in the smaller dataset. The corresponding CV and QCV values, again, indicate low variance across the 1000 trials. For substrings, the largest observed CV was 0.06 (with a mean of 0.02 across species), and the largest QCV was 0.09 (also averaging 0.02). Subsequences show consistently lower variability, with a maximum CV of 0.05 (mean of 0.02) and a maximum QCV of 0.04 (mean of 0.01). These findings reinforce the interpretation that subsequences, while less precise due to the loss of local structure, offer greater stability with respect to sampling length, whereas substrings preserve more of the geometric information relevant to organismal complexity.
As in Stage 1, n l m effectively separates random sequences from biological ones but remains less informative for ordering species by complexity.

7.2. Distinction Between DNA Sequences and Random Sequences

Life is thought to have arisen from organic matter that initially existed as disordered collections of molecules characterized by chaotic, noisy, and highly variable chemistry—conditions that displayed a significant degree of randomness. From this environment, biological sequences gradually emerged through a complex evolutionary process whose precise mechanisms remain only partially understood.
A variety of earlier studies have used quantitative approaches to explore what, if anything, differentiates biological sequences from random ones, with mixed results. Although this objective has proven inherently difficult to pin down [59], considerable evidence suggests that biosequences exhibit many characteristics typically associated with randomness. One of the most striking is their extremely high—almost complete—incompressibility [60]. By many measures, biological sequences are barely distinguishable from their randomly permuted counterparts, even though such permutations are clearly incompatible with living systems [61,62,63]. While this observation may appear straightforward from a biological standpoint, numerous computational studies have also reinforced it. For instance, Pande et al. [64] mapped certain protein sequences onto Brownian bridges and observed subtle departures from pure randomness. Weiss et al. [65] estimated differential entropy and context-free grammar complexity, showing that large collections of non-homologous proteins are roughly 1% less complex than comparable sets of random strings. In the same study, the authors argue that DNA sequences can be viewed as “slightly edited random sequences” and that modern proteins may represent “memorized” ancestral random polypeptides that have been modestly refined by evolutionary pressures to enhance their stability under specific physiological conditions [29].
A key finding of our study is the clear separation between the linearity properties of biological and random sequences. Across all experiments, biological sequences exhibit higher average and maximum deviation from linearity than random sequences. This difference is especially pronounced when substrings are analyzed. In Table A1 referring to Stage 1, the normalized a d v and m d v for the random sequence are both 12.8, whereas the smallest corresponding values among biological sequences are 29.3 and 29.7, respectively. Although the gap narrows when subsequences are used, random sequences still show the lowest deviation from linearity in every case. Figure 2 illustrates this distinction: in the substring-based plots (top row), the curve for the random sequence lies clearly below all others. In the subsequence-based plots (bottom row), the random sequence, again, forms the lowest curve, though the separation is less dramatic. The n l m measure amplifies this contrast even further. For substrings, the normalized n l m of the random sequence is 100.0, while the highest value among biological sequences is 39.8. For subsequences, the random sequence, again, attains 100.0, with the highest biological value being 96.2.
The results of Stage 2 (see Table A3) closely mirror those of Stage 1, confirming the robustness of the geometric trends identified above. More specifically, the normalized a d v and m d v for the random sequence are 12.1 and 11.8, respectively, whereas the smallest corresponding values among biological sequences are 28.6 and 27.6, respectively. Although the gap narrows when subsequences are used, random sequences still show the lowest deviation from linearity. The only exception is the placozoan, whose subsequences yield a d v and m d v values that approach—and, in the case of a d v , slightly fall below—those of random subsequences. This behavior is expected because subsequences eliminate local genomic structure by sampling non-consecutive positions, thereby erasing motifs, repeats, and other contiguous patterns that normally increase deviation from linearity in biological DNA. Once this structure is removed, the descriptors depend primarily on global nucleotide composition. For organisms with relatively uniform and compositionally simple genomes, such as placozoans, subsequences can therefore produce geometric measures that are nearly indistinguishable from or marginally lower than those of fully random sequences.
The n l m measure accentuates the separation between random and biological sequences. For substrings, the normalized n l m of the random sequence is 100.0, whereas the highest biological value is 37.7. For subsequences, the random sequence, again, reaches 100.0, with the highest biological value at 98.3.
Given the approximately 1% difference reported in [65], our experiments revealed discrepancies of several hundred percent when comparing DNA sequences to random sequences constructed over the same alphabet and of identical length. This leads to the following conclusion:
DNA sequences exhibit substantially higher average and maximum deviation from linearity than random sequences.
These striking differences support the following assertion:
DNA sequences cannot be regarded as minor perturbations of random sequences. The degree of departure from randomness is substantial, even though many randomness-related properties—such as incompressibility—have been largely preserved.
See Figure 5 for a conceptual illustration.
We conclude this section with one additional remark concerning the necessity of testing other random models.
Real DNA sequences cannot be captured by any single simple stochastic model. Rather than behaving like one homogeneous i.i.d. (i.i.d. stands for independent and identically distributed sequence of random variables), Markov, or Poisson process, genomes resemble mixtures of many short-range models whose parameters vary along the sequence. Different regions exhibit different base compositions, transition patterns, and local statistical structures, so the overall sequence is only “quasi-random” rather than stationary. Models like the ones listed above can each describe local behavior, but none can account for the full heterogeneity of genomes.
For random sequences with fixed letter probabilities, the monotone path q i , i = 1 , 2 , , m , behaves like a random walk with drift. The segment q 0 q m ¯ is the drift line, and the deviations form a centered fluctuation process and are measured relative to the single global-drift line segment. According to classical limit theorems, these fluctuations scale as m and converge to a Gaussian bridge. Consequently, for a strng s of a length m both m d v ( s ) and a d v ( s ) grow on the order of m , with constants determined by the covariance of the underlying distribution. This m scaling holds for a broad class of short-range dependent models including those listed above because they all satisfy functional central-limit theorems. Therefore, it is not necessary to repeat our experiments for each of these models individually, since they all produce the same m -scale deviations and yield results comparable to the uniform random case.

7.3. Rise of Deviation from Primitive to Biologically Complex Organisms

Given the widely accepted principles of evolutionary theory and the available theoretical and experimental evidence, it is reasonable to hypothesize that as organisms evolved from primitive forms to more biologically complex ones, their corresponding DNA sequences also evolved from sequences that were random or nearly random toward sequences exhibiting progressively greater deviation from randomness.
For the purposes of our analysis, we treated bacteria and microscopic organisms as the most primitive life forms. Plants were considered the next level of complexity, followed by fish, reptiles, and other egg-laying vertebrates. Mammals and primates were placed at the highest end of the evolutionary spectrum. We anticipated that the degree of deviation from linearity would increase in a manner consistent with this hierarchy of biological complexity. Our experiments, based on the measures introduced earlier, largely supported this expectation (though not with the same level of decisiveness observed when comparing random sequences to DNA sequences).
More specifically, in Stage 1, humans, neanderthals, gorillas, and chimpanzees exhibit the highest values of a d v and m d v . At the opposite end, bacterium Chlamydia shows the smallest a d v and m d v after the random sequence. Other organisms with low deviation from linearity include tuberculosis, sea squirt, and lizard (Green anole). Organisms such as the fruit fly, medaka fish, and soybean fall into the mid–low range, while zebrafish, chicken, and mouse occupy the mid–high range (see Table A1 in Appendix B).
Another noteworthy observation is that primates tend to have lower n l m (number of local maxima) values than other animals. Beyond this, however, the  n l m measure appears less effective than the other metrics for distinguishing organisms by biological complexity.
The analysis of Stage 2 produces patterns that remain consistent with those of Stage 1, aside from small variations attributable to the broader species set, thereby confirming the robustness of the geometric trends identified above. The highest a d v and m d v values were observed in dolphin, whale, humans, neanderthals, gorillas, and chimpanzees, whereas Chlamydia, again, shows the smallest values after the random sequence. Other organisms with low deviation from linearity include fission yeast, as well as sea squirt and lizard. Primates, whale, and dolphin typically have lower n l m values than other animals. Beyond this observation, this metric provides limited resolution compared with the other metrics.
See Figure 6 for illustration. Species are ordered by their m d v and n l m values to make the distribution and outliers visually clear. This ordering is purely for readability and does not affect the underlying comparisons.

7.4. Evolutionary Interpretation of Multiscale Genomic Geometry

As predicted by stochastic models, random sequences exhibited the smallest m d v and a d v values, consistent with the m -scale fluctuations of a drifted random walk. In contrast, real genomes showed substantially larger deviations, with a general trend of rising m d v and a d v from more primitive to more complex organisms, albeit with exceptions. These results indicate that genomes of complex organisms contain stronger long-range compositional heterogeneity and more coherent structural domains, producing large divergence away from the global-drift line in the cumulative-count geometry.
Complementing these global measures, the  n l m statistic captured local irregularity of the monotone paths. Here, the trend was reversed: primate genomes exhibited markedly fewer local maxima than those of the other examined species, indicating smoother local trajectories despite their large-scale deviations. This combination—high m d v / a d v but low n l m —suggests that complex genomes deviate strongly from linearity but do so through broad, smooth compositional arcs rather than through short, jagged fluctuations. In contrast, simpler organisms exhibit smaller global deviations but more fragmented local structure. Together, m d v , a d v , and  n l m provide a multiscale geometric characterization of genomic organization, revealing systematic differences in both global and local structure across the evolutionary spectrum.
Within the Graph of Life framework, these findings integrate naturally. Each genome contributes a monotone path embedded in the shared four-dimensional viability region of sequence space. Species with low m d v / a d v and high n l m values (e.g., bacteria) traverse short, locally irregular trajectories, while species with high m d v / a d v and low n l m values (e.g., primates) follow long, smooth, heterogeneous arcs through the same constrained manifold. Thus, the three deviation measures enrich the interpretation of the Graph of Life by quantifying how different lineages navigate the global genomic landscape with distinct geometric signatures. While these patterns are consistent with stronger or more coordinated evolutionary constraints in complex organisms, they do not, by themselves, distinguish the relative contributions of mutational biases, selection, and genomic architecture. Instead, they provide a geometric framework within which such mechanisms can be investigated.
For illustration, see the scatterplot of m d v versus n l m for the 50 examined species and random sequences presented in Figure 7.
  • Exceptions to the Trend and Their Genomic Basis
Although the overall pattern showed increasing m d v and a d v from bacteria to primates, several notable deviations were observed, each reflecting lineage-specific genomic architecture rather than contradictions of the general trend.
A first class of exceptions arises from plant genomes: when measured over substrings, rice and corn show relatively high m d v and a d v values—higher even than those of mouse and rat. This is consistent with the distinctive structure of many plant lineages. Cereals and other large plant genomes often contain extensive transposable element expansions, long blocks of repetitive sequence, and remnants of ancient whole-genome duplications, all of which create broad compositional domains and pronounced long-range excursions in the cumulative-count geometry. These features can inflate global deviation measures independently of organismal complexity.
A second class of exceptions emerged only after expanding the dataset to 50 species. Several mammals outside the primate clade—most notably, dolphin and whale—now occupy the extreme high- m d v / a d v end, surpassing gorilla, which was the top species in the smaller dataset. This shift is biologically consistent: cetaceans are known for unusually homogeneous, compositionally stable genomes. Group-level averages further illustrate these architectural influences: primates show the highest mean m d v and a d v , followed by other mammals, plants, protists, and other vertebrates, whereas invertebrates, fungi, and bacteria occupy progressively lower ranges. These differences reflect well-established genomic properties: mammals and many plants possess long, compositionally uniform domains; protists include lineages with large, smooth genomes, while invertebrates and bacteria typically exhibit more fragmented, heterogeneous sequence organization.
A third apparent exception was the elephant, which showed relatively low m d v and a d v and a comparatively high n l m values. This is consistent with the known architecture of elephant genomes, which are highly repeat-rich and compositionally heterogeneous; moreover, our sampling covered only a portion of the genome, likely capturing a structurally complex region. Thus, the elephant does not represent an anomaly but reflects genuine genomic heterogeneity within mammals.
Taken together, these exceptions highlight that m d v and a d v capture a combination of evolutionary, architectural, and compositional influences, and they underscore the importance of interpreting geometric signatures within the broader context of each lineage’s genomic organization rather than as direct proxies for organismal complexity.
The evolution of different species is influenced by a variety of factors and ultimately represents a sequence of events, some of which occur by chance. This naturally leads to numerous exceptions to general evolutionary trends. At times, such anomalies serve as catalysts for new hypotheses and scientific debate.

7.5. DNA-Based Viruses

As noted earlier, our study focuses primarily on species from the Tree of Life, and viruses are therefore not a central object of analysis. Nevertheless, we conducted additional experiments on five DNA-based viruses whose genomes are sufficiently large to allow for application of our geometric framework. These viruses were evaluated alongside the other 50 species, enabling direct comparison. Results for the more informative substring-based analysis are presented in Table A5, while Table A6 reports the corresponding CV and QCV values.
Three of the viruses (Acanthamoeba polyphaga, Pandoravirus dulcis, and Pandoravirus salinus) display the lowest a d v and m d v values—comparable only to mouse chlamydia—and the substantially highest n l m values, whereas the remaining two viruses (Acanthamoeba castellanii and Megavirus chiliensis) do not exhibit such extreme behavior. Overall, the average m d v (34.62) and a d v (34.52) values across the five viruses are the lowest among all species groups, while their average n l m (45.1) is, by far, the highest.
We also observed that in this experimental round, the gorilla, again, attains the highest a d v (100), followed closely by the dolphin (99.9), whereas the dolphin achieves the highest m d v (100), followed by the gorilla (97.0).

7.6. Deviation of Biosequences Relative to Permutational Extremes

As mentioned previously, it is natural to ask how far a DNA sequence s departs from linearity relative to the greatest possible deviation among all permutations of its symbols (denoted m a x m d v ( s ) ). Given the patterns observed in the preceding sections, our goal is to determine whether DNA sequences could, through further evolution, develop even greater deviation from linearity or whether their current structures are already close to the theoretical maximum. Applying Theorem 1 to a four-letter alphabet (the case relevant for DNA) yields the following result.
Corollary 1.
For n = 4 , max m d v ( L ( s ) ) is the maximum of the terms:   
a 1 2 ( a 2 2 + a 3 2 + a 4 2 ) a 1 2 + a 2 2 + a 3 2 + a 4 2 , a 2 2 ( a 1 2 + a 3 2 + a 4 2 ) a 1 2 + a 2 2 + a 3 2 + a 4 2 , a 3 2 ( a 1 2 + a 2 2 + a 4 2 ) a 1 2 + a 2 2 + a 3 2 + a 4 2 , a 4 2 ( a 1 2 + a 2 2 + a 3 2 ) a 1 2 + a 2 2 + a 3 2 + a 4 2 ,
( a 1 2 + a 2 2 ) ( a 3 2 + a 4 2 ) a 1 2 + a 2 2 + a 3 2 + a 4 2 , ( a 1 2 + a 3 2 ) ( a 2 2 + a 4 2 ) a 1 2 + a 2 2 + a 3 2 + a 4 2 , and  ( a 1 2 + a 4 2 ) ( a 2 2 + a 3 2 ) a 1 2 + a 2 2 + a 3 2 + a 4 2 .
If, for instance, the first of these seven quantities is the largest for the given values of a i , then the strings whose discrete monotone paths pass through either ( a 1 , 0 , 0 , 0 ) or ( 0 , a 2 , a 3 , a 4 ) achieve m d v ( L ( s ) ) . Concretely, this includes all strings beginning with x 1 a 1 and all strings ending with x 1 a 1 . Using this result, we computed the ratio of m a x m d v / m d v for DNA sequences from human, gorilla, mouse, zebrafish, and tuberculosis and for a random sequence, considering substrings of lengths 1 , , 10 × 10 5 (see Table A7 in Appendix B). Interpreting deviation from linearity as an indicator of structural organization, the data suggest the following:
All organisms retain substantial potential for further structural complexity.
As expected, the ratios vary across organisms, with more primitive species exhibiting larger values than more complex ones. The smallest observed ratio of m a x m d v / m d v = 10.5 occurs for a 600,000-base substring of the human genome. In contrast, the largest ratio among biological sequences, i.e., 144.9, is found in a 1,000,000-base substring of the tuberculosis genome. For comparison, a random 1,000,000-base substring reaches an even higher value of 397.

8. Hypotheses and Biological Interpretations Derived from the Results

It is worth noting that progress in the natural sciences has often relied on a productive interplay between speculative theorizing and empirical validation. Many foundational ideas, such as general relativity, antimatter, dark matter, and gravitational waves, were initially proposed on theoretical grounds long before direct evidence became available. Biology, too, has seen influential theories rise and fall as new data have emerged, reflecting the evolving nature of empirical understanding. In this broader context, speculative frameworks can play a constructive role by articulating conceptual possibilities, identifying new questions, and guiding the search for empirical signatures. The geometric perspective developed here is offered in this spirit: as a theoretical lens that may help illuminate aspects of genomic organization and, potentially, the deeper transitions that shaped the emergence of life.
Based on the empirical patterns observed across the diverse set of genomes analyzed here, several broader conjectures naturally emerge. These hypotheses are not presented as definitive biological laws but, rather, as informed proposals suggested by the geometric regularities uncovered in our study. They are intended to guide future theoretical development and empirical testing and to highlight potential directions in which the geometric perspective on genomic organization may have explanatory value.

8.1. Equilibrium Condition for Life

Life emerged and evolved under conditions in which chemical reactions and mutations were not directed but largely stochastic. Beneficial changes were retained because they improved survival or replication, yet the underlying mutational input remained fundamentally random. From this perspective, the gradual accumulation of structural organization in early nucleic acids can be viewed as a form of increasing negentropy—a departure from randomness that enabled more complex and adaptive biological functions.
At the same time, this process could not proceed without constraint. If genomic sequences became too regular or too compressible, they would lose the capacity to encode large amounts of information and would become overly predictable, reducing their functional potential. Modern genomes are strikingly incompressible, and this near-maximal incompressibility appears to be a prerequisite for storing and maintaining the vast informational content required for cellular life. In information-theoretic terms, systems with low entropy are more predictable and carry less information, whereas systems with higher entropy can encode richer and more diverse messages. Therefore, genomes cannot drift toward excessive regularity without compromising their informational capacity. This suggests that the evolution of life required a balance: increasing structural organization and functional complexity, on one hand, and the preservation of high informational capacity through incompressibility, on the other. In this view, biological systems occupy a regime where randomness and structure coexist in a productive equilibrium—enough order to support function and adaptation yet enough unpredictability to encode large amounts of information efficiently and without redundancy.
Conjecture 1.
Life is possible only when a sensible equilibrium is maintained between negentropy (structural organization) and incompressibility (informational capacity).
This conceptual perspective aligns with broader ideas in the literature that describe life as a process shaped by thermodynamic constraints (see, e.g.,  [66]) and by the need to maintain functional organization, securing higher complexity and adaptivity. It is offered here not as a definitive explanation but as a hypothesis motivated by the geometric and statistical properties observed in genomic sequences.

8.2. Local Maxima as a Geometric Descriptor

Recall that local maxima were identified using MATLAB’s peakfinder function, which detects peaks that stand out relative to their local neighborhoods. In this characterization, a higher n l m corresponds to a more irregular or “jagged” cumulative path, whereas a lower n l m indicates a smoother trajectory with fewer noticeable deviations from the global drift line determined by the endpoints of the path. It is known that random (or closer to random) sequences are geometrically “more jagged” than nonrandom sequences that would have more predictable structures. With a reference to our experimental results, as expected, the normalized average of n l m on 1000 trials with random sequences equals 100 in both stages. In Stage 1, those for the primates, i.e., human, neanderthal, gorilla, and chimp, are 14.0, 14.8, 16.5, and 15.8, respectively. These values are substantially lower than all of the other species that we have tested (the average n l m for all 25 examined species being 21.9). See Table A1 in Appendix B. Stage 2 (Table A3) shows a similar trend. The four primates have n l m values of 13.6, 14.0, 17.5, and 15.0 (mean of 15.0), which are generally lower than those of the remaining species (overall mean of 24.4). Dolphin and whale exhibit similarly low values (16.0 and 15.3) comparable to those of the primates.
The above indicates that in the cumulative-count geometry, the primate genomes possess smoother large-scale compositional trajectories than the genomes of the other species that were analyzed. This pattern suggests the following:
Conjecture 2.
Primate genomes possess a comparatively more homogeneous or less fragmented compositional structure.
Various factors may contribute to this difference, such as the inclusion of lineage-specific mutational biases, constraints on the variation of GC content, the distribution of repetitive elements, or differences in genome architecture. One possible hypothesis is that the smoother monotone paths reflect stronger or more consistent evolutionary constraints acting on primate genomes. However, further analysis is needed in order to distinguish the relative contributions of mutational processes and natural selection. Presently, this observation could be interpreted as a sturdy geometric signature rather than as straight evidence for a specific evolutionary mechanism.
Taken together, these hypotheses highlight a unifying theme (extending the idea introduced in the previous section): genomic sequences appear to be shaped by a balance between structural organization and informational capacity, a balance that manifests geometrically in the cumulative-count representation. The smoothness of cumulative-count paths, the bounded range of deviation from linearity, and the apparent equilibrium between negentropy and incompressibility all suggest that genomes are neither maximally random nor maximally regular but, instead, reside in an intermediate regime that supports both functional complexity and high information density. All these point toward a constrained region of genomic configuration space that living systems occupy. These constraints do not arise from the geometric framework itself; rather, the framework makes them visible by translating sequence composition into measurable geometric traits.
Therefore, the conjectures proposed here serve two complementary purposes. First, they provide a conceptual interpretation of the empirical signatures uncovered by our geometric descriptors. Second, they outline a set of testable predictions about how genomes may be organized across species and through evolutionary time. Future work may refine these ideas by incorporating phylogenetic modeling, examining additional clades, or extending the geometric analysis to other biological polymers. For now, the hypotheses underscore the potential of geometric approaches to reveal broad organizing principles of genomic architecture.

8.3. Borders of Life

The results of Section 7.6 can be interpreted in another way. Consider the value of 100 × m d v / max m d v (the percent of the reciprocal of max m d v / m d v ), which represents the percentage that the deviation of a species constitutes with respect to the possible maximum range of the species’ deviation. The maximum and minimum percentages are reached for the same species: 0.69% for tuberculosis and 9.5% for human, while for a random substring, the average ratio equals 0.32%. These values are obtained for sufficiently long substrings (6,000,000, resp. 1,000,000 nucleotide bases for the two species and 1,000,000 for the random sequence). This provides evidence that the degree of deviation of species is within certain bounds: a lower bound that is a fraction smaller than 1 (lower-bounded by the percentage for the random substring) and an upper bound (over all species) that seems to be not considerably greater than 10%. More precise bounds could be established through extensive experimentation, but the initial results indicate that these limits are unlikely to shift dramatically.
This raises a natural question: Can deviation from linearity increase indefinitely as evolution proceeds, or are there intrinsic limits? Our interpretation is that genomic deviation from linearity is constrained by two opposing requirements. A lower bound is imposed by the level of structural organization (negentropy) necessary to sustain biological complexity and functionality. A genome that is too close to randomness would lack the long-range structure required for coordinated regulation and organismal viability. Conversely, an upper bound is imposed by the need to maintain high incompressibility. As discussed earlier, excessive regularity would reduce informational capacity, increase predictability, and ultimately compromise the ability of the genome to encode diverse and finely tuned biological functions.
Conjecture 3.
These two constraints define a region of permissible genomic organization— “borders of life” measured in terms of deviation from linearity. Exceeding these borders in either direction would lead to gradual degradation and, eventually, to loss of viability.
This perspective is compatible with broader ideas in the literature that view life as a system maintained by thermodynamic flows and constrained by the need to balance order and unpredictability. In this sense, the deviation-from-linearity trait may reflect a deeper equilibrium between structural organization and informational capacity, a balance that appears necessary for the persistence of living systems.

9. Speculative Outlook: Geometric Irregularity, Transitional States, and the Emergence of Biological Complexity

A broader speculative perspective suggested by our findings is that the origins and subsequent evolution of life may be viewed through the geometry of self-replicating macromolecules. Early replicators in most origins-of-life scenarios are assumed to have been short, structurally simple, and close to linear chains, whereas increasing structural irregularity is associated with the emergence of catalytic pockets, binding specificity, and functional diversification. From this viewpoint, geometric deviation from linearity can be interpreted as a proxy for informational capacity and biochemical versatility, reflecting the transition from simple templated replication to molecules capable of folding, regulation, and interaction. Although highly conjectural, this geometric framing is compatible with established theoretical models in which viable sequences occupy a constrained, structured subset of sequence space and where evolutionary innovation corresponds to the exploration of increasingly irregular regions of that space. The idea does not claim direct empirical support but offers a unifying conceptual lens through which the geometric descriptors introduced here may be connected to deep questions about the emergence and expansion of biological complexity.
The geometric descriptors developed in this work naturally invite broader reflection on the transition from inanimate organic matter to the earliest living systems. While the emergence of new organisms today follows well-understood biological processes, the origin of the first autonomous replicators remains far less clear. Contemporary definitions of life typically require cellular organization, metabolism, replication, responsiveness, and adaptive capacity, yet is not obvious that these criteria must appear simultaneously or that the boundary between non-living and living matter is sharply defined.
One possibility is that early prebiotic systems passed through intermediate states—entities that possessed some but not all of the features associated with life. Such “protoliving” structures may have exhibited primitive forms of metabolism, compartmentalization, or information storage without yet achieving full biological autonomy. In this view, the transition from inanimate to animate matter would not be a discrete event but a gradual process unfolding over extended periods of chemical evolution.
A complementary perspective is that the emergence of life may be associated with surpassing a threshold of structural organization or negentropy. Within the geometric framework developed here, such organization could be reflected in quantitative measures such as deviation from linearity or average deviation. Although it is impossible to determine the precise numerical value that might correspond to such a threshold, the conceptual point is that a minimal degree of informational structure may be required before a system can sustain the functions typically associated with living matter.
These two perspectives are not mutually exclusive. A protoliving entity may have gradually accumulated structural organization until it crossed a threshold compatible with autonomous replication and adaptive evolution. While these ideas extend beyond the empirical scope of the present study, they illustrate how geometric and information-theoretic descriptors might contribute to future theoretical discussions about the origin of life. They are offered here as speculative remarks intended to stimulate further thought rather than to propose definitive mechanisms.

10. Concluding Remarks

This work has developed a geometric framework for analyzing biological sequences and has demonstrated that large-scale geometric descriptors—such as deviation from linearity, local irregularity, and path-based negentropy—capture fundamental structural properties of genomes across the Tree of Life. The presented results obtained here indicate that these measures are not merely mathematical curiosities but reflect genuine biological organization. The consistent patterns observed across diverse species suggest that geometric constraints may play a central role in shaping viable sequence space and in distinguishing biological sequences from random polymers.
The evidence accumulated in this study points toward a coherent picture: genomic sequences occupy a highly structured region of sequence space, and their geometric properties correlate with broad evolutionary and organizational features. The monotone path representation and the monotone Graph of Life, together, provide a unified mathematical substrate for comparing species, quantifying complexity, and exploring the boundaries of biological organization. These constructions offer a promising alternative to traditional combinatorial approaches and open the possibility of a genuinely geometric theory of genomic structure.
Beyond the specific results obtained here, the work introduces a broadly applicable geometric technique for genome analysis. By treating DNA sequences as monotone paths and monotone graphs and by defining associated quantitative descriptors, the framework provides a reusable methodological tool that can be applied to a wide range of comparative, evolutionary, and structural questions. This positions the approach not only as a conceptual contribution but also as a technique that other researchers can adopt, extend, and integrate into their own analyses.
Several directions for future research follow naturally from the present work. A priority is the development of rigorous metrics on monotone paths and monotone graphs, which would enable quantitative comparison of species and support clustering, phylogenetic reconstruction, and the study of evolutionary trajectories within the geometric framework. Equally important is the systematic application of these descriptors to large genomic datasets including hundreds or thousands of species to determine the extent to which the observed trends generalize and to identify lineage-specific signatures. A deeper biological interpretation of local maxima and minima of geometric deviation (potentially linked to functional domains, regulatory regions, or evolutionary hotspots) represents another promising avenue.
The conceptual implications are equally significant. The results suggest that increasing geometric irregularity may accompany—and perhaps enable—the transition from simple replicators to complex organisms. This raises the possibility that the emergence of life and the evolution of complexity can be viewed as trajectories within a geometric landscape constrained by the structure of viable sequence space. Such a perspective offers a unifying lens through which questions about the origin of life, evolutionary innovation, and the theoretical limits of biological complexity may be reconsidered.
The present work is primarily descriptive and methodological: it introduces and characterizes geometric measures of genomic organization across a broad phylogenetic range. Analyses relating these geometric descriptors to specific biological covariates, such as base composition, coding density, or effective population size, would address a different level of explanation and would require additional data, careful model specification, and explicit control for phylogenetic relatedness. Such investigations are scientifically interesting and may help illuminate potential evolutionary drivers of the observed geometric patterns, but they constitute a substantial extension of the current framework. We therefore view them as promising directions for future research rather than components of the present study, whose aim is to establish and explore the geometric properties themselves.
This study is intended as a foundation for a broader research program. The mathematical structures introduced here, together with the empirical patterns they reveal, provide a basis for a more comprehensive interdisciplinary effort involving mathematics, computational biology, molecular evolution, and origins-of-life research. The framework is flexible, extensible, and capable of supporting both theoretical development and large-scale empirical analysis. We believe that the geometric viewpoint advanced in this work has the potential to reshape how genomic organization is quantified and understood and to contribute meaningfully to the long-standing effort to uncover the principles that govern the emergence and evolution of living systems.

Author Contributions

Conceptualization, V.E.B.; methodology, V.E.B.; software, R.P.B.; validation, V.E.B.; formal analysis, R.P.B. and V.E.B.; investigation, V.E.B.; data curation, R.P.B.; writing—original draft preparation, V.E.B.; writing—review and editing, V.E.B.; visualization, R.P.B. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Data Availability Statement

The original data presented in the study are openly available in Ensembl Genomes available at https://ensemblgenomes.org/.

Acknowledgments

Data were downloaded from online public repositories such as Ensembl and GenBank. The authors used MathLab R2025b software to run experiments. Excel was used for visualization of data, and Copilot was used for text editing of some parts of the paper. The authors thank the four anonymous reviewers for their valuable comments and suggestions, which substantially improved both the content and the presentation of the paper.

Conflicts of Interest

The authors declare no conflicts of interest.

Appendix A

Appendix A.1. Proof of Theorem 1

Lemma A1.
Let P be a convex polytope of dimension n and S be a closed convex proper subset of P. Then, some vertex v * of P satisfies d ( v * , S ) = max p P d ( p , S ) .
Proof. 
Let a be a point of P that satisfies d ( a , S ) = max p P d ( p , S ) , and suppose that a cannot be chosen to be a vertex. Let x be the (unique) point of S for which d ( a , S ) = d ( a , x ) = | | a x | | , so | | a x | | | | p x | | for every p P and | | a x | | > | | v x | | for every vertex v of P.
Clearly, a must belong to the boundary of P, as otherwise, the segment a x ¯ could be extended until it intersects the boundary of P. Let F be the face of P that contains a, with F having a dimension between 1 and n 1 . Let v 1 , , v k be the vertices of F. Since F is a convex polytope, any point in F can be expressed as a convex combination of its vertices; thus, we can write a = λ 1 v 1 + + λ k v k for some λ 1 , , λ k with 0 λ i < 1 , i = 1 k λ i = 1 (where λ i < 1 because we assume a is not a vertex). Also, by assumption, | | x v i | | < | | x a | | for all 1 i k . Thus, we have
| | x a | | = | | x ( λ 1 v 1 + + λ k v k ) | | = | | ( λ 1 + + λ k ) x ( λ 1 v 1 + + λ k v k ) | | = | | λ 1 ( x v 1 ) + + λ k ( x v k ) | | λ 1 | | x v 1 | | + + λ k | | x v k | | < λ 1 | | x a | | + + λ k | | x a | | = ( λ 1 + + λ k ) | | x a | | = | | x a | | .
This is a contradiction. □
Corollary A1.
Let a = ( a 1 , a 2 , , a n ) R + n and consider the parallelepiped K with a main diagonal of the segment σ = o a ¯ , where o is the origin of the coordinate system, and edges parallel to the coordinate axes. Then, max x σ , y K d ( x , y ) is reached at some vertex of K.
Now, let s be a string of length m over the alphabet X = { x 1 , , x n } , and let L ( s ) be the corresponding monotone path. Let letter x i appear a i times in s for 1 i n . Taking the σ segment connecting the first and last points of L ( s ) —the origin and the vector a = ( a 1 , , a n ) —we can consider the associated parallelepiped K with a main diagonal σ and axes parallel to the coordinate axes. We have the following plain fact.
Fact A1.
For every vertex v of K, there is a permutation of the elements of s such that the corresponding monotone path contains v.
With this preparation, we can finalize the proof of Theorem 1.
Proof of Theorem 1.
With reference to Corollary 1 and Fact A1 and the related notations, max m d v ( s ) is attained at some of the vertices of the parallelepiped K. These are in the form of ( a ¯ 1 , a ¯ 2 , , a ¯ n ) , where a ¯ i equals either a i or 0. Since max m d v ( s ) is clearly not attained for points o and a, we may exclude them from consideration. Thus, a particular vertex v of K different from o and a corresponds to a partition of the set A = { 1 , 2 , , n } into two nonempty sets N and Z, the former containing the indexes of the nonzero components of v and the latter containing the indexes of its zero components.
Keeping in mind Formula (A1), the distance from v to σ = o a ¯ satisfies
d ( v , σ ) = i N ( k a i a i ) 2 + i Z ( k a i ) 2 , where k = i N a i 2 i A a i 2 .
We then obtain
d ( v , σ ) = i N a i i N a i 2 i A a i 2 i A a i 2 i A a i 2 2 + i Z a i 2 i N a i 2 i A a i 2 2 = i N a i 2 i Z a i 2 i A a i 2 2 + i Z a i 2 ( i N a i 2 ) 2 ( i A a i 2 ) 2 = i N a i 2 ( i Z a i 2 ) 2 ( i A a i 2 ) 2 + i Z a i 2 ( i N a i 2 ) 2 ( i A a i 2 ) 2 = i N a i 2 i Z a i 2 ( i Z a i 2 + i N a i 2 ) ( i A a i 2 ) 2 = i N a i 2 i Z a i 2 i A a i 2 ( i A a i 2 ) 2 = i N a i 2 i Z a i 2 i A a i 2 .
The above considerations also imply the validity of the second part of the theorem. □

Appendix A.2. Proof of Theorem 2

Proof. 
Let l be the straight line through the first and last points of path L, which is the axis of an enclosing cylinder for L.
Let p = ( p 1 , , p n ) be a point of L that maximizes the distance to l, and let r = d ( p , l ) .
Let C * be an enclosing cylinder for L of minimal radius r * centered about an axis l * . We have r * r 2 r * , as the second inequality holds from the following argument. Let p l be the foot of the perpendicular line from p to l, i.e., d ( p , p ) = r . According to the construction of path L and plain geometric arguments, a minimal enclosing cylinder C ¯ for the three points q 0 , q m , and p has an axis l ¯ that is parallel to line l and passes through the midpoint of segment p p ¯ ; obviously, the radius r ¯ of C ¯ satisfies r ¯ = r / 2 . Since { q 0 , q m , p } L , we have r ¯ r * and, thus, r 2 r * .
The estimate of the time necessary to compute r follows from related calculus formulas. Let l have a parametric equation x = t a , where x , a R n , a = ( a 1 , , a n ) , and let point p be as defined above. Point p belongs to l, so p = t a for some t R . Vector p p is orthogonal to vector a, i.e., their scalar product satisfies a · ( p p ) = 0 . From the last equality, one easily obtains
t = a 1 p 1 + + a n p n a 1 2 + + a n 2 and d ( p , l ) = ( a 1 t p 1 ) 2 + + ( a n t p n ) 2 .
Obviously, t and r = d ( p , l ) can be found with O ( n ) arithmetic operations and a single square-root computation; the latter is not necessary to perform when comparing distances from points of L to line l. □

Appendix A.3. Proof of Property 4

Proof. 
We use induction on the number of monotone paths that compose a monotone graph. Obviously, the statement holds for m = 1 , i.e., if G consists of a single monotone path. Let the statement be true for any monotone graph composed of m 1 distinct monotone paths, each with a length of at most r, where r is a positive integer. Such a graph has no more than n = m r vertices and, owing to its sparsity, no more than c · n edges for some absolute positive constant c (i.e., G has O ( n ) edges).
Now, let G be a monotone graph made of m + 1 distinct monotone paths ( m 1 ) with lengths of at most r, where r is a positive integer. Let P be the shortest among these paths. Let G = G P be the graph obtained by removing P from G. G is monotone formed by monotone paths and has no more than n = m r vertices. Moreover, G is sparse according to the induction hypothesis, i.e., G has O ( n ) edges.
Now, consider graph G and path P within it. Parts of path P may be parts of other paths of G. Let S 1 , S 2 , , S k be other portions of P that are not parts of other paths, and let s 1 , s 2 , , s k be the number of vertices in these portions. Since all paths composing G are distinct, k 1 . Also, clearly, k n and s 1 + s 2 + + s k n . Moreover, the number t i of edges in S i ( 1 i k ) satisfies t i = s i + 1 . Then, for the total number of edges in S 1 , S 2 , , S k , we have t 1 + t 2 + + t s s 1 + s 2 + + s k + k 2 n = O ( n ) . (Note that some pairs of sets S i may have a common vertex, so the above inequality may be strict.) Recalling that G is sparse and has O ( n ) edges, we find that the total number of edges in G is of the order of O ( n ) + O ( n ) = O ( n ) ; hence, G is sparse. □

Appendix B

Table A1. Normalized averages of 1000 trials.
Table A1. Normalized averages of 1000 trials.
Scientific NameCommon NameSubstringsSubsequences
advmdvnlmadvmdvnlm
Random SequenceRandom12.812.8100.09.48.9100.0
H. SapiensHuman88.989.914.084.287.87.9
H. NeanderthalensisNeanderthal85.585.114.8100.0100.07.7
G. GorillaGorilla100.0100.016.561.869.44.6
P. TroglodytesChimp81.281.015.840.539.311.8
C. FamiliarisDog71.874.318.737.233.817.4
G. GallusChicken59.157.222.592.873.54.5
C. JacchusMarmoset55.055.824.714.313.756.3
R. NorvegicusRat68.268.022.931.630.513.9
M. MusculusMouse62.564.919.815.514.949.7
O. AnatinusPlatypus41.241.727.145.435.05.4
A. CarolinensisLizard31.632.339.811.711.067.1
D. RerioZebrafish59.658.525.632.930.111.1
O. LatipesMedaka fish46.647.534.731.929.310.9
D. MelanogasterFruit fly49.149.331.117.116.136.3
C. IntestinalisSea squirt32.832.232.011.210.674.7
C. ElegansNematode51.652.422.718.519.826.3
S. CerevisiaeYeast39.439.827.412.711.774.1
C. MuridarumChlamydia29.729.327.972.364.94.5
M. TuberculosisTuberculosis33.834.023.29.79.396.2
P. GingivalisGingivalis50.048.118.518.716.931.9
S. ThermophilusStreptococcus37.236.522.647.843.14.5
O. SativaRice73.876.924.514.914.063.5
Z. MaysCorn76.180.318.918.518.344.5
A. ThalianaCress44.844.427.811.210.878.8
G. MaxSoybean52.253.022.723.018.725.4
Maximum pre-normalized value:554.51100.821.9739.51563.622.4
Table A2. Coefficient of variation (CV) and quartile coefficient of variation (QCV) of the 1000 trials.
Table A2. Coefficient of variation (CV) and quartile coefficient of variation (QCV) of the 1000 trials.
Common NameSubstringsSubsequences
CVQCVCVQCV
advmdvnlmadvmdvnlmadvmdvnlmadvmdvnlm
Random0.240.210.450.160.140.270.250.220.490.170.150.30
Human0.430.420.710.290.280.330.050.030.250.030.020.33
Neanderthal0.430.420.720.260.270.330.040.030.280.030.020.33
Gorilla0.520.480.860.350.340.670.060.050.170.040.030.00
Chimp0.430.430.780.270.280.330.100.060.270.060.040.20
Dog0.430.410.600.280.260.500.070.080.200.050.060.14
Chicken0.480.460.970.370.360.500.050.040.030.030.030.00
Marmoset0.440.420.610.260.240.400.200.190.460.140.140.28
Rat0.590.610.760.410.390.560.130.110.290.080.070.00
Mouse0.330.360.570.230.260.250.220.200.490.150.130.33
Platypus0.420.380.580.240.230.450.090.100.380.060.070.00
Lizard0.330.320.510.230.220.380.190.180.440.130.120.29
Zebrafish0.410.360.570.280.260.400.120.100.480.080.070.20
Medaka fish0.420.370.540.260.250.430.120.110.380.090.080.20
Fruit fly0.480.460.690.350.320.500.130.140.330.090.090.25
Sea squirt0.340.310.590.210.200.380.250.230.470.160.160.31
Nematode0.360.350.550.260.210.330.190.170.460.130.120.33
Yeast0.330.300.400.180.200.270.250.220.460.170.150.27
Chlamydia0.340.280.490.230.200.330.050.050.000.030.030.00
Tuberculosis0.420.410.710.250.260.400.230.200.400.150.130.27
Gingivalis0.460.450.580.300.280.430.190.180.480.130.120.29
Streptococcus0.370.340.540.220.220.400.080.070.030.060.050.00
Rice0.470.450.540.300.290.400.210.190.320.140.130.21
Corn0.450.450.560.310.350.500.160.160.300.100.100.16
Cress0.290.290.490.200.180.270.220.210.420.150.140.27
Soybean0.340.330.480.190.210.330.140.130.410.100.090.27
Table A3. Substring and subsequence statistics across species.
Table A3. Substring and subsequence statistics across species.
GroupScientific NameCommon NameSubstringsSubsequences
advmdvnlmadvmdvnlm
Random Sequence12.111.8100.09.58.9100.0
PrimatesH. sapienshuman85.982.613.684.287.87.6
H. neanderthalensisNeanderthal84.882.514.0100.0100.07.5
G. gorillagorilla93.891.017.561.869.14.5
P. troglodyteschimpanzee79.076.715.040.539.411.5
Average 85.983.215.071.674.17.8
MammalsT. truncatusbottlenose dolphin100.0100.016.027.740.36.6
B. musculusblue whale88.285.615.325.324.526.0
C. familiarisdog66.666.919.037.533.917.1
R. norvegicusbrown rat65.163.021.731.730.613.7
M. musculushouse mouse61.062.618.515.614.948.9
C. jacchuscommon marmoset52.652.424.014.313.855.0
L. africanaAfrican elephant41.240.626.325.925.215.4
O. anatinusplatypus40.339.926.945.534.95.4
Average 64.463.921.027.927.323.5
Other vertebratesG. galluschicken56.852.921.492.673.34.4
C. moneduloidesNew Caledonian crow60.356.322.838.032.822.8
P. majorgreat tit52.751.528.717.820.723.2
T. guttatazebra finch50.748.929.325.323.917.3
M. aterbrown-headed cowbird45.144.230.123.921.118.4
A. carolinensisgreen anole30.330.337.711.610.966.9
G. morhuaAtlantic cod68.766.118.713.012.267.5
D. reriozebrafish55.652.625.232.429.711.4
O. latipesmedaka43.944.033.331.629.110.6
Average 51.649.627.531.828.226.9
InvertebratesA. californicaCalifornia sea hare52.950.624.115.514.248.6
C. elegansnematode49.549.022.418.419.825.7
D. pulexwater flea48.846.723.015.115.043.6
B. morisilkworm44.747.028.219.216.334.4
D. melanogasterfruit fly45.945.430.117.116.036.0
T. adhaerensplacozoan33.432.524.69.49.298.3
C. intestinalissea squirt31.229.731.011.110.771.9
Average 43.843.026.215.114.551.2
PlantsZ. mayscorn/maize71.173.219.118.618.443.0
O. sativarice69.369.924.114.814.062.6
P. trichocarpablack cottonwood52.952.220.513.112.265.3
G. maxsoybean50.649.922.123.018.724.7
B. distachyonpurple false brome50.649.828.912.011.186.5
A. thalianathale cress43.642.125.911.411.076.8
Average 56.456.223.415.514.259.8
FungiN. crassared bread mold62.361.333.318.519.034.8
S. cerevisiaeyeast37.236.526.712.611.673.1
S. pombefission yeast33.234.230.310.39.886.1
Average 44.244.030.113.813.564.7
ProtistsD. discoideumslime mold59.760.629.015.513.153.7
C. reinhardtiigreen alga52.751.224.113.113.652.3
L. majorleishmania parasite52.048.125.515.814.838.5
G. lambliagiardia45.744.416.813.012.862.8
Average 52.551.123.914.413.651.8
BacteriaP. gingivalisoral bacterium48.245.217.318.716.831.9
S. coelicolorsoil actinomycete40.739.819.219.617.824.3
D. radioduransradiation-resistant bacterium36.835.921.215.213.143.8
P. fluorescensfluorescent pseudomonad35.434.822.325.723.88.1
V. choleraecholera bacterium36.034.022.623.223.311.2
S. thermophilusthermophilic streptococcus34.433.023.447.743.04.4
N. punctiformecyanobacterium34.033.126.510.19.691.9
M. tuberculosistuberculosis bacterium33.632.622.09.89.394.5
C. muridarummouse chlamydia28.627.626.972.064.64.4
Average 36.435.122.426.924.634.9
Table A4. Coefficient of variation (CV) and quartile coefficient of variation (QCV) for substrings and subsequences.
Table A4. Coefficient of variation (CV) and quartile coefficient of variation (QCV) for substrings and subsequences.
GroupScientific NameCommon NameSubstringSubsequence
CVQCVCVQCV
advmdvnlmadvmdvnlmadvmdvnlmadvmdvnlm
Random Sequence0.260.230.480.180.160.330.250.220.470.170.150.32
PrimatesH. sapienshuman0.430.420.690.280.280.330.050.030.260.030.020.33
H. neanderthalensisNeanderthal0.420.410.750.270.270.330.040.030.270.030.020.33
G. gorillagorilla0.530.500.850.350.360.430.060.050.170.040.030.00
P. troglodyteschimpanzee0.410.420.790.270.270.330.090.050.250.060.040.20
MammalsT. truncatusbottlenose dolphin0.880.900.740.380.290.330.130.080.530.090.050.33
B. musculusblue whale0.540.510.620.380.380.430.120.090.230.090.060.17
C. familiarisdog0.420.410.570.260.250.500.080.080.200.050.060.14
R. norvegicusbrown rat0.540.560.750.370.360.500.120.110.300.080.070.33
M. musculushouse mouse0.350.390.640.240.280.430.200.190.470.140.130.33
C. jacchuscommon marmoset0.420.400.570.220.220.400.210.190.420.140.130.28
L. africanaAfrican elephant0.340.340.550.230.220.270.140.130.330.100.090.14
O. anatinusplatypus0.430.410.590.260.240.450.100.090.390.060.060.00
Other vertebratesG. galluschicken0.480.460.970.350.360.500.040.040.000.030.030.00
C. moneduloidesNew Caledonian crow0.650.600.860.430.410.560.090.080.140.060.050.09
P. majorgreat tit0.640.670.730.400.370.500.120.130.330.070.090.20
T. guttatazebra finch0.650.640.650.340.350.500.120.120.370.080.090.25
M. aterbrown-headed cowbird0.470.460.680.310.310.500.090.100.310.060.070.25
A. carolinensisgreen anole0.350.330.530.240.210.380.190.180.420.130.120.24
G. morhuaAtlantic cod0.660.600.660.240.250.430.190.160.350.130.110.24
D. reriozebrafish0.410.360.540.280.240.400.120.110.480.090.070.20
O. latipesmedaka0.410.380.540.260.260.430.130.110.390.090.070.20
InvertebratesA. californicaCalifornia sea hare0.400.380.670.270.270.430.230.220.450.170.150.27
C. elegansnematode0.360.340.550.240.210.330.180.150.440.130.100.27
D. pulexwater flea0.420.410.770.290.280.330.200.200.390.140.130.26
B. morisilkworm0.490.460.590.290.290.380.180.150.410.120.100.25
D. melanogasterfruit fly0.490.470.680.360.330.500.130.140.330.090.090.25
T. adhaerensplacozoan0.330.320.540.220.210.400.250.230.460.180.150.29
C. intestinalissea squirt0.320.300.570.200.180.380.240.220.440.170.150.29
PlantsZ. mayscorn/maize0.440.440.540.330.340.330.150.160.310.100.110.16
O. sativarice0.470.440.560.300.280.400.200.190.320.130.140.21
P. trichocarpablack cottonwood0.300.280.490.210.200.330.200.180.390.130.120.24
G. maxsoybean0.370.340.480.190.180.330.150.130.380.100.080.27
B. distachyonpurple false brome0.270.250.440.170.170.330.250.220.470.170.150.30
A. thalianathale cress0.280.270.450.170.170.270.230.210.420.150.140.29
FungiN. crassared bread mold1.041.040.610.300.290.430.150.150.220.110.110.13
S. cerevisiaeyeast0.320.290.400.180.190.270.250.220.460.170.150.29
S. pombefission yeast0.360.350.450.200.210.290.260.240.470.180.160.30
ProtistsD. discoideumslime mold0.380.350.560.260.240.330.220.190.330.150.110.22
C. reinhardtiigreen alga0.360.340.510.210.200.400.210.180.380.140.120.22
L. majorleishmania parasite0.540.490.830.420.380.600.180.150.320.120.110.18
G. lambliagiardia0.400.410.570.270.280.430.240.220.440.170.160.29
BacteriaP. gingivalisoral bacterium0.470.450.570.280.260.430.180.170.470.120.110.29
S. coelicolorsoil actinomycete0.460.480.590.270.260.500.210.190.590.140.130.40
D. radioduransradiation-resistant bacterium0.460.460.550.220.220.330.210.170.400.140.110.26
P. fluorescensfluorescent pseudomonad0.500.480.570.270.280.400.170.150.690.110.100.33
V. choleraecholera bacterium0.540.490.600.270.250.400.150.140.370.100.100.20
S. thermophilusthermophilic streptococcus0.360.330.560.220.210.400.080.070.000.050.050.00
N. punctiformecyanobacterium0.330.310.490.220.190.270.230.210.440.160.140.30
M. tuberculosistuberculosis bacterium0.420.420.720.250.270.500.230.210.420.150.130.27
C. muridarummouse chlamydia0.320.280.470.220.200.330.050.050.000.040.040.00
Table A5. Substring statistics including a group of DNA-based viruses.
Table A5. Substring statistics including a group of DNA-based viruses.
GroupScientific NameCommon Nameadvmdvnlm
Random Sequence12.912.6100.0
PrimatesH. sapienshuman91.088.513.9
H. neanderthalensisNeanderthal89.186.813.8
G. gorillagorilla100.097.016.7
P. troglodyteschimpanzee84.081.115.9
Average 91.088.415.1
MammalsT. truncatusbottlenose dolphin99.9100.016.4
B. musculusblue whale95.693.415.3
C. familiarisdog70.470.319.6
R. norvegicusbrown rat67.966.722.5
M. musculushouse mouse64.465.719.3
C. jacchuscommon marmoset54.854.825.1
L. africanaAfrican elephant43.143.026.5
O. anatinusplatypus42.942.326.7
Average 67.467.021.4
Other vertebratesG. galluschicken61.057.221.1
C. moneduloidesNew Caledonian crow63.759.123.1
P. majorgreat tit56.155.429.4
T. guttatazebra finch55.353.329.1
M. aterbrown-headed cowbird48.046.829.7
A. carolinensisgreen anole31.831.938.9
G. morhuaAtlantic cod73.070.918.4
D. reriozebrafish60.057.324.8
O. latipesmedaka46.946.933.8
Average 55.153.227.6
InvertebratesA. californicaCalifornia sea hare56.954.524.0
C. elegansnematode52.952.122.3
D. pulexwater flea50.348.424.1
B. morisilkworm49.151.628.6
D. melanogasterfruit fly49.248.430.9
T. adhaerensplacozoan35.034.324.7
C. intestinalissea squirt33.531.931.4
Average 46.745.926.6
PlantsZ. mayscorn/maize77.180.518.8
O. sativarice74.475.624.6
P. trichocarpablack cottonwood56.355.120.9
G. maxsoybean53.553.122.5
B. distachyonpurple false brome54.153.529.2
A. thalianathale cress46.445.226.9
Average 60.360.523.8
FungiN. crassared bread mold68.167.032.8
S. cerevisiaeyeast39.338.727.3
S. pombefission yeast34.436.131.9
Average 47.347.330.7
ProtistsD. discoideumslime mold64.665.328.6
C. reinhardtiigreen alga59.157.423.4
L. majorleishmania parasite53.450.526.5
G. lambliagiardia49.148.217.3
Average 56.655.424.0
BacteriaP. gingivalisoral bacterium50.347.518.5
S. coelicolorsoil actinomycete43.742.818.6
D. radioduransradiation-resistant bacterium39.438.522.2
P. fluorescensfluorescent pseudomonad38.838.022.2
V. choleraecholera bacterium39.937.722.0
S. thermophilusthermophilic streptococcus38.236.622.5
N. punctiformecyanobacterium36.535.926.6
M. tuberculosistuberculosis bacterium34.433.623.3
C. muridarummouse chlamydia30.729.627.4
Average 39.137.822.6
DNA-VirusesAcanthamoeba castellanii 45.445.721.1
Acanthamoeba polyphaga 27.527.159.5
Megavirus chiliensis 42.443.724.4
Pandoravirus dulcis 29.929.860.6
Pandoravirus salinus 27.426.859.9
Average 34.534.645.1
Table A6. Coefficient of variation (CV) and quartile coefficient of variation (QCV) for substrings including DNA-based viruses.
Table A6. Coefficient of variation (CV) and quartile coefficient of variation (QCV) for substrings including DNA-based viruses.
GroupScientific NameCommon NameCVQCV
advmdvnlmadvmdvnlm
Random 0.250.220.460.170.150.29
PrimatesH. sapienshuman0.420.420.750.270.260.33
H. neanderthalensisNeanderthal0.410.410.700.260.270.33
G. gorillagorilla0.520.480.820.350.350.43
P. troglodyteschimpanzee0.440.440.760.290.280.33
MammalsT. truncatusbottlenose dolphin0.760.760.720.370.260.33
B. musculusblue whale0.510.480.650.360.360.33
C. familiarisdog0.420.420.580.260.250.50
R. norvegicusbrown rat0.540.570.790.370.380.56
M. musculushouse mouse0.340.370.590.250.270.43
C. jacchuscommon marmoset0.420.400.590.230.230.40
L. africanaAfrican elephant0.310.310.510.220.210.27
O. anatinusplatypus0.430.400.560.260.240.45
Other vertebratesG. galluschicken0.470.441.020.320.330.50
C. moneduloidesNew Caledonian crow0.640.590.870.410.380.56
P. majorgreat tit0.660.700.740.420.400.50
T. guttatazebra finch0.660.650.670.350.360.50
M. aterbrown-headed cowbird0.470.460.680.300.300.50
A. carolinensisgreen anole0.330.310.530.230.210.38
G. morhuaAtlantic cod0.630.580.620.230.230.43
D. reriozebrafish0.390.350.540.280.250.40
O. latipesmedaka0.410.370.540.270.270.43
InvertebratesA. californicaCalifornia sea hare0.380.360.660.260.250.40
C. elegansnematode0.350.330.530.240.190.33
D. pulexwater flea0.430.420.780.270.270.33
B. morisilkworm0.540.500.630.310.310.50
D. melanogasterfruit fly0.490.460.700.350.320.50
T. adhaerensplacozoan0.320.310.510.200.210.40
C. intestinalissea squirt0.340.310.550.200.200.38
PlantsZ. mayscorn/maize0.420.440.540.310.340.50
O. sativarice0.470.440.560.310.290.40
P. trichocarpablack cottonwood0.320.290.480.210.200.33
G. maxsoybean0.360.340.460.190.190.33
B. distachyonpurple false brome0.270.240.440.190.170.33
A. thalianathale cress0.290.280.480.200.180.33
FungiN. crassared bread mold1.041.040.590.300.280.43
S. cerevisiaeyeast0.300.280.390.170.180.27
S. pombefission yeast0.330.360.480.200.210.29
ProtistsD. discoideumslime mold0.370.340.520.250.230.33
C. reinhardtiigreen alga0.360.360.550.220.210.40
L. majorleishmania parasite0.550.500.830.400.360.60
G. lambliagiardia0.420.420.600.280.300.43
BacteriaP. gingivalisoral bacterium0.450.440.590.280.260.50
S. coelicolorsoil actinomycete0.450.450.600.260.270.43
D. radioduransradiation-resistant bacterium0.510.510.580.240.250.40
P. fluorescensfluorescent pseudomonad0.520.510.580.270.270.40
V. choleraecholera bacterium0.560.520.640.270.260.33
S. thermophilusthermophilic streptococcus0.390.360.560.230.210.33
N. punctiformecyanobacterium0.330.300.480.220.200.27
M. tuberculosistuberculosis bacterium0.410.410.680.250.270.40
C. muridarummouse chlamydia0.320.280.470.220.200.33
DNA-VirusesAcanthamoeba castellanii 0.510.530.480.250.260.33
Acanthamoeba polyphaga 0.400.370.440.210.200.31
Megavirus chiliensis 0.570.580.510.250.260.27
Pandoravirus dulcis 0.290.270.360.180.170.23
Pandoravirus salinus 0.370.330.430.220.190.31
Table A7. The max m d v / m d v ratio computed for substrings of length { 1 , , 10 } × 10 5 .
Table A7. The max m d v / m d v ratio computed for substrings of length { 1 , , 10 } × 10 5 .
Human17.315.012.613.215.610.513.111.414.312.2
Gorilla12.720.115.212.011.813.513.011.210.610.9
Mouse19.928.130.840.237.941.745.257.146.750.0
Zebrafish28.224.927.732.731.728.632.837.026.640.8
Tuberculosis44.664.866.584.688.196.5105.1116.3122.5144.9
Random Sequence126.7164.3205.7260.2280.7304.7325.1336.0351.2397.0

References

  1. Cairns-Smith, A.G. Seven Clues to the Origin of Life: A Scientific Detective Story; Cambridge University Press: Cambridge, UK, 1990. [Google Scholar]
  2. Darwin, C.M.A. On the Origin of Species by Means of Natural Selection, or the Preservation of Favoured Races in the Struggle for Life; D. Appleton & Co.: New York, NY, USA, 1859. [Google Scholar]
  3. Davies, P. The Fifth Miracle: The Search for the Origin and Meaning of Life; Simon & Schuster: New York, NY, USA, 2000. [Google Scholar]
  4. Deamer, D. First Life: Discovering the Connections Between Stars, Cells, and How Life Began; University of California Press: Berkeley, CA, USA, 2012. [Google Scholar]
  5. Dyson, F. Origins of Life; Cambridge University Press: Cambridge, UK, 1999. [Google Scholar]
  6. Fry, I. Emergence of Life on Earth: A Historical and Scientific Overview; Rutgers University Press: New Brunswick, NJ, USA, 2000. [Google Scholar]
  7. Hazen, R. Genesis: The Scientific Quest for Life’s Origin; National Academies Press: Washington, DC, USA, 2005. [Google Scholar]
  8. Lane, N. The Vital Question: Energy, Evolution, and the Origins of Complex Life; W. W. Norton and Company, Inc.: New York, NY, USA, 2015. [Google Scholar]
  9. Luisi, P.L. The Emergence of Life: From Chemical Origins to Synthetic Biology; Cambridge University Press: Cambridge, UK, 2010. [Google Scholar]
  10. Morowitz, H.J. Beginnings of Cellular Life: Metabolism Recapitulates Biogenesis; Yale University Press: New Haven, CT, USA, 2004. [Google Scholar]
  11. Pross, A. What Is Life?: How Chemistry Becomes Biology; Oxford University Press: Oxford, UK, 2012. [Google Scholar]
  12. Rasmussen, S.; Bedau, M.A.; Chen, L.; Deamer, D.; Krakauer, D.C.; Packard, N.H.; Stadler, P.F. (Eds.) Protocells: Bridging Nonliving and Living Matter; The MIT Press: Cambridge, MA, USA, 2008. [Google Scholar]
  13. Rauchfuss, H. Chemical Evolution and the Origin of Life; Springer: Berlin/Heidelberg, Germany, 2008. [Google Scholar]
  14. Russell, M. (Ed.) Abiogenesis: How Life Began. The Origins and Search for Life; Cosmology Science Publishers: Cambridge, UK, 2011. [Google Scholar]
  15. Schopf, J.W. (Ed.) Life’s Origin: The Beginnings of Biological Evolution; University of California Press: Berkeley, CA, USA, 2002. [Google Scholar]
  16. Schrodinger, E. What Is Life? Cambridge University Press: Cambridge, UK, 1944. [Google Scholar]
  17. Schulze-Makuch, D.; Irwin, L.N. Life in the Universe: Expectations and Constraints; Springer: Berlin/Heidelberg, Germany, 2008. [Google Scholar]
  18. Smith, E.; Morowitz, H. The Origin and Nature of Life on Earth: The Emergence of the Fourth Geosphere; Cambridge University Press: Cambridge, UK, 2016. [Google Scholar]
  19. Smith, J.M.; Szathmáry, E. The Origins of Life: From the Birth of Life to the Origin of Language; Oxford University Press: Oxford, UK, 2000. [Google Scholar]
  20. Sullivan, W.T., III; Baross, J.B. (Eds.) Planets and Life: The Emerging Science of Astrobiology; Cambridge University Press: Cambridge, UK, 2007. [Google Scholar]
  21. Wills, C.; Bada, J. The Spark of Life: Darwin and the Primeval Soup; Oxford University Press: New York, NY, USA, 2001. [Google Scholar]
  22. Zubay, G. Origins of Life: On Earth and in the Cosmos, 2nd ed.; Academic Press: San Diego, CA, USA, 2000. [Google Scholar]
  23. Brimkov, B. Geometric approach to string analysis for biosequence classification. J. Integr. Bioinform. 2014, 11, 72–87. [Google Scholar] [CrossRef]
  24. Brimkov, B.; Brimkov, V.E. Geometric approach to string analysis: Deviation from linearity and its use for biosequence classification. arXiv 2013, arXiv:1308.2885. [Google Scholar] [CrossRef]
  25. Brimkov, B.; Brimkov, V.E. Geometric Approach to Biosequence Analysis. In Proceedings of the 8th International Conference on Practical Applications of Computational Biology and Bioinformatics, Salamanca, Spain, 4–6 June 2014; pp. 97–104. [Google Scholar]
  26. Cormen, T.H.; Leiserson, C.E.; Rivest, R.L.; Stein, C. Introduction to Algorithms; MIT Press & McGraw Hill: Cambridge, MA, USA, 2001. [Google Scholar]
  27. Garey, M.; Johnson, D. Computers and Intractability; W.H. Freeman & Company: San Francisco, CA, USA, 1979. [Google Scholar]
  28. Apostolico, A.; Giancarlo, R. Sequence alignment in molecular biology. J. Comput. Biol. 1998, 5, 173–196. [Google Scholar] [CrossRef]
  29. Apostolico, A.; Cunial, F. Probing the randomness of proteins by their subsequence composition. In Proceedings of the Data Compression Conference DCC ’09, Washington, DC, USA, 16–18 March 2009; pp. 173–182. [Google Scholar]
  30. Apostolico, A.; Cunial, F. The subsequence composition of polypeptides. J. Comput. Biol. 2010, 17, 1011–1049. [Google Scholar] [CrossRef]
  31. Apostolico, A.; Galil, Z.Z. (Eds.) Pattern Matching Algorithms; Oxford University Press: New York, NY, USA, 1997. [Google Scholar]
  32. Crochemore, M.; Rytter, W. Text Algorithms; Oxford University Press: New York, NY, USA, 1994. [Google Scholar]
  33. Sankoff, D.; Kruskal, J.B. (Eds.) Time Warps, String Edits, and Macromolecules: The Theory and Practice of Sequence Computation; Addison-Wesley: Reading, MA, USA, 1983. [Google Scholar]
  34. Waterman, M.S. Introduction to Computational Biology. Maps, Sequences and Genomes; Chapman Hall: London, UK, 1995. [Google Scholar]
  35. Yoder, N. PeakFinder (Update 2011, File ID: #25500), Matlab Central. Available online: http://www.mathworks.com/matlabcentral/fileexchange/25500-peakfinder (accessed on 28 October 2025).
  36. Agarwal, P.; Aronov, B.; Sharir, M. Line transversals of balls and smallest enclosing cylinders in three dimensions. Discret. Comput. Geom. 1999, 21, 373–388. [Google Scholar] [CrossRef][Green Version]
  37. Chan, T. Approximating the diameter, width, smallest enclosing cylinder, and minimum-width annulus. In Proceedings of the 16th Annual Symposium on Computational Geometry, Clear Water Bay, Kowloon, Hong Kong, 12–14 June 2000; pp. 300–309. [Google Scholar]
  38. Devillers, O.; Preparata, F. Evaluating the cylindricity of a nominally cylindrical point set. In Proceedings of the SODA, San Francisco, CA, USA, 2000; SIAM Publisher: Philadelphia, PA, USA, 2000; pp. 518–527. [Google Scholar]
  39. Erdős, P.; Gyori, E.; Simonovits, M. How many edges should be deleted to make a triangle-free graph bipartite? In Sets, Graphs and Numbers, Budapest, Hungary, 1991; Colloquia Mathematica Societatis János Bolyai: Amsterdam, The Netherlands, 1992; Volume 60, pp. 239–263. [Google Scholar]
  40. Kell, D.B. Scientific discovery as a combinatorial optimization problem: How best to navigate the landscape of possible experiments? BioEssays 2012, 34, 163–251. [Google Scholar] [CrossRef]
  41. Barabasi, A.-L.; Oltvai, Z.N. Network Biology: Understanding the Cell’s Functional Organization. Nat. Rev. Genet. 2004, 5, 101–113. [Google Scholar] [CrossRef]
  42. Junker, B.H.; Schreiber, F. Analysis of Biological Networks; Wiley: Hoboken, NJ, USA, 2008. [Google Scholar]
  43. Newman, M. Networks; Oxford University Press: Oxford, UK, 2018. [Google Scholar]
  44. Proulx, S.R.; Promislow, D.E.L.; Phillips, P.C. Network Thinking in Ecology and Evolution; Princeton University Press: Princeton, NJ, USA, 2005. [Google Scholar]
  45. Albert, R.; Barabasi, A.-L. Statistical Mechanics of Complex Networks. Rev. Mod. Phys. 2002, 74, 47–97. [Google Scholar] [CrossRef]
  46. Jeong, H.; Mason, S.P.; Barabási, A.-L.; Oltvai, Z.N. Lethality and Centrality in Protein Networks. Nature 2001, 411, 41–42. [Google Scholar] [CrossRef]
  47. Palla, G.; Derenyi, I.; Farkas, I.; Vicsek, T. Uncovering the Overlapping Community Structure of Complex Networks in Nature and Society. Nature 2005, 435, 814–818. [Google Scholar] [CrossRef]
  48. Sharan, R.; Ulitsky, I.; Shamir, R. Network-Based Prediction of Protein Function. Mol. Syst. Biol. 2007, 3, MSB4100129. [Google Scholar] [CrossRef]
  49. Szachniuk, M. Graph-based and graph grammar approaches to modelling structure determination, development and evolution. In Handbook of Graph Grammars and Computing by Graph Transformation; Applications, Languages, and Tools; Ehrig, H., Engels, G., Kreowski, H.-J., Rozenberg, G., Eds.; World Scientific: Singapore, 2004; Volume 3, pp. 639–675. [Google Scholar]
  50. Brun, C.; Chevenet, F.; Martin, D.; Wojcik, J.; Guenoche, A.; Jacq, B. Functional classification of proteins for the prediction of cellular function from a protein–protein interaction network. Genome Res. 2003, 13, 1373–1380. [Google Scholar] [CrossRef]
  51. Gavin, A.-C.; Bösche, M.; Krause, R.; Grandi, P.; Marzioch, M.; Bauer, A.; Schultz, J.; Rick, J.M.; Michon, A.-M.; Cruciat, C.-M.; et al. Functional organization of the yeast proteome by systematic analysis of protein complexes. Nature 2002, 415, 141–147. [Google Scholar] [CrossRef]
  52. Jothi, R.; Balaji, S.; Wuster, A.; Grochow, J.A.; Gsponer, J.; Przytycka, T.M.; Aravind, L.; Babu, M.M. Genomic analysis reveals a tight link between transcription factor dynamics and regulatory network architecture. Mol. Syst. Biol. 2009, 5, MSB200952. [Google Scholar] [CrossRef]
  53. Bascompte, J.; Jordano, P.; Melian, C.J.; Olesen, J.M. The nested assembly of plant-animal mutualistic networks. Proc. Natl. Acad. Sci. USA 2003, 100, 9383–9387. [Google Scholar] [CrossRef]
  54. Peñ a, J.; Rochat, Y. Bipartite graphs as models of population structures in evolutionary multiplayer games. PLoS ONE 2012, 7, E44514. [Google Scholar] [CrossRef][Green Version]
  55. Calladine, C.R.; Drew, H.R.; Luisi, B.F.; Travers, A.A. Understanding DNA: The Molecule and How It Works, 3rd ed.; Elsevier/Academic Press: London, UK, 2004. [Google Scholar]
  56. Wang, G.; Vasquez, K.M. Non-B DNA structure-induced genetic instability. Nat. Rev. Genet. 2017, 18, 469–484. [Google Scholar] [CrossRef]
  57. Oparin, A.I. Vozniknovenie Zhizni na Zemle (The Origin of Life on Earth); Izd. Akad. Nauk SSSR: Moscow, Russia, 1936. [Google Scholar]
  58. Haldane, J.B.S. The Origin of Life. Ration. Annu. 1929, 148, 3–10. [Google Scholar]
  59. Broox, F.P., Jr. Three great challenges for half-century-old computer science. J. ACM 2003, 50, 25–26. [Google Scholar]
  60. Nevil-Manning, C.; Witten, I. Protein is incompressible. In Proceedings of the Data Compression Conference, Snowbird, UT, USA, 29–31 March 1999; p. 257. [Google Scholar]
  61. Monod, J. Chance and Necessity; Collins: London, UK, 1972. [Google Scholar]
  62. Schwartz, R.; King, J. Sequences of hydrophobic and hydrophilic runs and alternations in proteins of known structure. Protein Sci. 2006, 15, 102–112. [Google Scholar] [CrossRef]
  63. White, S.; Jacobs, R. Statistical distribution of hydrophobic residues along the length of protein chains. Biophys. J. 1990, 57, 911–921. [Google Scholar] [CrossRef] [PubMed]
  64. Pande, V.; Grosberg, A.; Tanaka, T. Nonrandomness in protein sequences: Evidence for a physically driven stage of evolution. Proc. Natl. Acad. Sci. USA 1994, 91, 12972–12975. [Google Scholar] [CrossRef] [PubMed]
  65. Weiss, O.; Jiménez-Montañgo, M.; Herzel, H. Information content of protein sequences. J. Theoret. Biol. 2000, 206, 379–386. [Google Scholar] [CrossRef] [PubMed]
  66. Vanchurin, V.; Wolf, Y.I.; Koonin, E.V.; Katsnelson, M.I. Thermodynamics of evolution and the origin of life. Proc. Natl. Acad. Sci. USA 2022, 119, e2120042119. [Google Scholar] [CrossRef]
Figure 1. Left: String s = x 4 y 10 x 13 y 13 x 5 y 5 and its corresponding monotone discrete path. Right: Another string that has nearly the same a d v and m d v as s but a different n l m .
Figure 1. Left: String s = x 4 y 10 x 13 y 13 x 5 y 5 and its corresponding monotone discrete path. Right: Another string that has nearly the same a d v and m d v as s but a different n l m .
Mathematics 14 01760 g001
Figure 2. Left: Regular triangular helix (regular skew-apeirogon). Right: Irregular triangular helix (irregular skew-apeirogon).
Figure 2. Left: Regular triangular helix (regular skew-apeirogon). Right: Irregular triangular helix (irregular skew-apeirogon).
Mathematics 14 01760 g002
Figure 3. Conceptual illustration of the bounded region of geometric space compatible with living genomes. Real genomes occupy a narrow intermediate band between highly random and overly regular sequences, shown schematically as a life-permissible region. This diagram summarizes the structural constraints suggested by the empirical results.
Figure 3. Conceptual illustration of the bounded region of geometric space compatible with living genomes. Real genomes occupy a narrow intermediate band between highly random and overly regular sequences, shown schematically as a life-permissible region. This diagram summarizes the structural constraints suggested by the empirical results.
Mathematics 14 01760 g003
Figure 4. Absolute and normalized a v d and m d v measured for substrings and subsequences of increasing length for the 25 species of Stage 1. These graphs show that a d v and m d v are essentially independent of length, since the lines remain in the same relative positions as length varies. Note also that the lowest line in each graph (which stands out significantly when the samples are substrings) corresponds to the random sequence.
Figure 4. Absolute and normalized a v d and m d v measured for substrings and subsequences of increasing length for the 25 species of Stage 1. These graphs show that a d v and m d v are essentially independent of length, since the lines remain in the same relative positions as length varies. Note also that the lowest line in each graph (which stands out significantly when the samples are substrings) corresponds to the random sequence.
Mathematics 14 01760 g004
Figure 5. Conceptual comparison of deviation from linearity m d v and local maxima counts n l m between real genomes and random sequences. Real genomes occupy a narrow intermediate range for both statistics, while random sequences lie far outside this region, illustrating the geometric bounds compatible with living genomes.
Figure 5. Conceptual comparison of deviation from linearity m d v and local maxima counts n l m between real genomes and random sequences. Real genomes occupy a narrow intermediate range for both statistics, while random sequences lie far outside this region, illustrating the geometric bounds compatible with living genomes.
Mathematics 14 01760 g005
Figure 6. Left: Deviation from linearity m d v values for all examined species, shown as individual points along the horizontal axis. Species are arranged along the axis in descending order of their m d v values, with primates highlighted. The random sequence appears as an extreme outlier. Right: Number of local maxima n l m in cumulative count paths for each species, shown as individual points along the horizontal axis. Species are arranged along the axis in descending order of their n l m values, with primates highlighted. The random sequence appears as a high- n l m outlier.
Figure 6. Left: Deviation from linearity m d v values for all examined species, shown as individual points along the horizontal axis. Species are arranged along the axis in descending order of their m d v values, with primates highlighted. The random sequence appears as an extreme outlier. Right: Number of local maxima n l m in cumulative count paths for each species, shown as individual points along the horizontal axis. Species are arranged along the axis in descending order of their n l m values, with primates highlighted. The random sequence appears as a high- n l m outlier.
Mathematics 14 01760 g006
Figure 7. Joint distribution of deviation from linearity m d v and local maxima n l m across 50 species. Each point (blue disc) represents one genome, with primates highlighted (purple diamonds). Real genomes cluster in a compact region of geometric space, while the random sequence (orange square) appears as a distant outlier.
Figure 7. Joint distribution of deviation from linearity m d v and local maxima n l m across 50 species. Each point (blue disc) represents one genome, with primates highlighted (purple diamonds). Real genomes cluster in a compact region of geometric space, while the random sequence (orange square) appears as a distant outlier.
Mathematics 14 01760 g007
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Brimkov, V.E.; Barneva, R.P. Geometric Structure of Genomes Across the Tree of Life: Toward a Geometric Theory of Sequence Structure. Mathematics 2026, 14, 1760. https://doi.org/10.3390/math14101760

AMA Style

Brimkov VE, Barneva RP. Geometric Structure of Genomes Across the Tree of Life: Toward a Geometric Theory of Sequence Structure. Mathematics. 2026; 14(10):1760. https://doi.org/10.3390/math14101760

Chicago/Turabian Style

Brimkov, Valentin E., and Reneta P. Barneva. 2026. "Geometric Structure of Genomes Across the Tree of Life: Toward a Geometric Theory of Sequence Structure" Mathematics 14, no. 10: 1760. https://doi.org/10.3390/math14101760

APA Style

Brimkov, V. E., & Barneva, R. P. (2026). Geometric Structure of Genomes Across the Tree of Life: Toward a Geometric Theory of Sequence Structure. Mathematics, 14(10), 1760. https://doi.org/10.3390/math14101760

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop