2.1. Shannon’s Algebra for Population Genetics
In Shannon’s genetic algebra [
1], as in any algebra, it is important to understand the notation used to describe the operations. Shannon introduces a notation in his work capable of representing the probability of observing each allele in the genetic loci of interest in the context of the population’s genotype. A typical representation of the probabilities involved in two genetic loci is of the form
, a symbolic representation that can be interpreted in the following manner:
Within Shannon’s algebra, Greek letters are used, as base letters, to describe each population. In this case, the population’s base letter is “λ”, whereas commonly used letters within Shannon’s original work are “μ, υ, θ”. Furthermore, for the base letter, in order to describe genetic information “stored” in S loci, there are pairs of indices. Each pair consist of a subscript and a superscript describing one of the S loci (i.e., in the symbolism
, “
” and “
” are used to identify the two allele types in the first locus). Each pair of indices can take a range of different values for each locus, equal to the number of different allele types that can appear in the specific population at that specific locus. Therefore, in the symbolism of the genetic information stored in S gene loci in the form of the frequency of observing a specific combination of an allele types in a population, S columns and two rows are used, with the exception of the sex chromosomes. The two rows, one written as superscripts and one as subscripts, represent the genetic information inherent to each of the parents. Thus, in the notation
and
correspond to the two homologous chromosomes at the first locus, while
and
correspond to the two homologous chromosomes at the second locus. Consequently,
and
originate from different parental chromosomes, and the same is true for
and
. It should be noted, however, that Shannon’s notation does not explicitly encode chromosomal location. Therefore, the loci represented by
and
, and similarly
and
, may or may not reside on the same physical chromosome. Likewise, the notation does not specify whether the alleles
and
or
and
were inherited from the same ancestral chromosome, as this information is captured separately through the recombination probabilities used in Shannon’s genetic cross-product. Each pair of
(with s being either 1 or 2 in this case) is located in homologous chromosomes, with each column representing a genetic locus. Indices of the same column have the same range of values that they can take since the same type of alleles can be inherited from either parent in a given population. For example, in order to describe the genetic information stored in the form of the genotypic frequencies on two loci in a population where there are two possible allele types s for the first genetic locus and three allele types s for the second, the indexes
,
can take the values 1 and 2, identifying each allele type at the first locus, whereas the index
can take the values 1, 2, 3, identifying each allele at the second locus. It should be stressed that these numbers identify a specific allele type in a specific locus and have no relation whatsoever to the numbering of the other loci. One may appreciate the compactness of Shannon’s formalism even in this simple example since the genetic information is stored in only two loci, corresponding to the 2
2 × 3
2 = 36 frequencies of observing each possible combination of allele types, i.e.,
with
When specific integer values are assigned to all indices, the resulting symbol represents the frequency (or, equivalently, the probability) of observing a particular multilocus genotype in the population. The values assigned to the indices specify the allelic variants present at each genetic locus and therefore uniquely identify the corresponding genotype configuration., e.g., represents the frequency of observing in the first locus the allele type 1 (the first allele type at the first locus) in both homologous chromosomes and the alleles of type 2 and 3 in the second locus (the second and third allele type of the second locus). It is important to understand that not all values on those frequencies are initially considered as independent, and Shannon imposes dependences and indistinguishability in the genetic information in the form of equations expressed as theorems. The easiest such dependence comes from the fact that all frequencies should sum to one.
In the language of linear algebra, the knowledge of the number of possible allele types at each genetic locus under consideration is sufficient to establish an upper bound on the dimensionality of the vector space that embeds such genetic information. This dimensionality is equal to the square of the product of the numbers of possible allele types at each locus. In the previous example, this dimensionality was (2 × 3)2. Any additional independent linear constraint that results from applying Shannon’s theorems reduces the dimensionality of the embedding vector space by one (for each resulting equation), significantly reducing the submanifold that embeds the span of the vectors that describe the genetic information. There are two key concepts underlying most of the equations that constitute Shannon’s theorems. The first is that genotypes that are indistinguishable in terms of the genetic information transmitted from one generation to the next generation impose symmetries that can be expressed as algebraic equations. The second is that summing the frequencies of a set of outcomes yields the probability of observing any of the genotypes that belong to that set. Consequently, summing over all possible combinations of allele types must result in a total probability equal to one 1. Because summation over sets of possible genotypes occurs frequently in Shannon’s algebra, he introduced a compact notation for summing over specific indices. This operation corresponds to a projection onto a lower-dimensional vector space representing the outcomes that can be observed when no information is available about one or more of the genetic loci being summed over. Such projections from higher- to lower-dimensional spaces are represented through an abbreviated notation in which a given index is replaced by a dot symbol •. The dot indicates summation over all possible allelic variants associated with that index.
Suppose as in the previous example that the indices
and
can take values in the ranges
and
, respectively. Then summation over the index
is represented by replacing i
2 with dot symbol •, yielding
which represents the probability to observe all possible events for the three remaining allele types, independent of the value of
type. Similarly, summing over both
and
, we obtain the probabilities for all possible combinations involving the remaining allele types described by
and
represented in Shannon’s notation as
Similarly, summing over
and
results in Equation (3):
It should be noted that using a subset of loci in the notation is equivalent to summing over all remaining loci in a representation that explicitly includes every genetic locus. For example, since the omitted loci are implicitly summed over all possible allelic configurations. Most importantly, summations over any index can be interpreted as the projections of the corresponding probability vector from a higher-dimensional space to a lower-dimensional space, retaining information only about the remaining allele types. In this sense, Shannon’s summation notation provides a compact representation of marginalization over selected genetic loci. To avoid confusion, we should note that Shannon’s symbolic use of subscripts and superscripts is unrelated to the modern representation of tensors.
Having defined a symbolic representation for genetic information in terms of the allele frequencies for any population and any set of genetic loci, Shannon introduces his main algebra operation, which relates the genetic information of the offspring population
to that of the two parental populations λ and μ in the form
Shannon showed how his definition of a genetic product can be derived by enumerating and appropriately weighting all possible events that result in each specific combination of allele types in each genetic locus as the result of random mating between the parental population, deriving from the probabilities associated with the transmission of gametes. By construction, the genetic product on which Shannon’s algebra is based relies on the same basic assumption in Mendelian genetics and, in particular, on the assumptions used to derive the Hardy–Weinberg genetic equilibrium in infinite large populations undergoing random mating in the absence of mutations. To be more precise, the basic assumptions in Shannon’s algebra are random mating, independent association for genes in different chromosomes, and deviation from independent association due to genetic linkage between genes located on the same chromosomes, as well as the absence of mutations, gene flow, genetic drift, natural selection and selective breeding. The absence of mutations and gene flow imposes that no introduction of new alleles or loss of existing alleles is possible, either through mutation in existing alleles or through the migration of individuals into or out of the population. Similarly, the absence of genetic drift implies that the population is sufficiently large to prevent the disappearance of a low-frequency allele from the population. The absence of natural selection and selective breeding implies that all genotypes have an equal probability of producing offspring. It should be noted that these assumptions are essential for constructing the main mathematical framework of Shannon’s algebra. Once this framework has been established, it can be extended to incorporate additional biological processes. For example, Shannon proposed in his thesis a way to incorporate natural selection via an additional step in his formulation. In this work, we extend Shannon’s algebra by showing how it is possible to extend his formalism in dealing with samples of limited size.
In this section, we present a summary of Shannon’s most important theorems. Readers interested in detailed derivations are referred to the Master’s thesis of Georgios Kousoulis, “On Shannon’s Algebra for Theoretical Genetics” [
1].
According to Shannon’s first theorem, the genetic information in each of the homologous chromosomes on all autosomal pairs of healthy offspring is equal and indistinguishable.
Theorem 1. Theorem 1 (Equation (5)) dictates that the genetic information inherited from each of the parents is indistinguishable in terms of its probability of being transmitted to the next generation and is therefore treated identically within the context of Shannon’s genetic algebra. It is important to note that this theorem requires the simultaneous interchange of all superscript and subscript indices. However, the interchange of only some, but not all, alleles does not generally hold. In other words, indices cannot be freely exchanged between rows without restrictions. As is shown later, all loci located on the same chromosome must participate in such an interchange. This requirement arises because the inheritance of alleles located on the same chromosome introduces correlations between them that depend on the distance separating the loci, a phenomenon commonly referred to as genetic linkage [
1], measured in centimorgan, making alleles that come from the same parent placed in the vicinity of the same chromosome have a higher probability of being passed together in the next generation, creating a correlation. In contrast, loci located on different chromosomes are inherited independently through independent assortment. Consequently, the genetic information associated with such loci is transmitted independently, giving rise to the corresponding indistinguishability relations within Shannon’s algebra.
For example, in the case of a given population where three loci are under consideration, with the first two genetic loci residing on the same chromosome pair and the third locus located on a different chromosome pair, the following equalities can be used to represent the indistinguishability of the corresponding genetic information.
In Equation (6), the indices are rearranged in groups that include all alleles located on the same chromosome. Consequently, alleles on the same chromosome must be reversed collectively as a single unit. As mentioned before, it would be incorrect to impose that
is always equal to
, because the alleles
h,
i and
k,
l belong to the same chromosomes, and the column containing allele
h was not reversed. To allow for the reversal between
i and
l, the alleles
h and
k must also be reversed. In Shannon’s original work, the information on the location of loci at different chromosomes has not been included, but it can be added as additional equations imposing the known symmetries in the elements of the population frequencies. In order to systematically include directly in the formulation this additional information as part of the encoding, for the cases in which the genetic loci of alleles are known beforehand, a new notation can be introduced, as proposed by G. Kousoulis [
15]. This involves modifying Shannon’s original formalism by introducing a comma to separate genetic loci that are known to reside on different chromosomes. Using this notation, the above equations can be rewritten as
This revised notation conveys additional information compared to the original, allowing it to be used not only for distinguishing alleles on different chromosomes but also as the result of an inferential problem when analyzing a population to determine whether genetic loci reside on the same or on different chromosomes.
According to Shannon’s second theorem reported in Equation (8), the summation over all indices must equal to 1, implying that the summation over all indices is the sum of normalized probabilities over all possible combinations of events.
Theorem 2. As previously discussed, the “•“ symbol represents the summation of all possible combinations of a specific allele but equivalently describes no knowledge of a specific allele in the corresponding locus. As a consequence, representing only a subset of loci is equivalent to using the “•“ symbol in all the other loci. As a third theorem, Shannon introduces his main result in the definition of a cross-product operation. Shannon’s cross-product operation expressed the population frequencies of the offspring population based on the two parent populations () assuming random mating.
Theorem 3. Theorem 3 (Equation (9)) presents Shannon’s cross-product in the simplest case involving two genetic loci, but it can be extended to describe the general case of an arbitrary number of genetic loci following a standard procedure. The essential concept behind Shannon’s approach is the probability of producing each possible genotype based on the gametes generated by the parent population and accounting for all possible outcomes that come from all other crossovers or the absence of crossovers between the genetic loci under study. Note that, at this level, a crossover event is not been distinguished from a shuffling in the creation of the gametes due to the independent assortment process when the loci belong to different chromosomes, except for the values that can take, since the case of independent assortment should always have= 0.5. For the simplest case of two genetic loci, one needs to define either the probability of zero or an odd number of crossovers p0 or the probability of one or even number of crossovers p1, since p0 + p1 = 1 accounts for all possible events. One way to rationalize Equation (9) that can also be extended in the general case of s loci is to realize that Equation (9) describes the various ways through which an offspring can inherit a specific genotype. Much like an algebraic expression of a function f(x) = x + 2 that provides expressions for each value of x in a domain of definitions, Equation (9) represents an algebraic expression for each possible specific allele type, with domains of definition, different for each locus, for the specific allele types that are under consideration in each of the loci, in the offspring population and the parental populations and . We should also know that Equation (9) is agnostic as to whether the loci are on the same chromosome or not, or if they are close to each other. This information is actually embedded in the values p0 and p1, with both of them equal to 0.5 in the absence of genetic linkage, either because the loci are in deferent chromosomes or because the two loci are at the same chromosome but too far from each other. In the context of Shannon’s formalism, what in population genetics is described as genetic linkage corresponds to values of p1 in the range 0.5 > p1 > 0, where p1 = 0 represents the cases where both loci are coinherited, always passing from generation to generation as a set, with no probability of even crossovers between them. Therefore, Equation (9) states that represents the frequency of members in the population in the algebraic sense, i.e., the symbols h, j, I, k are algebraic variables that take independent values from the set of the possible allelic types in each locus, with h and j having the domain of definition of the allelic types of the first locus and l and k being the allelic types of the second locus. In Shannon’s formalism, represents a situation in which the first locus has a specific allele type h in one of the homologue chromosomes and an allele type I in the same locus in the homologue chromosome, and the second locus has a specific allele type j in one of the homologue chromosomes (not necessarily the same as for the precise locus) and the specific allele type k in its homologue chromosome. As mentioned above, the reader should keep in mind when reading Equation (9) that the inability to distinguish between homologous chromosomes within Shannon’s formulation is manifested in the form of equations that impose symmetries between all undistinguishable genotypes. Although Equation (9) can be written in a much simpler form, the current form depicts most clearly the enumeration of all possible inheritance paths that lead to the members of the population that are described by . Since each ancestor can provide an offspring with either the allele h or the allele j but not both, because they are located at the same locus, and similarly either the allele l or k in the second locus, each offspring contributing to the frequency can only be the outcome of the following inheritance paths:
Given that no preference is expected, both scenarios have equal weight, resulting in the terms in Equation (9). Having separated the alleles inherited from each parent, one must then account for the possibilities of receiving them with or without a crossover event, resulting in two alternative scenarios that lead to the observed “gametes”. In the first scenario, both alleles that pass to the offspring came from the same chromosome of the parental population, and either no recombination occurs or an even number of recombination events occurs. This is represented in Shannon’s formula with the terms , , , , with representing the probability of no recombination and , , , representing the frequency of the parental genotypes that, without recombination, can provide the offspring with the specific set of alleles. In the second scenario, the alleles that are being transmitted to the offspring originate from homologue chromosomes in the parental genotype, that is, from a different grandparent, due to recombination. The probabilities of these events are represented by the terms , , , , where denotes the probability of recombination and . Combining all possible inheritance “scenarios” yields a total of eight alternative combinations that contribute to the frequency of members in the offspring population represented by We should also stress that the inability to distinguish between certain hereditary configurations is expressed through additional algebraic equations in Shannon’s formulation. For example, the indistinguishability of the two homologue chromosomes is represented by the relations and =.
Similarly to Equation (9), Shannon’s cross-product for the case of three genetic loci is given by
In Equation (9), one has to account for the presence or absence of a single recombination event between two loci. In the general case of S loci, the number of possible recombination events that must be considered is S-1. It should be noted that the proposed approach of representing all possible events is sufficient but not a unique representation. Most importantly, there is no a priori information on the order of the loci, apart from the condition of consistency between the order in the locus as used in the population frequencies and the probability of recombination . In order to distinguish between the probabilities of observing or not observing S-1 recombinations between adjacent loci, is assigned S-1 subscripts. Each subscript represents the presence, with subscript 1, or absence, with subscript 0, of recombination between adjacent loci as they are listed in the population frequency superscripts and subscripts. For example, in the case of three loci, one has to define ,,,, in order to represent the probabilities of no recombination in either adjacent pair of loci , recombination only in the second pair of loci , recombination only in the first per of loci pair , and finally recombination only in both pairs of loci . As before, the values of these variables are constrained since they represent probabilities and must sum to one (), as in the case of Equation (9), constructed through the enumeration of the inheritance “scenarios” that lead to a specific genotype .
The same procedure can be used to drive Shannon’s genetic cross-product for any number of loci by appropriately defining the recombination probabilities between all adjacent loci, with the exception of the simplest case—that of a single genetic locus—for which the concept recombination does not apply. For the simplest case of a single genetic locus, there are only two possibilities: the offspring may inherit the chromosome-carrying allele
h from the maternal population and the chromosome with the
i allele from the paternal population, and vice versa. Because they are equally possible, the equation giving the probability of the offspring population having the genotype equals
As in the previous cases, the use of a • in the algebraic representation of a population frequency means that knowing the type of the allele is irrelevant to the inheritance process, for example, provides the probability of observing the allele type in the maternal population by summing over all combinations of allele types on the homologue’s chromosome. Note that not listing a locus is equivalent to using • for both homologues’ chromosomes. By construction, Shannon’s cross-product is commutative for any number of loci S, meaning that there is symmetry in swapping the population frequencies of the parental populations.
Theorem 4. As mentioned before, Shannon’s original formalisms deals with loci that are not located on the sex chromosome. Any extension of the formalism to include the sex chromosome must take into account that the frequencies of loci located on the sex chromosome will not be commutative. Having defined an algebraic operation capable of describing the population frequencies of an offspring compatible with the Mendelian inheritance, including the effect of the recombination process, Shannon provides a number of interesting theorems satisfied by his cross-product: Theorem 5 (Equation (13)) states that, when two populations share identical frequencies for all the alleles at a given genetic locus and when each is crossed with another population, they will exhibit identical breeding characteristics, as long as the specific locus is taken into account.
Theorem 5. Another interesting property of Shannon’s cross-product is that it is distributive with respect to the addition of population frequencies and the multiplications by real numbers. This means that one can express a population as the weighted sum over more than one population,, and estimate the offspring population resulting from crossing with another population such as the weighted sum of the pair of populations shown in Equation (14). Theorem 6. where the multiplication of a real number with the population frequencies, and the additions between population frequencies, are defined in the same manner as scalar vector multiplication and vector addition. A direct consequence of Shannon’s product is the existence and stability of the Hardy–Weinberg genetic equilibrium, stating that, if a population has reached equilibrium, then further mating will result in offspring with the same genetic frequencies. For example, in the case of a single locus, this equilibrium population will be given by Equation (15).
where Hardy–Weinberg genetic equilibrium is defined as the condition where, in the absence of perturbations (such as mutations or genetic drift), the allele’s frequencies remain constant during successive generations [
5]. For the simple case of a single locus, the concept of equilibrium in Shannon’s genetic algebra is identical to the Hardy–Weinberg equilibrium, but it becomes more interesting when one considers the approach to equilibrium in cases with more than one locus. The Hardy–Weinberg equilibrium for a locus with two alleles A and a is represented by (A + a)
2 = 1 → A
2 + 2Aa + a
2 = 1 and A + a = 1. Let the frequency for allele A be equal to 0.7 in the population; then, a = 0.3. Plugging these values into the first equation gives the solution (0.7 + 0.3)
2 = 0.49 + 0.42 + 0.09 = 1. This is translated to the frequencies of AA = 0.49, Aa + aA = 2aA = 0.42, and aa = 0.09. Furthermore, for the same example illustrated using Genetic Shannon notation, let allele A in the population be denoted by
and allele a by
; additionally,
=
. From the above, it can be deduced that AA =
, 2aA =
and aa =
. For this simple case of a single locus, the equilibrium is reached in a single generation, provided that both parental populations have identical genotypic frequencies
for each locus. Using identical parental populations in Equation (11) results in Equation (16):
It is then easy to confirm that the further offspring of this population when mated with itself in subsequent generations will continue to satisfy the genetic equilibrium conditions in Equation (17).
Another interesting concept of Shannon’s algebra comes when one considers
as a regular vector and
as a two-dimensional matrix in the context of linear algebra. In this setting, for a population in genetic equilibrium and with a single locus under consideration, the following three statements are equivalent:
Let , h and i take values from the set (1, 2, 3); then, . Furthermore, from Equation (15), which leads to the matrix .
Additionally, matrix operations allow the multiplication or division of all the elements of a line, leading to a Row-Equivalent Matrix, which does not change the rank of a matrix. Under these terms, if row 1 is divided by
, row 2 is divided by
, and row 3 is divided by
, this will lead to a new matrix
that has all rows identical, leading to the conclusion that it has rank 1 [
16].
As mentioned before, Shannon’s approach not only results in the Hardy–Weinberg equilibrium for an arbitrary number of loci, but most importantly it also describes the steps leading to the equilibrium when the initial populations are out of equilibrium.
For example, Shannon has shown that, if a population reproduces randomly, then the offspring population after n generations will be given by Equation (19).
Consistent with the concept of the Hardy–Weinberg equilibrium, the moment that even the smallest probability of recombination is possible,
(and
), and the offspring population will tend to the equilibrium population
ας described by Hardy–Weinberg equilibrium [
17] This is because the number of offspring generations
n appears as a power of the probability of no recombination, whose limit (
) is either 0 for
or 1 if
. Therefore, if
, the population will approach the equilibrium population
. Note that, if
, then
, which implies that crossover is not allowed, and the two loci are always passed from generation to generation together, resulting in a genetic equilibrium that can be described by the knowledge of allele type in one of the loci,
. Another way to rationalize this limit is to realize that, if no recombination is allowed, then the two distinct loci can be represented by a single locus that encodes the information from both loci, and
is the Hardy–Weinberg equilibrium of that single locus. After all, this is the reason that grouping part of the genetic code into allele types is sufficient for describing Mendelian inheritance.
Similarly, one can construct the approach to equilibrium for an arbitrary number of loci by recursively applying Shannon’s cross-product
n times. While the expression becomes intrinsically lengthy as the number of loci increases, with the case of three loci depicted in Equation (20), spanning several lines, it can be easily derived using a computer.
As in the case of two loci, as the limit n increases, the Hardy–Weinberg equilibrium is observed. If recombination is possible between all loci the genetic equilibrium is . In this case, the genetic frequencies at equilibrium will be the product of the allele type frequencies. If recombination is impossible between certain sets of loci, the genetic equilibrium will have contributions from a single representative of each set.
Another interesting property of the genetic frequency pointed out by Shannon is that the allele type frequencies within the framework of linear algebra form a vector space (i.e., act like usual multi-dimensional vectors) equipped with addition and multiplication with a real number. Interestingly, within this context, Shannon has also shown that any given population φ can be uniquely represented as a linear combination of n linearly independent populations, where n is the number of different genetic formulas under consideration.
As mentioned above, in this work, we review the essential aspects of Shannon’s work related to Shannon’s genetic cross-product and its relation to linear algebra. Beyond these, Shannon has also introduced other interesting concepts, such a time derivative of the population frequencies and attempts to incorporate mutations in his model. These topics go beyond the scope of the present work and the reader is referred to Shannon’s original publication for further details.
2.4. Shannon’s Algebra and Analytical Constraints on Offspring Phenotypic Frequencies
In the spirit of Claude Shannon’s algebraic approach, we have recently proposed a linear-algebraic framework that analytically relates phenotypic frequencies to compatible genotypic and allelic frequencies [
14]. Within this framework, genotypic information, expressed at the level of allele frequency distributions, is directly linked to phenotypic information, expressed at the level of phenotypic trait frequencies, thereby enabling both the inference and validation of genotype–phenotype mappings. Furthermore, we have shown [
14] that it is possible to construct analytical solutions for all genotypic frequency distributions compatible with a given phenotypic distribution.
The reconstruction of genotype–phenotype relationships through phenotype-constrained bootstrapped subsampling (Constrained Observation and Null Space-based Inference, CONSPIN) [
14] was originally derived for cases of complete penetrance, where the genotype determines the phenotype with 100% certainty, but it has also been shown to be applicable in cases of stochastic expression. In our previous work [
14], we focused on genotype–phenotype mapping within a single generation and examined the analytical constraints imposed on genotypic frequencies when the phenotypic frequencies of a sample are known. It is straightforward to demonstrate that, by combining this analytical genotype–phenotype mapping approach with Shannon’s genetic algebra, analogous analytical relations can be derived that link the phenotypic information of a parental generation to that of one or more offspring generation(s), at the level of compatible phenotypic frequencies.
Following the derivation in [
14], phenotypic information is encoded as the distribution of phenotypic traits in a population and is represented by a column vector
, where the dimensionality corresponds to the total number of phenotypes. Each component
represents the frequency of the associated phenotype and is defined as the ratio of individuals expressing that phenotype to the total population size:
Similarly, genotypic information is represented [
14] by a column vector
, where the dimensionality corresponds to the total number of distinct genotypes. Each component
denotes the frequency of the corresponding genotype and is defined as
In relation to Shannon’s representation, can be viewed as an enumeration of all possible genotypes in vector form, where each component corresponds to a specific fixation of indices in Shannon’s notation, analogous to the construction of in Equation (18), with the difference that all genotypes are listed in a vector rather than a matrix.
Finally, the frequency of observing each allele in the population is represented by the vector
, whose dimensionality equals the total number of distinct alleles. Each component
represents the frequency of the corresponding allele and is defined as
In Shannon’s notation,
, where
denotes the fixation of the index corresponding to allele k. Based on the definitions in Equations (22)–(24), allelic interactions (i.e., the expression of a phenotypic trait given a genotype) can be represented by the allelic interaction matrix
Here,
is an m ×
n matrix, where m is the number of phenotypes (rows) and
n is the number of genotypes (columns). Each row corresponds to a phenotype, and each column corresponds to a genotype [
14].
Similarly, allele frequencies can be expressed via a matrix–vector product:
where
accounts for the contribution of each genotype to the allele pool [
14]. Its elements take values of 1 for homozygous genotypes, 0.5 for heterozygous genotypes, and 0 for genotypes that do not contain the corresponding allele.
As we have shown [
14], for any sample with fixed phenotypic frequencies
, Equations (25) and (26) imply that
and
cannot, in general, be uniquely determined, since the number of unknowns exceeds the number of equations. However, all compatible solutions can be expressed analytically as functions of a set of auxiliary independent variables
. The number of such variables is given by the difference between the number of unknowns and the number of equations in Equation (25). Thus, we may write
.
As we show in the Results section, combining this analytical framework with Shannon’s genetic algebra allows us to derive analytical expressions for all possible offspring phenotypic frequency vectors that are compatible with a given parental phenotypic distribution .
The key idea is as follows: given the allelic interaction matrix and parental phenotypic frequencies , one can determine the set of compatible parental genotypic frequencies as which, by virtue of Equation (26), leads to . Using Shannon’s algebra to estimate the offspring genotype frequencies of the offspring generation as a function of the parental haplotype allele frequencies and the appropriate set of recombination parameters , we can estimate the expected values for the offspring genotypic frequencies (and thus ):. Finally, applying Equation (25), the offspring phenotypic frequencies are obtained as
The overall process can be summarized as
For a given choice of auxiliary variables , this procedure yields one possible offspring phenotypic distribution that is consistent with the parental phenotypic frequencies . In the Results section, we present the simplest case of a single locus with two alleles.