Next Article in Journal
Aberrant Expression of Human Endogenous Retroviruses and SETDB1 in Adolescents with Anorexia Nervosa
Previous Article in Journal
Vascular Endothelial Growth Factor and Placental Growth Factor in Conjunction with Vascular Endothelial Growth Factor Receptor-1 May Exert Dual Effects Within the Kidney and Brain in Patients with Type 2 Diabetes Mellitus and Normoalbuminuric Diabetic Kidney Disease
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

The Distribution of Average Pairwise Distances Among Human Pre-miRNAs for Disease Association Analysis

Institute of Statistics, National Yang Ming Chiao Tung University, Hsinchu 300093, Taiwan
*
Author to whom correspondence should be addressed.
Int. J. Mol. Sci. 2026, 27(9), 3750; https://doi.org/10.3390/ijms27093750
Submission received: 11 March 2026 / Revised: 9 April 2026 / Accepted: 21 April 2026 / Published: 23 April 2026

Abstract

MicroRNAs (miRNAs) play essential roles in cell differentiation, development, gene regulation, and apoptosis, and have been widely implicated in numerous disease mechanisms. Owing to their regulatory importance, miRNAs are increasingly recognized as valuable disease biomarkers. Previous studies have used nucleotide sequence pairwise distances between miRNAs to explore disease associations and have derived the distribution of pairwise distances to assess the percentile rank of an observed miRNA pair. However, because a single disease may involve multiple miRNA biomarkers, evaluating the percentile rank of an average pairwise distance is often more appropriate than focusing on individual pairs. In this study, we established percentile distributions for the average pairwise nucleotide distances corresponding to different numbers of miRNAs. Applying this framework to 51 diseases and several groups of related diseases, we observed that miRNA biomarkers associated with the same disease, as well as with related diseases, often exhibit low-percentile average pairwise distances under the reference distribution. While the present study does not directly evaluate whether precursor miRNA (pre-miRNA) sequence similarity is associated with shared biological function or regulatory targets, the proposed framework provides a systematic approach for quantifying such similarity among disease-associated miRNAs and may serve as a useful foundation for future studies integrating functional and clinical validation.

1. Introduction

MicroRNA (miRNA) is a small RNA molecule, typically 20–23 nucleotides long, that plays an important role in various biological functions such as gene regulation, cell differentiation, apoptosis, and immune response [1,2,3]. The first miRNA, lin-4, was discovered in the early 1990s while researching the nematode Caenorhabditis elegans. Until 1993, it was revealed that lin-4 does not encode a protein but produces a ~22-nucleotide RNA that is partially complementary to the 3′ untranslated region (UTR) of lin-14 mRNA [4]. In 2000, the second miRNA, let-7, was identified, which regulated the translation of the lin-41 gene via similar mechanisms [5]. These findings suggest that the regulation of mRNA by miRNAs is a common feature.
The biogenesis of miRNAs has multiple processes. miRNA biogenesis occurs through two main pathways: the canonical and non-canonical pathways [6]. The canonical pathway is the primary route. In this process, the primary miRNA transcript (pri-miRNA) is first cleaved in the nucleus by the ribonuclease (RNase) III enzyme Drosha, producing a precursor miRNA (pre-miRNA). This pre-miRNA, characterized by a hairpin-shaped secondary structure, is then exported to the cytoplasm, where it is further processed by another RNase III enzyme, Dicer, into the mature miRNA [7]. The maturation of miRNAs thus involves two sequential steps carried out by Drosha in the nucleus and Dicer in the cytoplasm [8,9]. Drosha functions in complex with DGCR8 to initiate miRNA processing within the nucleus. In contrast, non-canonical miRNA biogenesis pathways diverge from this standard route and involve alternative combinations of canonical components. These non-canonical pathways can generally be categorized as either Drosha/DGCR8-independent or Dicer-independent [8]. Although the mature miRNA is the functional form responsible for gene regulation, the pre-miRNA can influence its function, as discussed in detail in the Discussion section. In this study, we focused on pre-miRNA sequence analysis.
miRNAs also play a vital role in stem cell maintenance and function and are strongly linked to cancer development and pathology [10,11]. Therefore, miRNAs can be used as useful biomarkers for various cancers such as lung, gastric, and brain cancers [12,13,14]. miRNA dysregulation in cancer is driven by complex mechanisms such as genomic alterations, transcriptional regulation, epigenetic modifications, and interactions with target mRNAs and cellular signaling pathways [15]. In addition to cancer, miRNAs have served as useful biomarkers for various diseases, including neurological diseases and cardiovascular diseases, because they are involved in these disease pathologies [16,17]. miRNAs have strong potential as sensitive and specific biomarkers for kidney diseases due to their tissue specificity and stability across various biological materials [18]. Salivary miRNAs show promise as biomarkers for the detection of sports-related concussion [19].
Since miRNAs can serve as useful biomarkers for diseases, and many distinct miRNAs have been associated with specific diseases, it is worth investigating whether miRNA biomarkers for the same disease exhibit a high level of similarity. In this study, high similarity between two miRNAs corresponds to a low pairwise distance between their nucleotide sequences. Therefore, we explore whether the pairwise distance between two miRNA biomarkers for the same disease is smaller than that between any two arbitrary miRNAs. A single disease may have multiple miRNA biomarkers. Because this results in many pairwise distances among the biomarkers for the same disease, we can use the average of the pairwise distances to assess the similarity level of miRNA biomarkers for a given disease.
To examine whether miRNA biomarkers for the same disease exhibit a high level of similarity, distributions of the pairwise distance of two miRNAs have been derived, enabling the estimation of the percentile rank corresponding to a specified pairwise distance value [20,21]. However, since the average pairwise distance is typically used and the existing distributions are based on individual pairwise distances rather than their averages, a modified distribution for the average pairwise distance may be more appropriate for estimating the corresponding percentile rank. When a disease has only two miRNA biomarkers, there is only one pairwise distance. In contrast, for n miRNA biomarkers, there are n ( n 1 ) / 2 pairwise distances. Therefore, the distribution of the average pairwise distance should depend on the number of miRNA biomarkers. In this study, we provide percentiles of the distribution of average pairwise distances for miRNA counts ranging from 10 to 200, which can help determine whether the miRNA biomarkers of a disease exhibit high similarity.
Previous studies have also shown that miRNA biomarkers of the same disease tend to have a smaller average pairwise distance compared to the overall average pairwise distance of all miRNAs [20,21]. In this study, we used the proposed method to examine the average pairwise distances of 51 diseases and several associated diseases. The results demonstrate that the average pairwise distances of miRNA biomarkers for most of these diseases have low percentile ranks, indicating a high similarity of miRNA biomarkers.

2. Results

2.1. Percentile Table

For a given disease, if there are n miRNA biomarkers associated with it, then the number of pairwise distances among these n miRNAs is
C 2 n = ( n × n 1 ) / 2
The average of these pairwise distances can be calculated using the Jukes and Cantor (JC) one-parameter model. To examine whether the average pairwise distance of a disease has a low percentile rank, the percentiles corresponding to percentile ranks 1, 5, 10, 20, …, and 90 are obtained for different numbers of miRNAs using the miRBase dataset containing 1917 human miRNAs [22,23] (see Section 5 for details).
Table 1 presents the percentiles for n = 10 ,   20 ,   40 ,   60 ,   80 , and 100 . Table S1 provides the results for additional values of n from 120 to 200. For example, if a disease has 10 miRNA biomarkers and the average of their pairwise distances is close to 0.9, then from the column corresponding to n = 10 in Table 1, we observe that this value is close to 0.910, which corresponds to the 5th percentile. This indicates that the average pairwise distance has a low percentile rank, suggesting a high similarity among the miRNA biomarkers.

2.2. Applications

In this section, we first use the derived percentiles to examine the percentile rank of the average pairwise distance for single diseases, and then examine it for certain associated diseases.

2.2.1. Single Disease

In this study, we use the Human miRNA Disease Database (HMDD) [24] to find miRNA biomarkers for various conditions. HMDD is a database that curated experiment-supported evidence for human miRNA and disease associations. For a single disease, after searching the miRNA biomarkers in HMDD, the JC one-parameter model is used to calculate the pairwise distances for these miRNAs. Then we can obtain the average of these pairwise distances and find the corresponding percentile ranks in Table 1 or Table S1. The range of percentile rank can be from 0 to 100. An average pairwise distance corresponding to a low percentile rank means that this distance is small compared to the other distances.
Note that Table 1 and Table S1 do not provide the percentiles for any number of miRNA biomarkers. For the number not in these tables, the nearest number in these tables can be used to estimate the percentile of the average pairwise distances.
The miRNA biomarkers for 51 conditions were obtained from HMDD (Table S2). A small proportion of these miRNAs could not be found in miRBase. Therefore, only miRNAs listed in miRBase or their closely related counterparts were considered. The percentiles of the average pairwise distances for these 51 conditions are provided in Table 2.
In Table 2, the first disease is anovulation, which has 10 miRNA biomarkers. There are 45 pairwise distances among these miRNAs, and the average of these distances is 0.8669. Referring to the column corresponding to n = 10 in Table 1, 0.8669 is smaller than 0.883, which corresponds to the 1st percentile. Consequently, the percentile rank of the average pairwise distance for anovulation is less than 1.
If 10 miRNAs are randomly selected from the dataset, rather than using these disease-associated biomarkers, the probability of obtaining an average pairwise distance as low as 0.8669 is very small, as indicated by the column corresponding to n = 10 in Table 1. This result suggests that the miRNA biomarkers associated with anovulation exhibit a high level of sequence similarity.
The results in Table 2 show that the percentile ranks of the average pairwise distances for most of the diseases are less than 1, indicating a high level of pre-miRNA sequence similarity among the miRNA biomarkers for these conditions. Figure 1 displays the count distribution of the percentile ranks across the 51 diseases. Notably, 38 diseases have percentile ranks below 1, and the distribution is strongly right-skewed. Furthermore, most of the percentile ranks of the average pairwise distances are small. Therefore, we conclude that the average pairwise distance among miRNA biomarkers for a single disease is likely to be smaller than that among randomly selected miRNAs.
Several conditions in Table 2 have percentile ranks higher than 10. The relatively high percentile ranks observed for brucellosis, myopia, necrosis, and seizures may reflect the heterogeneous or non-specific nature of these conditions. Unlike more well-defined diseases, these conditions involve diverse biological processes or clinical manifestations, which may result in more variable and less similar sets of associated miRNA biomarkers.
In addition, the number of miRNA biomarkers for these four conditions is smaller than that for most other conditions in Table 2. Since the percentiles in Table 1 were derived for different numbers of miRNA biomarkers, the results remain valid even when the number of miRNA biomarkers is small.

2.2.2. Associated Diseases

In addition to investigating the similarity level of miRNA biomarkers for a single disease, the similarity levels of miRNA biomarkers for associated diseases are also explored, including diabetes mellitus (DM) and glaucoma, DM and Parkinson’s disease (PD), migraines and depression, and stroke and hypertension.
DM is an endocrine disorder caused either by insufficient insulin production by the pancreas or by the body’s inadequate response to insulin [25,26]. It can lead to various complications, including retinopathy, nephropathy, and peripheral neuropathy. The increasing global burden of DM has become a significant public health concern, placing immense and often unsustainable pressure on individuals, caregivers, healthcare systems, and society as a whole [27]. DM, longer diabetes duration, elevated intraocular pressure, and higher fasting glucose levels are significantly associated with an increased risk of glaucoma and are considered among its leading risk factors [28,29,30]. The miRNA biomarkers of DM and those of glaucoma are obtained from HMDD, and the average pairwise distance of these miRNAs is calculated. A total of 107 miRNA biomarkers for DM and glaucoma were obtained from HMDD. The average value of the pairwise distance of these 107 miRNAs is 0.93343. Since n = 107 in this case, by checking the percentiles for n = 100 and n = 110 in Table 1 and Table S1, this average pairwise distance falls within the bottom 1%, indicating a percentile rank below 1. This suggests that the miRNA biomarkers of these two associated diseases have high similarity.
PD is a neurological disorder that mainly affects older adults, characterized by a combination of motor and non-motor difficulties [31]. This progressive neurodegenerative condition arises from the deterioration of dopamine-producing neurons in a specific region of the brain known as the substantia nigra pars compacta [32]. DM may increase the risk of developing a Parkinson-like condition, and when it coexists with PD, it can lead to a more severe form of the disease [33]. The miRNA biomarkers of DM and those of PD are obtained from HMDD, and the average pairwise distance of these miRNAs is calculated. A total of 100 miRNA biomarkers for DM and PD were obtained from HMDD. The average pairwise distance of these 100 miRNAs is 0.89359. Since n = 100 in this case, by checking the percentiles for n = 100 in Table 1, the percentile rank of this average pairwise distance is less than 1. This suggests that the miRNA biomarkers of these two associated diseases have high similarity.
Migraines frequently co-occur with depression and are commonly seen in clinical settings [34,35]. This comorbidity can result in more severe conditions, accompanied by additional symptoms and prolonged treatment periods [36]. Individuals with migraine are 2.5 times more likely to develop a depressive disorder, with the risk being even greater among those with chronic migraine or migraine with aura. This association is bidirectional, as depression also increases the likelihood of earlier onset and greater severity of migraine, contributing to a higher risk of chronic migraine and, in turn, greater healthcare costs compared to migraine alone [37]. There are a total of 44 miRNA biomarkers for migraine disorders and depression obtained from HMDD. The average value of the pairwise distance of these 44 miRNAs is 0.948. Since n = 44 in this case, by checking the percentiles for n = 40 and 50 in Table S1, the percentile rank of this average pairwise distance is less than 5. This suggests that the miRNA biomarkers of these two associated diseases also exhibit high similarity, although the similarity level is not as high as in the two cases above.
Hypertension is the most prevalent risk factor for stroke [38,39]. There are a total of 76 miRNA biomarkers for stroke and hypertension obtained from HMDD. The average value of the pairwise distance of these 76 miRNAs is 0.905. Since n = 76 in this case, by checking the percentiles for n = 70 and 80 in Table S1, the percentile rank of this average pairwise distance is less than 1. This suggests that the miRNA biomarkers of these two associated diseases have high similarity.
Table 3 summarizes the total number of miRNA biomarkers, the average pairwise distances, and the corresponding percentile ranks for these four cases. The details of these miRNA biomarkers are provided in Table S3. Figure 2 provides the number of miRNA biomarkers for the four cases.

3. Discussion

miRNAs are evolutionarily conserved across species, meaning that many miRNAs have similar sequences and functions in different organisms, from plants to animals [5,8,40,41]. This conservation means that the structure of the miRNA, which includes its sequence, remains similar in different organisms, highlighting its important regulatory roles in gene expression, development, and disease. In general, miRNAs with highly similar sequences tend to have similar functions because if two miRNAs have very similar nucleotide sequences, they are likely to target similar mRNAs and regulate similar genes, leading to similar biological effects. As a result, the similarity degree of miRNA biomarkers between two diseases might reflect the association between these two diseases [20,21]. Additionally, various similarity measurements in miRNAs for applications in miRNA-disease association predictions have been studied [42].
This study uses stem-loop (pre-miRNA) sequences to measure the similarity between miRNAs. Although mature miRNAs are the functional molecules involved in gene regulation, using pre-miRNA sequences to assess miRNA similarity is also reasonable. Various studies have demonstrated that nucleotides within the loop region of pre-miRNAs can control the production and activity of mature miRNAs. For example, the two mature miRNAs miR-181a-1 and miR-181c both belong to the same miR-181 family and only have a one-nucleotide difference (Table 4), but have different functions.
Ectopic expression of miR-181a-1, but not miR-181c, promotes the development of CD4 and CD8 double-positive T cells from thymic progenitor cells [43]. The differing functions of miR-181a-1 and miR-181c are primarily driven by the unique nucleotides in their pre-miRNA loop regions, rather than by the single-nucleotide variation in their mature sequences [43]. Another study demonstrated that regulatory signals embedded in the loop nucleotides of pri- and pre-miRNAs modulate the function of both pri-miRNAs and mature let-7 by affecting the formation of pri-miRNA-target complexes and ensuring the accuracy of mature miRNA seed region generation [44]. Therefore, it is likely that pre-miRNA sequences and structures might also play roles in regulating miRNA expression, localization, or stability.
In addition, the JC model used in this study has been applied in molecular evolution and provides a simple correction for multiple substitutions at the same site over time [45]. It is especially useful when comparing homologous nucleotide sequences, like miRNA precursors, to estimate evolutionary divergence. It has also been used to estimate the nucleotide divergence of miRNAs in humans and animals. For Drosophila miRNAs, nucleotide divergence was estimated by calculating the mean JC distance within each orthologous miRNA group [46]. The association between Coronavirus disease 2019 (COVID-19) and anti-NMDA receptor encephalitis has been investigated by using the JC distance to measure the similarity of the pre-miRNA biomarker sequence [47]. Pre-miRNAs are often conserved among functionally similar miRNAs, and the JC model was used to measure the miRNA sequence distance to reflect functional similarity or evolutionary relatedness of diploid Oryza species [48].
Therefore, in this study, we apply the JC model to calculate the pairwise distance of miRNAs and derive the percentiles for the average pairwise distance distributions corresponding to different miRNA counts. The average pairwise distances of miRNA biomarkers for 51 diseases were calculated based on their pre-miRNA sequences and the JC model, as well as four associated disease cases. Most of the corresponding percentile ranks of these 51 values are small, with many below 1 and the maximum equal to 45. These results indicate that the pre-miRNA sequences may influence the function of mature miRNAs. However, the present analysis focuses only on sequence similarity and does not examine functional effects.

Limitation

In this study, the results show that, for many diseases, the average pairwise distances among pre-miRNAs corresponding to disease-associated miRNAs tend to fall within low percentile ranges under the proposed reference distribution. This suggests that disease-associated miRNAs exhibit sequence-level clustering relative to random expectation. However, several limitations should be acknowledged.
The observed low percentile outcomes may reflect precursor-level sequence similarity; however, they should not be directly interpreted as evidence of functional similarity. Although the use of pre-miRNA sequences provides a consistent and well-annotated basis for sequence comparison, the relationship between global precursor sequence similarity and biological function remains indirect. Sequence similarity may reflect evolutionary or structural relatedness but does not necessarily imply shared regulatory roles or disease mechanisms. miRNA function is influenced by multiple factors, including the seed region, target gene networks, expression profiles, arm selection (5p or 3p), and cellular context. While certain regions of pre-miRNAs, such as loop structures, can affect processing and maturation, functional similarity is more often associated with mature miRNA sequences, particularly the seed region, as well as downstream target interactions and regulatory pathways. Therefore, the present framework primarily captures sequence-level similarity rather than functional equivalence, and any biological interpretation should be made with caution.
In addition, several factors may contribute to the observed clustering patterns. These include conserved miRNA families, evolutionary constraints on pre-miRNA structure, and database-related biases, such as the overrepresentation of well-studied miRNAs. The current analysis does not distinguish among these factors, which introduces uncertainty in interpretation. Therefore, the results should be interpreted carefully.
Another limitation lies in the use of the JC one-parameter model for distance calculation. While the JC model provides a simple and widely used approach for estimating nucleotide substitution distances, it assumes equal base frequencies and substitution rates across all nucleotides. This assumption may not adequately reflect the structural and functional constraints of pre-miRNAs, which are characterized by stem-loop secondary structures and non-uniform evolutionary pressures. Future studies may explore alternative distance measures.
Additionally, all miRNA biomarkers associated with a disease are treated equally in the current framework. In reality, miRNAs may differ substantially in their biological importance, expression levels, and regulatory impact. Averaging pairwise distances provides a global summary of sequence similarity but may obscure heterogeneity within miRNA sets. Incorporating weighting schemes based on functional relevance or expression data represents a potential extension of this work. Moreover, reliance on HMDD without incorporating other databases may introduce bias. Furthermore, the term “biomarker,” as used in this study, encompasses a range of evidence types, including expression associations and experimentally validated functional roles, which may vary in reliability and biological significance.
Despite these limitations, the percentile framework proposed in this study provides a useful reference for assessing whether a set of miRNAs exhibits unusually high or low sequence similarity. It may serve as a complementary tool for exploratory analysis.

4. Materials and Methods

4.1. Materials

miRBase is a database that compiles published miRNA sequences [22]. It contains miRNA nucleotide sequences from 271 organisms, including both animals and plants. Specifically, miRBase provides access to 1917 human miRNA sequences. The 1917 miRNA stem-loop sequences (pre-miRNA sequences) are used in this study for pre-miRNA distance calculation. Pairwise distances among these miRNAs, totaling 1,836,486, were computed using the JC one-parameter model, a widely used method for estimating nucleotide sequence distances. The pairwise distance calculations were performed using the Bioinformatics Toolbox in MATLAB (version R2021a, MathWorks, Natick, MA, USA) [49].

4.2. Nucleotide Substitution Models

The JC one-parameter model is reviewed in this subsection. The JC one-parameter model is a frequently used model assuming that substitutions occur with equal probability among the four nucleotide types, A, T, C, and G [50]. Let K denote the number of substitutions per site since the time of divergence between two sequences with length L . Let X denote the number of different sites between these two sequences. Under the JC one-parameter model, we have
K 1 = 3 4 l n ( 1 4 3 p ^ )
where p ^ = X / L is the observed proportion of different nucleotides between two sequences. The value K 1 is used as a distance of two miRNA sequences in this study. An approximated estimator for the sampling variance of K 1 is [51]
V K 1 = p ^ p ^ 2 L 1 4 3 p ^ 2

4.3. Pairwise Distance

The JC one-parameter model was adopted to calculate the sequence distance between two miRNAs. There are 1917 human miRNA sequences used in this study. Consequently, a total of C 2 1917 = ( 1917 1916 ) / 2 = 1,836,486 pairwise distances can be calculated based on the 1917 miRNA sequences. These pairwise distances are directly used to construct the distance distribution of pre-miRNA sequences in a previous study [21].
Using the constructed distance distribution, the percentile in which this calculated average falls can be determined. If the percentile is low, it indicates that these miRNA sequences have high similarity. This method is used to explore the association between two diseases by calculating the average distance between their miRNA biomarkers. If the miRNA biomarkers of two diseases have high similarity, the association between these two diseases cannot be excluded. While this existing distance distribution can provide a useful approach to examining the level of similarity of miRNAs, it was developed to examine a single distance. However, for the average pairwise distances of n miRNAs, a more appropriate distance distribution needs to be constructed depending on the number n . Therefore, in this paper, the percentiles of the average pairwise distance distributions corresponding to different miRNA counts are established.

4.4. Percentiles of Average Pairwise Distance

This section provides an algorithm to calculate the percentiles of the average pairwise distance corresponding to different numbers of miRNAs. Let F n denote the distribution of the average pairwise distance for n miRNAs. The algorithm for finding the percentiles of the sampling distribution for a fixed n is provided below. The MATLAB codes are provided as Supplementary Materials.
In this analysis, all pre-miRNA sequences were obtained from miRBase and analyzed at full precursor length without additional preprocessing. miRNA identifiers from HMDD were mapped to corresponding entries in miRBase to ensure consistency; unmatched entries were either excluded or replaced with closely related counterparts. All sequences were obtained from a single version of miRBase. In Step 1, pairwise distances were calculated directly using MATLAB’s seqpdist function under the JC model. Given the large number of human pre-miRNAs and the extensive pairwise computations required, explicit sequence alignment is computationally challenging. Accordingly, the proposed method provides a simplified measure of similarity based on direct comparison of nucleotide sequences, rather than an alignment-optimized analysis. A limitation of this approach is that pre-miRNA sequences vary in length and secondary structure; therefore, alignment-based methods may provide more biologically meaningful similarity estimates and should be considered in future work.
By using Algorithm 1 with k = 10,000 , the ith percentiles can be obtained for different values of n. Note that other methods, such as empirical cumulative distribution and the kernel density estimation method, can be used in Step 4 to estimate the percentiles of the average pairwise distance distribution [20].
Algorithm 1. Procedure for estimating the percentiles of the average pairwise distance distribution for a given number of miRNAs.
Step 1. Among the 1917 human miRNAs, randomly select n miRNAs. Calculate the C 2 n pairwise distances, and average the C 2 n distances.
Step 2. Repeat Step 1 for k times for a large k, say 10,000. Then k average values can be obtained.
Step 3. Rank the k average values from the smallest to the largest, say x 1 , , x k .
Step 4. The ith percentile of the average pairwise distance distribution is x ( [ k i 100 ] ) ,
where [ ] denotes the floor function, indicating the greatest integer less than or equal to the given value.

5. Conclusions

MiRNAs are involved in various disease mechanisms and have potential as valuable biomarkers. Sequence similarity among miRNA biomarkers has been used to explore associations between diseases. Previous studies have characterized the distribution of pairwise distances between miRNAs, which can be used to measure the similarity between two miRNAs. However, because a single disease is often associated with multiple miRNA biomarkers, it may be more appropriate to assess the percentile rank of the average pairwise distance among all biomarker pairs, rather than that of a single pair.
Therefore, this study establishes percentile distributions of average pairwise distances for varying numbers of miRNAs. Using this framework, the percentile ranks of the average pairwise distances were evaluated for 51 diseases and several groups of related diseases. The results indicate that, for many diseases, the average pairwise distances of their associated miRNAs tend to fall within low percentile ranges under the proposed reference distribution.
These findings suggest that disease-associated miRNAs often exhibit sequence-level clustering at the precursor level relative to random expectation. However, this observation should be interpreted with caution. The proposed framework captures global sequence similarity and does not directly reflect functional similarity, regulatory mechanisms, or disease causality, which depend on additional factors such as seed sequences, target interactions, and expression context. In addition, integrating sequence-based similarity with functional and biological information, including shared target genes, pathway enrichment, and expression profiles, remains an important direction for future work.
Overall, the proposed method provides a descriptive statistical framework for assessing sequence-level similarity patterns among miRNA sets and may serve as a complementary approach in miRNA-related bioinformatics research.

Supplementary Materials

The following supporting information can be downloaded at: https://www.mdpi.com/article/10.3390/ijms27093750/s1.

Author Contributions

H.W. conceived the presented idea, Y.-S.J. and W.-C.H. collected the data, Y.-S.J. and W.-C.H. analyzed the data, and H.W. wrote the paper. All authors have read and agreed to the published version of the manuscript.

Funding

This research was funded by National Science and Technology Council 113-2118-M-A49-004-MY2, Taiwan.

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

The original contributions presented in this study are included in the article/Supplementary Material. Further inquiries can be directed to the corresponding author.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Wang, H. Anti-NMDA Receptor Encephalitis, Human Papillomavirus, and microRNA. Curr. Med. Chem. 2025, 32, 771–787. [Google Scholar] [CrossRef]
  2. Wang, H. Predicting cancer-related MiRNAs using expression profiles in tumor tissue. Curr. Pharm. Biotechnol. 2014, 15, 438–444. [Google Scholar] [CrossRef]
  3. Väisänen, M.; Siukosaari, P.; Tjäderhane, L. How epigenetics and miRNA affect gene expression in dental pulp inflammation: A narrative review. Int. Endod. J. 2025, 58, 833–847. [Google Scholar] [CrossRef]
  4. Lee, R.C.; Feinbaum, R.L.; Ambros, V. The C. elegans heterochronic gene lin-4 encodes small RNAs with antisense complementarity to lin-14. Cell 1993, 75, 843–854. [Google Scholar] [CrossRef]
  5. Pasquinelli, A.E.; Reinhart, B.J.; Slack, F.; Martindale, M.Q.; Kuroda, M.I.; Maller, B.; Hayward, D.C.; Ball, E.E.; Degnan, B.; Muller, P.; et al. Conservation of the sequence and temporal expression of let-7 heterochronic regulatory RNA. Nature 2000, 408, 86–89. [Google Scholar] [CrossRef]
  6. Treiber, T.; Treiber, N.; Meister, G. Regulation of microRNA biogenesis and its crosstalk with other cellular pathways. Nat. Rev. Mol. Cell Biol. 2019, 20, 5–20, Correction in Nat. Rev. Mol. Cell Biol. 2019, 20, 321. [Google Scholar] [CrossRef] [PubMed]
  7. De Rie, D.; Abugessaisa, I.; Alam, T.; Arner, E.; Arner, P.; Ashoor, H.; Åström, G.; Babina, M.; Bertin, N.; Burroughs, A.M. An integrated expression atlas of miRNAs and their promoters in human and mouse. Nat. Biotechnol. 2017, 35, 872–878. [Google Scholar] [CrossRef] [PubMed]
  8. O’Brien, J.; Hayder, H.; Zayed, Y.; Peng, C. Overview of microRNA biogenesis, mechanisms of actions, and circulation. Front. Endocrinol. 2018, 9, 402. [Google Scholar] [CrossRef] [PubMed]
  9. Kim, V.N.; Han, J.; Siomi, M.C. Biogenesis of small RNAs in animals. Nat. Rev. Mol. Cell Biol. 2009, 10, 126–139. [Google Scholar] [CrossRef]
  10. Hatfield, S.; Ruohola-Baker, H. microRNA and stem cell function. Cell Tissue Res. 2008, 331, 57–66. [Google Scholar] [CrossRef]
  11. Sevcikova, A.; Fridrichova, I.; Nikolaieva, N.; Kalinkova, L.; Omelka, R.; Martiniakova, M.; Ciernikova, S. Clinical significance of microRNAs in hematologic malignancies and hematopoietic stem cell transplantation. Cancers 2023, 15, 2658. [Google Scholar] [CrossRef]
  12. Qian, H.; Cui, N.; Zhou, Q.; Zhang, S. Identification of miRNA biomarkers for stomach adenocarcinoma. BMC Bioinform. 2022, 23, 181. [Google Scholar] [CrossRef]
  13. Petrescu, G.E.; Sabo, A.A.; Torsin, L.I.; Calin, G.A.; Dragomir, M.P. MicroRNA based theranostics for brain cancer: Basic principles. J. Exp. Clin. Cancer Res. 2019, 38, 231. [Google Scholar] [CrossRef]
  14. Ishiguro, H.; Kimura, M.; Takeyama, H. Role of microRNAs in gastric cancer. World J. Gastroenterol. WJG 2014, 20, 5694. [Google Scholar] [CrossRef]
  15. Hussen, B.M.; Hidayat, H.J.; Salihi, A.; Sabir, D.K.; Taheri, M.; Ghafouri-Fard, S. MicroRNA: A signature for cancer progression. Biomed. Pharmacother. 2021, 138, 111528. [Google Scholar] [CrossRef]
  16. Wang, H.; Taguchi, Y.; Liu, X. miRNAs and neurological diseases. Front. Neurol. 2021, 12, 662373. [Google Scholar] [CrossRef]
  17. Nishiguchi, T.; Imanishi, T.; Akasaka, T. MicroRNAs and cardiovascular diseases. BioMed Res. Int. 2015, 2015, 682857. [Google Scholar] [CrossRef] [PubMed]
  18. Franczyk, B.; Gluba-Brzózka, A.; Olszewski, R.; Parolczyk, M.; Rysz-Górzyńska, M.; Rysz, J. miRNA biomarkers in renal disease. Int. Urol. Nephrol. 2022, 54, 575–588. [Google Scholar] [CrossRef] [PubMed]
  19. Hicks, S.D.; Onks, C.; Kim, R.Y.; Zhen, K.J.; Loeffert, J.; Loeffert, A.C.; Olympia, R.P.; Fedorchak, G.; DeVita, S.; Gagnon, Z. Refinement of saliva microRNA biomarkers for sports-related concussion. J. Sport Health Sci. 2023, 12, 369–378. [Google Scholar] [CrossRef]
  20. Wang, H. The distance distribution of human microRNAs in MirGeneDB database. Sci. Rep. 2022, 12, 17696. [Google Scholar] [CrossRef] [PubMed]
  21. Wang, H.; Ho, C. The Human Pre-miRNA Distance Distribution for Exploring Disease Association. Int. J. Mol. Sci. 2023, 24, 1009. [Google Scholar] [CrossRef] [PubMed]
  22. Kozomara, A.; Birgaoanu, M.; Griffiths-Jones, S. miRBase: From microRNA sequences to function. Nucleic Acids Res. 2019, 47, D155–D162. [Google Scholar] [CrossRef]
  23. Huang, L.; Zhang, L.; Chen, X. Updated review of advances in microRNAs and complex diseases: Experimental results, databases, webservers and data fusion. Brief. Bioinform. 2022, 23, bbac397. [Google Scholar] [CrossRef]
  24. Cui, C.; Zhong, B.; Fan, R.; Cui, Q. HMDD v4. 0: A database for experimentally supported human microRNA-disease associations. Nucleic Acids Res. 2024, 52, D1327–D1332. [Google Scholar] [CrossRef]
  25. Wang, H. MicroRNA, Diabetes Mellitus and Colorectal Cancer. Biomedicines 2020, 8, 530. [Google Scholar] [CrossRef]
  26. Jadon, A.S.; Kaushik, M.P.; Anitha, K.; Bhatt, S.; Bhadauriya, P.; Sharma, M. Types of diabetes mellitus, mechanism of insulin resistance and associated complications. In Biochemical Immunology of Diabetes and Associated Complications; Elsevier: Amsterdam, The Netherlands, 2024; pp. 1–18. [Google Scholar]
  27. Forouhi, N.G.; Wareham, N.J. Epidemiology of diabetes. Medicine 2019, 47, 22–27. [Google Scholar] [CrossRef]
  28. Zhao, D.; Cho, J.; Kim, M.H.; Friedman, D.S.; Guallar, E. Diabetes, fasting glucose, and the risk of glaucoma: A meta-analysis. Ophthalmology 2015, 122, 72–78. [Google Scholar] [CrossRef]
  29. AlDarrab, A.; Al Jarallah, O.J.; Al Balawi, H.B. Association of diabetes, fasting glucose, and the risk of glaucoma: A systematic review and meta-analysis. Eur. Rev. Med. Pharmacol. Sci. 2023, 27, 2419–2427. [Google Scholar] [CrossRef]
  30. Li, Y.; Mitchell, W.; Elze, T.; Zebardast, N. Association between diabetes, diabetic retinopathy, and glaucoma. Curr. Diabetes Rep. 2021, 21, 38. [Google Scholar] [CrossRef] [PubMed]
  31. Saini, N.; Singh, N.; Kaur, N.; Garg, S.; Kaur, M.; Kumar, A.; Verma, M.; Singh, K.; Sohal, H.S. Motor and non-motor symptoms, drugs, and their mode of action in Parkinson’s disease (PD): A review. Med. Chem. Res. 2024, 33, 580–599. [Google Scholar] [CrossRef]
  32. Church, F.C. Treatment options for motor and non-motor symptoms of Parkinson’s disease. Biomolecules 2021, 11, 612. [Google Scholar] [CrossRef]
  33. Pagano, G.; Polychronis, S.; Wilson, H.; Giordano, B.; Ferrara, N.; Niccolini, F.; Politis, M. Diabetes mellitus and Parkinson disease. Neurology 2018, 90, e1654–e1662. [Google Scholar] [CrossRef]
  34. Chen, Y.-H.; Wang, H. The association between migraine and depression based on miRNA biomarkers and cohort studies. Curr. Med. Chem. 2021, 28, 5648–5656. [Google Scholar] [CrossRef]
  35. McCracken, H.T.; Thaxter, L.Y.; Smitherman, T.A. Psychiatric comorbidities of migraine. Handb. Clin. Neurol. 2024, 199, 505–516. [Google Scholar]
  36. Zhang, Q.; Shao, A.; Jiang, Z.; Tsai, H.; Liu, W. The exploration of mechanisms of comorbidity between migraine and depression. J. Cell. Mol. Med. 2019, 23, 4505–4513. [Google Scholar] [CrossRef]
  37. Viudez-Martínez, A.; Torregrosa, A.B.; Navarrete, F.; García-Gutiérrez, M.S. Understanding the biological relationship between migraine and depression. Biomolecules 2024, 14, 163. [Google Scholar] [CrossRef] [PubMed]
  38. Wajngarten, M.; Silva, G.S. Hypertension and stroke: Update on treatment. Eur. Cardiol. Rev. 2019, 14, 111. [Google Scholar] [CrossRef] [PubMed]
  39. Li, A.-L.; Ji, Y.; Zhu, S.; Hu, Z.-H.; Xu, X.-J.; Wang, Y.-W.; Jian, X.-Z. Risk probability and influencing factors of stroke in followed-up hypertension patients. BMC Cardiovasc. Disord. 2022, 22, 328. [Google Scholar] [CrossRef] [PubMed]
  40. Pietrykowska, H.; Sierocka, I.; Zielezinski, A.; Alisha, A.; Carrasco-Sanchez, J.C.; Jarmolowski, A.; Karlowski, W.M.; Szweykowska-Kulinska, Z. Biogenesis, conservation, and function of miRNA in liverworts. J. Exp. Bot. 2022, 73, 4528–4545. [Google Scholar] [CrossRef]
  41. Wang, Y.; Tang, X.; Lu, J. Convergent and divergent evolution of microRNA-mediated regulation in metazoans. Biol. Rev. 2024, 99, 525–545. [Google Scholar] [CrossRef]
  42. Chen, H.; Guo, R.; Li, G.; Zhang, W.; Zhang, Z. Comparative analysis of similarity measurements in miRNAs with applications to miRNA-disease association predictions. BMC Bioinform. 2020, 21, 176. [Google Scholar] [CrossRef]
  43. Liu, G.; Min, H.; Yue, S.; Chen, C.-Z. Pre-miRNA loop nucleotides control the distinct activities of mir-181a-1 and mir-181c in early T cell development. PLoS ONE 2008, 3, e3592. [Google Scholar] [CrossRef]
  44. Yue, S.-B.; Trujillo, R.D.; Tang, Y.; O’Gorman, W.E.; Chen, C.-Z. Loop nucleotides control primary and mature miRNA function in target recognition and repression. RNA Biol. 2011, 8, 1115–1123. [Google Scholar] [CrossRef][Green Version]
  45. Nei, M.; Kumar, S. Molecular Evolution and Phylogenetics; Oxford University Press: Oxford, UK, 2000. [Google Scholar]
  46. Jovelin, R. Pleiotropic constraints, expression level, and the evolution of miRNA sequences. J. Mol. Evol. 2013, 77, 206–220. [Google Scholar] [CrossRef] [PubMed]
  47. Wang, H. COVID-19, anti-NMDA receptor encephalitis and microRNA. Front. Immunol. 2022, 13, 825103. [Google Scholar] [CrossRef]
  48. Ganie, S.A.; Debnath, A.B.; Gumi, A.M.; Mondal, T.K. Comprehensive survey and evolutionary analysis of genome-wide miRNA genes from ten diploid Oryza species. BMC Genom. 2017, 18, 711. [Google Scholar] [CrossRef] [PubMed]
  49. Singh, G.B. Processing Biological Sequences in MATLAB. In Fundamentals of Bioinformatics and Computational Biology: Methods and Exercises in MATLAB; Springer: Berlin/Heidelberg, Germany, 2025; pp. 37–67. [Google Scholar]
  50. Kimura, M.; Ohta, T. On the stochastic model for estimation of mutational distance between homologous proteins. J. Mol. Evolution 1972, 2, 87–90. [Google Scholar] [CrossRef] [PubMed]
  51. Wang, H.; Tzeng, Y.H.; Li, W.H. Improved variance estimators for one- and two-parameter models of nucleotide substitution. J. Theor. Biol. 2008, 254, 164–167. [Google Scholar] [CrossRef]
Figure 1. The count distribution of the percentile rank for the 51 diseases.
Figure 1. The count distribution of the percentile rank for the 51 diseases.
Ijms 27 03750 g001
Figure 2. (a) The numbers of miRNA biomarkers for diabetes mellitus and glaucoma; (b) the numbers of miRNA biomarkers for diabetes mellitus and Parkinson’s disease; (c) the numbers of miRNA biomarkers for migraine disorders and depression; (d) the numbers of miRNA biomarkers for stroke and hypertension.
Figure 2. (a) The numbers of miRNA biomarkers for diabetes mellitus and glaucoma; (b) the numbers of miRNA biomarkers for diabetes mellitus and Parkinson’s disease; (c) the numbers of miRNA biomarkers for migraine disorders and depression; (d) the numbers of miRNA biomarkers for stroke and hypertension.
Ijms 27 03750 g002
Table 1. The percentiles for n   =   10 ,   20 ,   40 ,   60 ,   80 ,   a n d   100 .
Table 1. The percentiles for n   =   10 ,   20 ,   40 ,   60 ,   80 ,   a n d   100 .
Percentile Rank n
1020406080100
10.8830.9180.9390.9480.9540.958
50.9100.9360.9520.9590.9630.967
100.9250.9470.9590.9650.9680.972
200.9360.9600.9690.9730.9760.978
300.9450.9700.9760.9800.9810.983
400.9530.9790.9830.9850.9860.987
500.9600.9880.9890.9900.9910.991
600.9660.9970.9960.9960.9960.996
700.9731.0071.0041.0021.0021.001
800.9791.0191.0131.0111.0101.009
900.9851.0401.0331.0281.0261.023
Table 2. Numbers of miRNA biomarkers, average pairwise distances, and percentiles of the average pairwise distances for 51 diseases.
Table 2. Numbers of miRNA biomarkers, average pairwise distances, and percentiles of the average pairwise distances for 51 diseases.
DiseaseNumber of miRNA
Biomarkers
Average Pairwise DistancesPercentile Rank
Anovulation100.8669<1
Brucellosis90.925810
Cardiomegaly1030.909<1
Cataract340.8918<1
Choriocarcinoma350.8776<1
Coronary Occlusion110.92095
Gout290.76<1
Infertility, Female70.88581
Infertility, Male520.95161
Kidney Calculi110.91015
Laryngeal Neoplasms1120.9147<1
Mastitis240.8918<1
Melanoma3020.9163<1
Myocarditis330.9305<1
Myopia70.982645
Necrosis170.949710
Nerve Degeneration390.9124<1
Neuroendocrine Tumors180.93895
Norrie Disease410.9362<1
Obesity1640.9238<1
Ovarian epithelial cancer1080.9323<1
Pancreatitis520.9206<1
Papillomaviridae220.9104<1
Parkinson Disease1580.9123<1
Pediatric Obesity220.8988<1
Pre Eclampsia2020.9215<1
Precancerous Conditions360.9309<1
Prediabetic State230.9069<1
Premature Birth540.8978<1
Pterygium320.95175
Radiation Injuries150.8953<1
Rectal Neoplasms830.9188<1
Renal Insufficiency, Chronic370.8893<1
Reperfusion Injury830.9076<1
Sarcoma300.9244<1
Schizophrenia660.9198<1
Seizures110.934210
Sepsis1590.9231<1
Skin Neoplasms200.902<1
Solitary Fibrous Tumor, Pleural430.9156<1
Spinal Cord Injuries1250.9124<1
Stroke860.9299<1
Subarachnoid Hemorrhage380.8941<1
Syphilis200.8884<1
Teratoma80.86779<1
Testicular Diseases140.91515
Testicular Germ Cell
Tumor
240.9261
Thrombocytopenia280.9168<1
Thrombosis190.92891
Thyroid Carcinoma, Anaplastic290.898<1
Thyroid Neoplasms 1720.9123<1
Table 3. The total number of miRNA biomarkers, the average pairwise distance, and the corresponding percentile.
Table 3. The total number of miRNA biomarkers, the average pairwise distance, and the corresponding percentile.
Associated DiseasesTotal Number of miRNA Biomarkers (n)Average Pairwise
Distance
Percentile Rank
Diabetes Mellitus, Glaucoma1070.93343<1
Diabetes Mellitus, Parkinson’s disease1000.89359<1
Migraine Disorders, Depression440.94824<5
Stroke, Hypertension760.90521<1
Table 4. The sequences of mature miR-181a-1 and miR-181c.
Table 4. The sequences of mature miR-181a-1 and miR-181c.
miRNAMature Sequence
has-miR-181a-1AACAUUCAACGCUGUCGGUGAGU
has-miR-181cAACAUUCAACCUGUCGGUGAGU
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Wang, H.; Jiang, Y.-S.; Huang, W.-C. The Distribution of Average Pairwise Distances Among Human Pre-miRNAs for Disease Association Analysis. Int. J. Mol. Sci. 2026, 27, 3750. https://doi.org/10.3390/ijms27093750

AMA Style

Wang H, Jiang Y-S, Huang W-C. The Distribution of Average Pairwise Distances Among Human Pre-miRNAs for Disease Association Analysis. International Journal of Molecular Sciences. 2026; 27(9):3750. https://doi.org/10.3390/ijms27093750

Chicago/Turabian Style

Wang, Hsiuying, You-Shan Jiang, and Wei-Ching Huang. 2026. "The Distribution of Average Pairwise Distances Among Human Pre-miRNAs for Disease Association Analysis" International Journal of Molecular Sciences 27, no. 9: 3750. https://doi.org/10.3390/ijms27093750

APA Style

Wang, H., Jiang, Y.-S., & Huang, W.-C. (2026). The Distribution of Average Pairwise Distances Among Human Pre-miRNAs for Disease Association Analysis. International Journal of Molecular Sciences, 27(9), 3750. https://doi.org/10.3390/ijms27093750

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop