Leveraging Fst and Genetic Distance to Optimize Reference Sets for Enhanced Cross-Population Genomic Prediction
Simple Summary
Abstract
1. Introduction
2. Materials and Methods
2.1. Data Simulation
- (1)
- Population A (PopA) was modeled to represent a linkage disequilibrium (LD) pattern that starts from a stable state and continues to expand. The historical population (HP) was initially set at 200 individuals, with this size sustained for 1000 generations. Thereafter, the population was scaled up to 1000 individuals (100 males and 900 females) over the next 95 generations, thereby establishing LD state and reaching mutation-drift equilibrium. From the terminal generation of the HP, 250 males and 3000 females were randomly chosen to initiate population expansion, with random mating performed for 10 generations (one offspring per female per generation). Subsequently, 50 males and 2000 females were randomly selected from the expanded cohort and simulated over the latest 100 generations, maintaining the same offspring production rate (one per female per generation). Selection was based on phenotypic performance and breeding values (EBVs) estimated via Best Linear Unbiased Prediction (BLUP), with replacement rates of 60% for bulls and 30% for cows, and a phenotypic record missingness rate of 5%.
- (2)
- Population B (PopB) was simulated to reflect LD pattern that initially contracts and then expands. The initial population comprised 500 individuals, which underwent a size reduction to 200 individuals over 1000 generations and was subsequently expanded to 1000 individuals (100 males and 900 females) during the next 95 generations, for the purpose of establishing the initial LD pattern and reaching mutation-drift equilibrium. From the terminal generation of the HP, 250 males and 2550 females were randomly chosen to initiate population expansion, with random mating performed for 10 generations (one offspring per female per generation). Subsequently, 55 males and 2100 females were randomly selected from the expanded cohort and simulated over the latest 100 generations, maintaining the same offspring production rate (one per female per generation). Selection was based on phenotypic performance and EBV estimated BLUP, with replacement rates of 50% for bulls and 30% for cows, and a phenotypic record missingness rate of 5%.
- (3)
- Population C (PopC) was modeled to simulate LD pattern that first expands and then contracts. The initial population comprised 1000 individuals, which underwent a size reduction to 200 individuals over 1000 generations and was subsequently expanded to 1000 individuals (100 males and 100 females) during the next 95 generations, for the purpose of establishing the initial LD pattern and reaching mutation-drift equilibrium. From the terminal generation of the HP, 250 males and 3000 females were randomly chosen to initiate population expansion, with random mating performed for 10 generations (one offspring per female per generation). Subsequently, 50 males and 2200 females were randomly selected from the expanded cohort and simulated over the latest 100 generations, maintaining the same offspring production rate (one per female per generation). Selection was based on phenotypic performance and EBV estimated BLUP, with replacement rates of 50% for bulls and 20% for cows, and a phenotypic record missingness rate of 5%.
2.1.1. Genome
2.1.2. Genomic Evaluation of Breeding Value
2.2. Scenarios
- Screening of highly differentiated SNPs: First, pairwise fixation index (Fst) values between the target population (PopA) and candidate reference populations (PopB, PopC) were calculated. Using Fst > 0.1 as the threshold, highly differentiated SNPs between populations were screened to provide a core marker basis for subsequent genetic similarity assessment;
- Calculation of Euclidean genetic distance: An individual genetic similarity matrix was constructed based on the genotypic data of the screened highly differentiated SNPs. The genetic distance between each individual in PopB/PopC and individuals in PopA was defined as the Euclidean distance between genotype vectors, quantifying the genetic similarity at the individual level across populations;
- Screening of highly similar individuals: Individuals in PopB and PopC were ranked by Euclidean distance in ascending order. The top 10%, 15%, and 20% of individuals with the highest genetic similarity to PopA were selected from PopB and PopC, respectively, forming subsets of highly similar individuals with different proportions;
- Construction of cross-population reference sets: A stratified mixing strategy was adopted. The subsets of highly similar individuals with different proportions screened from PopB and PopC were combined with the original reference individuals of PopA, respectively, to generate 6 cross-population reference sets specific combinations:
- (1)
- A + 10%B: Core population A supplemented with individuals showing the top 10% genetic distance from population B;
- (2)
- A + 15%B: Core population A supplemented with individuals showing the top 15% genetic distance from population B;
- (3)
- A + 20%B: Core population A supplemented with individuals showing the top 20% genetic distance from population B;
- (4)
- A + 10%C: Core population A supplemented with individuals showing the top 10% genetic distance from population C;
- (5)
- A + 15%C: Core population A supplemented with individuals showing the top 15% genetic distance from population C;
- (6)
- A + 20%C: Core population A supplemented with individuals showing the top 20% genetic distance from population C.
2.3. Model and Analysis
2.4. Genomic Prediction Accuracy
2.5. Two-Way Analysis of Variance (2-Way ANOVA)
3. Results
3.1. Population Genetic Characterization
3.2. Comparative Analysis of Prediction Accuracy of Single Population and Cross-Population Under Different Genetic Distance Levels Using Three Genomic Prediction Methods
3.3. Comparative Analysis of Accuracy Based on GBLUP, ssGBLUP, and wGBLUP Models Under Different Genetic Distance Levels
4. Discussion
5. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
Abbreviations
| GS | Genomic selection |
| TBV | True breeding value |
| EBV | Estimation of breeding value |
| BLUP | Best linear unbiased prediction |
| GEBV | Genomic estimated breeding value |
| SNP | Single-nucleotide polymorphism |
| MAF | Minor allele frequency |
| LD | Linkage disequilibrium |
| QTL | Quantitative trait loci |
| GBLUP | Genomic best linear unbiased prediction |
| ssGBLUP | Single-step genomic best linear unbiased prediction |
| wGBLUP | Weighted best linear unbiased prediction |
| GRM | Genomic relationship matrix |
| Fst | Fixation index |
References
- Hayes, B.J.; Corbet, N.J.; Allen, J.M.; Laing, A.R.; Fordyce, G.; Lyons, R.; McGowan, M.R.; Burns, B.M. Towards multi-breed genomic evaluations for female fertility of tropical beef cattle. J. Anim. Sci. 2019, 97, 55–62. [Google Scholar] [CrossRef]
- Lund, M.S.; Su, G.S.; Janss, L.; Guldbrandtsen, B.; Brøndum, R.F. Genomic evaluation of cattle in a multi-breed context. Livest. Sci. 2014, 166, 101–110. [Google Scholar] [CrossRef]
- Hosseini, S.; Foroutanifar, S.; Abdolmohammadi, A. Comparison of combined crossbred and purebred reference populations for genomic selection in small populations. Small Rumin. Res. 2020, 190, 106171. [Google Scholar] [CrossRef]
- Gao, H.; Christensen, O.F.; Madsen, P.; Nielsen, U.S.; Zhang, Y.; Lund, M.S.; Su, G. Comparison on genomic predictions using three GBLUP methods and two single-step blending methods in the Nordic Holstein population. Genet. Sel. Evol. 2012, 44, 8. [Google Scholar] [CrossRef] [PubMed]
- Silva, R.M.; Fragomeni, B.O.; Lourenco, D.A.; Magalhães, A.F.; Irano, N.; Carvalheiro, R.; Canesin, R.C.; Mercadante, M.E.; Boligon, A.A.; Baldi, F.S.; et al. Accuracies of genomic prediction of feed efficiency traits using different prediction and validation methods in an experimental Nelore cattle population1. J. Anim. Sci. 2016, 94, 3613–3623. [Google Scholar] [CrossRef] [PubMed]
- Putz, A.M.; Tiezzi, F.; Maltecca, C.; Gray, K.A.; Knauer, M.T. A comparison of accuracy validation methods for genomic and pedigree-based predictions of swine litter size traits using Large White and simulated data. J. Anim. Breed. Genet. 2018, 135, 5–13. [Google Scholar] [CrossRef]
- Londoño-Gil, M.; Cardona-Cifuentes, D.; Espigolan, R.; Peripolli, E.; Lôbo, R.B.; Pereira, A.S.C.; Aguilar, I.; Baldi, F. Genomic evaluation of commercial herds with different pedigree structures using the single-step genomic BLUP in Nelore cattle. Trop. Anim. Health Prod. 2023, 55, 95. [Google Scholar] [CrossRef]
- Wientjes, Y.C.; Veerkamp, R.F.; Bijma, P.; Bovenhuis, H.; Schrooten, C.; Calus, M.P. Empirical and deterministic accuracies of across-population genomic prediction. Genet. Sel. Evol. 2015, 47, 5. [Google Scholar] [CrossRef]
- Lourenco, D.A.L. Methods for genomic evaluation of a relatively small genotyped dairy population and effect of genotyped cow information in multiparity analyses. J. Dairy Sci. 2014, 97, 1742–1752. [Google Scholar] [CrossRef]
- Ghavi Hossein-zadeh, N. An overview of recent technological developments in bovine genomics. Vet. Anim. Sci. 2024, 25, 100382. [Google Scholar]
- Clasen, J.B.; Fikse, W.F.; Su, G.; Karaman, E. Multibreed genomic prediction using summary statistics and a breed-origin-of-alleles approach. Heredity 2023, 131, 33–42. [Google Scholar] [CrossRef]
- Hayes, B.J.; Copley, J.; Dodd, E.; Ross, E.M.; Speight, S.; Fordyce, G. Multi-breed genomic evaluation for tropical beef cattle when no pedigree information is available. Genet. Sel. Evol. 2023, 16, 5571. [Google Scholar] [CrossRef]
- Pszczola, M.; Strabel, T.; Mulder, H.A.; Calus, M.P. Reliability of direct genomic values for animals with different relationships within and to the reference population. J. Dairy Sci. 2012, 95, 389–400. [Google Scholar] [CrossRef] [PubMed]
- Yin, C.; Zhou, P.; Wang, Y.; Yin, Z.; Liu, Y. Using genomic selection to improve the accuracy of genomic prediction for multi-populations in pigs. Animal 2024, 18, 101062. [Google Scholar] [CrossRef]
- Chang, L.Y.; Toghiani, S.; Hay, E.H.; Aggrey, S.E.; Rekaya, R. A Weighted Genomic Relationship Matrix Based on Fixation Index (FST) Prioritized SNPs for Genomic Selection. Genes 2019, 10, 922. [Google Scholar] [CrossRef] [PubMed]
- De Roos, A.P.W.; Hayes, B.J.; Goddard, M.E. Reliability of Genomic Predictions Across Multiple Populations. Genetics 2009, 183, 1545–1553. [Google Scholar] [CrossRef]
- Ramstein, G.P.; Casler, M.D. Extensions of BLUP Models for Genomic Prediction in Heterogeneous Populations: Application in a Diverse Switchgrass Sample. G3 Genes Genomes Genet 2019, 9, 789–805. [Google Scholar] [CrossRef] [PubMed]
- Marjanovic, J.; Hulsegce, B.; Calus, M.P.L. Relatedness between numerically small Dutch Red dairy cattle populations and possibilities for multibreed genomic prediction. J. Dairy Sci. 2021, 104, 4498–4506. [Google Scholar] [CrossRef]
- Duarte, I.N.H.; Bessa, A.F.O.; Rola, L.D.; Genuíno, M.V.H.; Rocha, I.M.; Marcondes, C.R.; Regitano, L.C.A.; Munari, D.P.; Berry, D.P.; Buzanskas, M.E. Cross-population selection signatures in Canchim composite beef cattle. PLoS ONE 2022, 17, e0264279. [Google Scholar] [CrossRef]
- Toghiani, S.; Aggrey, S.E.; Rekaya, R. FST-Based Marker Prioritization Within Quantitative Trait Loci Regions and Its Impact on Genomic Selection Accuracy: Insights from a Simulation Study with High-Density Marker Panels for Bovines. Genes 2025, 16, 563. [Google Scholar] [CrossRef]
- Sargolzaei, M.; Schenkel, F.S. QMSim: A large-scale genome simulator for livestock. Bioinformatics 2009, 25, 680–681. [Google Scholar] [CrossRef] [PubMed]
- Brito, F.V.; Neto, J.B.; Sargolzaei, M.; Cobuci, J.A.; Schenkel, F.S. Accuracy of genomic selection in simulated populations mimicking the extent of linkage disequilibrium in beef cattle. BMC Genet. 2011, 12, 80. [Google Scholar] [CrossRef]
- Kambal, S.; Walsh, A.T.; Sivasankaran, S.K.; Adossa, N.A.; Skarlupka, J.H.; Hanotte, O.; Suen, G.; Elsik, C.G. Bovine Genome Database: New curated collection of selective sweeps in bovine populations across the world. Nucleic Acids Res. 2025, 54, 12–14. [Google Scholar]
- Zhuang, Z.; Xu, L.; Yang, J.; Gao, H.; Zhang, L.; Gao, X.; Li, J.; Zhu, B. Weighted Single-Step Genome-Wide Association Study for Growth Traits in Chinese Simmental Beef Cattle. Genes 2020, 11, 189. [Google Scholar]
- Aguilar, I.; Misztal, I.; Johnson, D.L.; Legarra, A.; Tsuruta, S.; Lawlor, T.J. Hot topic: A unified approach to utilize phenotypic full pedigree and genomic information for genetic evaluation of Holstein final score. J. Dairy Sci. 2010, 93, 743–752. [Google Scholar] [CrossRef]
- Henderson, C.R. Best linear unbiased estimation and prediction under a selection model. Biometrics 1975, 31, 423–447. [Google Scholar] [CrossRef] [PubMed]
- Cheng, J.; Maltecca, C.; VanRaden, P.M.; O’Connell, J.R.; Ma, L.; Jiang, J. SLEMM: Million-scale genomic predictions with window-based SNP weighting. Bioinformatics 2023, 39, btad127. [Google Scholar] [CrossRef]
- Steyn, Y.; Lourenco, D.A.L.; Misztal, I. Genomic predictions in purebreds with a multibreed genomic relationship matrix1. J. Anim. Sci. 2019, 97, 4418–4427. [Google Scholar] [CrossRef] [PubMed]
- Hayes, B.J.; Bowman, P.J.; Chamberlain, A.C.; Verbyla, K.; Goddard, M.E. Accuracy of genomic breeding values in multi-breed dairy cattle populations. Genet. Sel. Evol. 2009, 41, 51. [Google Scholar] [CrossRef]
- Mohammad Rahimi, S.; Rashidi, A.; Esfandyari, H. Accounting for differences in linkage disequilibrium in multi-breed genomic prediction. Livest. Sci. 2020, 240, 104165. [Google Scholar] [CrossRef]
- Karaman, E.; Su, G.; Croue, I.; Lund, M.S. Genomic prediction using a reference population of multiple pure breeds and admixed individuals. Genet. Sel. Evol. 2021, 53, 46. [Google Scholar] [CrossRef]
- Boerner, V.; Johnston, D.J.; Tier, B. Accuracies of genomically estimated breeding values from pure-breed and across-breed predictions in Australian beef cattle. Genet. Sel. Evol. 2014, 46, 61. [Google Scholar] [CrossRef]
- Zhou, L.; Zhu, L.; Ma, F.Y.; Gu, M.J.; Na, R.S.; Zhang, W.G. Assessing the Impact of Different Mixing Strategies on Genomic Prediction Accuracy for Beef Cattle Breeding Values in Multi-Breed Genomic Prediction. Animals 2025, 15, 2463. [Google Scholar] [CrossRef] [PubMed]
- Luan, T.; Woolliams, J.A.; Odegård, J.; Dolezal, M.; Roman-Ponce, S.I.; Bagnato, A.; Meuwissen, T.H. The importance of identity-by-state information for the accuracy of genomic selection. Genet. Sel. Evol. 2012, 44, 28. [Google Scholar] [CrossRef] [PubMed]
- Powell, J.E.; Visscher, P.M.; Goddard, M.E. Reconciling the analysis of IBD and IBS in complex trait studies. Nat. Rev. Genet. 2010, 11, 800–805. [Google Scholar] [CrossRef] [PubMed]
- Zhang, M.; Du, H.; Zhang, Y.; Zhuo, Y.; Liu, Z.; Xue, Y.; Zhou, L.; Zhou, S.; Li, W.; Liu, J.F. A high-throughput screening method for selecting feature SNPs to evaluate breed diversity and infer ancestry. Genome Res. 2025, 35, 1875–1886. [Google Scholar] [CrossRef]







| Parameter | PopA 1 | PopB 2 | PopC 3 |
|---|---|---|---|
| Step 1: Historical Generations (HG) | |||
| Phase 1 generation count (population size) | 0 (200) | 0 (500) | 0 (1000) |
| Phase 2 generation count (population size) | 1000 (200) | 1000 (200) | 1000 (1000) |
| Phase 3 generation count (population size) | 1095 (1000) | 1095 (1000) | 1095 (200) |
| Step 2: Recent Generations (RG) | |||
| Number of male founders from HG | 200 | 220 | 200 |
| Number of female founders from HG | 2800 | 2335 | 2800 |
| Selection and mating | ebv/h | ||
| Sire/Dam replacement | 0.6/0.3 | 0.5/0.3 | 0.5/0.2 |
| Mating strategy | Random | ||
| Culling scheme | ebv/L | ||
| Genome | |||
| Chromosome count | 29 (no X Chr) | ||
| Genome length | 2715.85 cM | ||
| Number of markers | 100 k | ||
| Marker/QTL genomic locations | Random | ||
| Marker/QTL allele count | 2/2 3 4 | ||
| Marker allele frequency distribution | Equal | ||
| QTL allelic effect sizes | Equal | ||
| Additive allelic effects of QTL | Gamma distribution (shape = 0.4) | ||
| Marker genotype missingness proportion | 0.01 | ||
| Marker genotyping error rate | 0.005 | ||
| Recurrent mutation frequency | 0.0001 | ||
| Pop | Percentage | GBLUP | ssGBLUP | wGBLUP | |||
|---|---|---|---|---|---|---|---|
| b 1 | SE 2 | b | SE | b | SE | ||
| PopB | 10% | 1.12 | 0.022 | 0.96 | 0.026 | 1.18 | 0.06 |
| 15% | 0.93 | 0.042 | 1.03 | 0.003 | 0.85 | 0.02 | |
| 20% | 1.23 | 0.021 | 1.23 | 0.064 | 1.39 | 0.017 | |
| PopC | 10% | 1.22 | 0.021 | 1.19 | 0.002 | 0.86 | 0.016 |
| 15% | 1.43 | 0.019 | 0.96 | 0.057 | 1.29 | 0.015 | |
| 20% | 1.34 | 0.017 | 1.24 | 0.054 | 1.15 | 0.014 | |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Zhou, L.; Zhu, L.; Ma, F.; Gu, M.; Na, R.; Zhang, W. Leveraging Fst and Genetic Distance to Optimize Reference Sets for Enhanced Cross-Population Genomic Prediction. Animals 2026, 16, 359. https://doi.org/10.3390/ani16030359
Zhou L, Zhu L, Ma F, Gu M, Na R, Zhang W. Leveraging Fst and Genetic Distance to Optimize Reference Sets for Enhanced Cross-Population Genomic Prediction. Animals. 2026; 16(3):359. https://doi.org/10.3390/ani16030359
Chicago/Turabian StyleZhou, Le, Lin Zhu, Fengying Ma, Mingjuan Gu, Risu Na, and Wenguang Zhang. 2026. "Leveraging Fst and Genetic Distance to Optimize Reference Sets for Enhanced Cross-Population Genomic Prediction" Animals 16, no. 3: 359. https://doi.org/10.3390/ani16030359
APA StyleZhou, L., Zhu, L., Ma, F., Gu, M., Na, R., & Zhang, W. (2026). Leveraging Fst and Genetic Distance to Optimize Reference Sets for Enhanced Cross-Population Genomic Prediction. Animals, 16(3), 359. https://doi.org/10.3390/ani16030359

