Next Article in Journal
Harnessing the Action Model of the Defense Responses Induced by UPSIDE® Against Plasmopara viticola in Grapevine
Next Article in Special Issue
Identification of HsfB Family in Peanut (Arachis hypogea) and Role of AhHsfB1-5A in High-Temperature Stress
Previous Article in Journal
Coordinated Ecophysiological Trait Shifts of Populus euphratica Along a Groundwater-Depth Gradient: From Carbon Acquisition Toward Water Conservation in an Arid Riparian Forest
Previous Article in Special Issue
Heterologous Expression of Sorghum bicolor PIP1-3 Gene Improves Drought Tolerance in Arabidopsis and Rapeseed
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

A Quantitative Trait Nucleotide-Based Genomic Selection Strategy for Seed Oil and Protein Content in Soybean

1
Key Laboratory of Soybean Molecular Design Breeding, Northeast Institute of Geography and Agroecology, Chinese Academy of Sciences, Changchun 130102, China
2
Jilin Academy of Agricultural Sciences (China Agricultural Science and Technology Northeast Innovation Center), Soybean Research Institute, Changchun 130033, China
3
National Maize Improvement Center, Department of Crop Genomics and Bioinformatics, College of Agronomy and Biotechnology, China Agricultural University, Beijing 100193, China
*
Authors to whom correspondence should be addressed.
These authors contributed equally to this work.
Plants 2026, 15(9), 1296; https://doi.org/10.3390/plants15091296
Submission received: 31 December 2025 / Revised: 6 April 2026 / Accepted: 18 April 2026 / Published: 22 April 2026
(This article belongs to the Special Issue Genetic Improvement of Oilseed Crops)

Abstract

In recent years, genomic selection (GS) has been widely adopted in plant breeding; however, its practical application is constrained by the high cost of genotyping large segregating populations. To address this issue, this study employed a Quantitative Trait Nucleotide (QTN)-assisted GS strategy to evaluate its efficiency in reducing genotyping costs for soybean seed oil content (OC) and protein content (PC). Based on six multi-parent F4 populations (n = 4404) derived from seven elite soybean cultivars, which were genotyped using a 20K SNP chip, we identified 83 and 110 QTNs that were significantly associated with OC and PC, respectively. Among these loci, 37 and 62 QTNs were specific to OC and PC, respectively. Genomic prediction accuracies were evaluated across different training population (TP) sizes using three marker panels: genome-wide SNPs, all detected QTNs, and trait-specific QTNs. The panel consisting of all detected QTNs exhibited significantly higher prediction accuracy than the other two panels, except for PC when using 90% of the population as the training set. Phenotypic verification of the selected individuals showed that the PC-specific QTN panel yielded higher PC values and increased OC + PC values compared with the other marker panels. These results demonstrate that a small set of QTNs provides a cost-effective approach for genomic selection in practical soybean breeding programs.

1. Introduction

Genomic selection (GS) is an effective breeding improvement method for complex quantitative traits in plants controlled by multiple genes [1,2]. Different GS models, including direct methods, such as genomic best linear unbiased prediction (GBLUP) [3], single-step best linear unbiased prediction (ssBLUP) [4,5], standard Best Linear Unbiased Prediction (sBLUP), and combined Best Linear Unbiased Prediction (cBLUP) [6], and indirect methods, such as ridge regression best linear unbiased prediction (rrBLUP) [7], BayesA, BayesB, BayesCπ, BayesDπ [8], and Bayesian LASSO [9], have been developed. Previous studies have demonstrated that distinct models yield divergent prediction values for various traits, primarily owing to differences in the assumptions underlying marker effect distributions, which in turn influence the total variance [1,2,10].
While numerous models are available for GS, its commercial application in plant breeding is hindered by a critical challenge i.e., high genotyping costs. Population size and marker density serve as key factors influencing both cost efficiency and prediction accuracy [11,12,13]. Research has demonstrated that in aquaculture species, single-nucleotide polymorphism (SNP) panels comprising 1000–2000 markers yield selection accuracies comparable to those attained via high-density genotyping [11]. An increasing number of studies have recently focused on employing low-density marker panels combined with small training populations (TPs) to reduce genotyping costs for GS in crops [14,15,16,17].
Soybean (Glycine max L.) is a major cash crop cultivated worldwide, primarily owing to its seeds being abundant in edible protein and oil. On average, soybean seeds contain approximately 40% protein content (PC) and 20% oil content (OC) [18]. The OC and PC traits in soybean are well characterized and exhibit a significant negative correlation, with the relevant functional genes exerting opposing regulatory effects on them [19,20,21]. For such antagonistic traits, the artificial selection for one trait invariably leads to the reduction in the other [22]. Both OC and PC are complex quantitative traits in soybean controlled by multiple genes [23,24]. Whole-genome marker association analysis has been effectively applied to identify QTLs and functional genes underlying seed PC and OC traits in soybean [25,26,27,28,29,30]. Although hundreds of QTLs associated with OC and PC have been documented in SoyBase (http://www.soybase.org), the genetic mechanisms underlying these traits remain to be fully elucidated. Thus, high-resolution dissection of the detailed genetic architecture of OC and PC will facilitate the development of strategies to mitigate their negative correlation.
In this study, our main objectives were to develop a strategy for reducing the genotyping cost of GS, thereby promoting its application in commercial breeding. To this end, we used a GS population consisting of 4404 individual F4 plants and selected two negatively correlated traits, OC and PC, to demonstrate the effectiveness of our strategy.

2. Results

2.1. Population Structure Analysis of the Genomic Selection Population

A total of 4404 soybean genotypes were used in this study, which were derived from six F4 single-hybrid populations (Table S1). These six populations, designated as F4GS1-6, were developed from six bi-parental crosses involving three male parents and four female parents (Table S1). Semi-sib relationships were prevalent among these populations, and the agronomic performance of quality-related traits of these parents is presented in Table S1. Two key traits, namely OC and PC, exhibited an approximately normal distribution across all six F4 segregating populations, as well as the combined population (Figure 1 and Figure S2, and Table S1). With respect to OC, the F4GS3 population had the highest average value (20.56%), while the F4GS4 population had the lowest (18.05%). For PC, the F4GS5 population displayed the highest average content (43.58%), and the F4GS3 population had the lowest (39.15%). OC and PC showed a significant negative correlation in each population, with Pearson’s correlation coefficients ranging from −0.76 to −0.86 (Table S1).
To address inter-population cross-prediction challenges, we pooled the six F4 populations for downstream marker-trait association (MTA) and GS analyses. Phylogenetic analysis and principal component analysis (PCA) revealed that individuals from the same hybrid clustered together, while F4 populations sharing a common parent and half-sib families (e.g., “F4GS1 and F4GS2”, “F4GS3 and F4GS4”) grouped closely (Figure 1C,D). PCA and phylogenetic results were consistent (Figure 1D), confirming the six F4 populations were genetically distinct but related. This structure characterized by distinct but related semi-sib subpopulations was incorporated into MTA and GS. It minimized false MTAs and facilitated the evaluation of GS model transferability, forming a feedback loop.

2.2. QTN Identification for OC and PC via Whole-Genome Marker Association Analysis

We genotyped the soybean population using a customized soybean genotyping panel, which included 20,659 single-nucleotide polymorphism (SNP) markers mainly derived from the genic/protein-coding regions of the soybean genome. After quality control filtering, we retained 9942 high-quality SNPs for further genetic analysis. For these 9942 SNPs, we estimated marker density as the number of SNPs per 1 Mb contiguous window in the soybean genome. Our results showed that these SNP markers almost covered the entire soybean genome, with relatively lower density near centromeres (Figure 2A).
To reduce the number of markers for GS, we performed MTA analysis on a combined population of 4404 lines from six individual populations, aiming to detect quantitative trait nucleotides (QTNs) associated with OC and PC (Table S1). Six multi-locus association models (mrMLM, FASTmrMLM, FASTmrEMMA, ISIS EM-BLASSO, pLARmEB, and pKWmEB) were used to enhance detection power and result reliability. Among the six models, OC-associated QTNs detected ranged from 11 to 41; and 83 non-redundant OC QTNs were retained (detected by ≥2 methods, or single-method LOD ≥ 3.0), that are distributed across all soybean chromosomes except chr.13 (Figure 2B; Table S2). For PC, QTNs detected ranged from 10 to 49; and 110 non-redundant PC QTNs were retained (detected by ≥2 methods, or single-method LOD ≥ 3.0), that are present on all chromosomes except chr.11 (Figure 2C; Table S2).
Comparison of the chromosomal positions of QTNs demonstrated that 46 out of 83 OC QTNs and 48 out of 110 PC QTNs were located within 25 shared genomic regions (genomic distance < 1 Mb), while 37 OC QTNs and 62 PC QTNs were identified as trait-specific loci (Table S2). The overlapping proportions, accounting for 55.40% of OC QTNs and 43.60% of PC QTNs, thus revealing the genetic basis underlying the antagonistic relationship between these two traits.

2.3. Effects of Model Selection on GS Prediction Accuracy

The prediction performance of six GS models, namely Bayes A, Bayes B, Bayes C, Bayesian Ridge Regression (BRR), Bayesian Lasso (BL), and rrBLUP, were evaluated using 10 repetitions of cross-validation. A total of 90% of the population (n = 3964) was assigned as the test set, with the correlation between predicted and observed phenotypic values serving as the key evaluation metric (Figure 3 and Figure S3).
For the OC trait, the average prediction accuracies achieved by the six models using all SNP markers were 0.749, 0.735, 0.748, 0.750, 0.733, and 0.731, respectively. In contrast, when the 83 OC-associated QTNs were employed instead of all SNP markers, the corresponding average prediction accuracies were 0.747, 0.753, 0.761, 0.749, 0.753, and 0.755, respectively (Figure 3A).
A similar trend was observed for the PC trait. When all SNP markers were used, the average prediction accuracies of the six models were 0.756, 0.764, 0.756, 0.756, 0.757, and 0.763, respectively. In comparison, the use of 110 PC-related QTNs resulted in average prediction accuracies of 0.777, 0.764, 0.765, 0.778, 0.772, and 0.778 (Figure 3B).
Furthermore, paired t-tests with Bonferroni correction revealed no significant differences in prediction performance among the six models, regardless of all SNPs or trait-associated QTNs used (Table S3). Notably, while the identified QTN sets included several variants with low PVE (<0.01%)—which could represent minor QTLs or statistical noise—their exclusion (2 for OC and 4 for PC) had negligible impacts on GS accuracy (change < 0.02, p > 0.01; Table S4). Robustness analysis (Figure S4) further revealed nearly identical CV% profiles between the full and filtered QTN sets across all TP ratios. This indicates that these minor-effect variants do not introduce detrimental noise or compromise model stability. Consequently, the full QTN sets were retained for all subsequent analyses to ensure a comprehensive evaluation of the genetic architecture.

2.4. Effects of Training Population Size and Genotyping Marker Types on Prediction Accuracy

Training population (TP) size and genotyping marker quantity are critical factors influencing the cost of GS. We applied the rrBLUP model to assess their impacts on the predictive accuracy of OC and PC, evaluating three marker panels: 9942 SNPs (GSall), 83 OC-associated and 110 PC-associated QTNs (GSQTN), together with 37 OC-specific and 62 PC-specific QTNs (Specific QTNs) (Figure 4; Table S5).
When the TP proportion rose from 10% to 90% (test plants from 440 to 3964), all marker sets presented elevated predictive accuracy. GSall boosted OC accuracy by 0.054 and PC by 0.064, while GSQTN improved OC and PC accuracy by 0.040 and 0.026, respectively (Figure 4A). The smaller accuracy increment of GSQTN with expanding TP confirmed that QTN markers can realize favorable predictive performance with a limited TP size. In contrast, specific QTNs showed overall low accuracy, with mean values ranging from 0.686 to 0.717 for both traits, and its predictive performance remained stable with negligible responses to TP size changes (Figure 4A,B).
According to Tukey’s HSD test (p < 0.05), GSQTN exhibited significantly higher accuracy than GSall for both traits at small TP sizes (10–30%), and this advantage faded as TP increased (Table S5). At TP = 90%, GSall and GSQTN performed similarly for PC, yet GSQTN still maintained significantly superior accuracy for OC. Specific QTNs fell into an independent statistical group in post hoc tests, with accuracy markedly lower than the other two panels. Nonetheless, its stable predictivity facilitates the effective capture of core genetic signals (Figure 4).

2.5. Verification of Trait-Associated and Trait-Specific QTNs in Genomic Selection

To verify the effectiveness of trait-associated QTNs and trait-specific QTNs in GS applications, half of the F4 population was assigned as the TP (n = 2202) using the rrBLUP model, while the other half served as the validation population (VP) to evaluate the performance differences between these two marker sets (Figure 5).
For the OC trait, 37 OC-specific QTNs and 83 OC-associated QTNs were used to compare the predicted OC values of the top 20% individuals (n = 440) with higher OC in the VP against their actual measured OC values. The results showed that both subgroups exhibited significantly higher OC levels than the average OC of the entire VP, confirming the effectiveness of GS using QTNs for oil trait selection (Figure 5A). No significant difference in OC was observed between the two subgroups, indicating that 83 OC-associated QTNs did not improve the prediction accuracy. Additionally, both subgroups displayed lower OC + PC values compared to the VP average (Figure 5B), reflecting a negative genetic correlation between OC and PC. This finding suggested that the limited number of OC-specific QTNs was insufficient to counteract the antagonistic effects of common QTNs.
For the PC trait, 62 PC-specific QTNs and 110 PC-associated QTNs were employed. Both selection panels successfully identified VP subgroups with significantly increased PC levels (Figure 5C). The subgroup selected based on PC-specific QTNs had an average PC of 44.10%, which was higher than the 43.80% observed in the subgroup selected using all PC-associated QTNs (p = 0.018). This result indicates that common QTNs attenuated the predictive signal of PC
The OC + PC value of the PC-specific QTN subgroup reached 61.80%, which was significantly higher than the 61.60% of the all-QTN subgroup (p = 0.013). This outcome is attributed to the greater increase in PC, thereby optimizing seed quality (Figure 5D). These findings demonstrate that PC trait-specific QTNs can alleviate the OC-PC trade-off, with a more prominent effect on OC + PC than OC trait-specific QTNs (Figure 5B,D).

3. Discussion

Cost-effective GS strategies are essential for the extensive implementation of GS in plant breeding programs [31,32], with their value assessed by the genetic gain attained per unit cost [33]. Population-specific linkage disequilibrium (LD) patterns enable the development of cost-efficient genomic prediction models employing low-density SNP panels [34]. The optimization of the TP is designed to achieve a balance between maximizing genetic gain via elevated genomic prediction precision and cutting phenotyping costs by shrinking the TP scale [35]. This study proposes an innovative approach: detecting QTNs via MTA in a moderately sized TP, and subsequently establishing low-density marker panels for high-throughput screening in practical GS applications.

3.1. Targeted QTN Sets as an Alternative to Dense Genome-Wide Markers for GS Application

A key advantage of QTN-based GS relative to genome-wide marker-based strategies resides in its maintenance of model simplicity and cost-effectiveness. Our findings revealed that six conventional GS models (including rrBLUP, BayesA, BayesB, BRR, BL and BayesC) exhibited comparable predictive performance for OC and PC, irrespective of whether they were trained using 9942 genome-wide markers (GSall) or 83–110 QTNs (GSQTN). Strikingly, GSQTN surpassed GSall in predictive accuracy when small TPs were utilized: with a 10% TP, GSQTN attained an OC prediction accuracy of 0.721, exceeding that of GSall implemented with larger training populations (Figure 4). This superiority can be ascribed to the enrichment of genetic signals derived from stable, trait-specific QTNs and the concomitant reduction in background noise [36,37,38]. Although the overall prediction accuracy (PA) of GSQTN was moderate, the stable and trait-targeted characteristics of specific QTN panels equip them with an enhanced ability to capture the core genetic determinants of target traits.

3.2. Trait-Specific QTNs in Mitigating OC-PC Trade-Off

Soybean OC and PC exhibit a strong negative correlation, which constitutes a major bottleneck in soybean breeding [22,39]. This negative correlation has traditionally been ascribed to pleiotropy and tight genetic linkage between loci regulating the two traits [40]. Similar trade-offs are prevalent in wheat, where GS indices have been employed to decouple grain yield from protein quality. By integrating dough rheological traits into GS models [41] or using genomic-assisted approaches to optimize nitrogen utilization [42], breeders can identify genotypes that balance yield and protein more effectively than classical phenotypic selection. These advancements underscore the potential of genomic strategies in mitigating negative correlations between complex traits. Trait-specific QTNs offer an alternative approach to breaking unfavorable linkages without relying on large segregating populations: PC-specific QTNs achieved a PC of 44.10% (vs. 43.80% with all PC QTNs, p = 0.018) with a comparable level of OC loss, and a higher combined OC + PC (61.80% vs. 61.60%, p = 0.013) (Figure 5D), thereby enhancing seed value by effectively balancing the inherent OC-PC trade-off.

4. Materials and Methods

A schematic workflow diagram (Figure S1) outlines the experimental design and analytical pipeline, including population development, phenotypic/genotypic data collection, MTA, GS model comparison, TP optimization, and QTN effect evaluation on the OC–PC trade-off.

4.1. Plant Materials

Seven soybean cultivars from Northeast China were used to construct six F4 populations (Table S1). Their PCs/OCs were: FNGS0852 (44.64%/19.35%), FNGS0225 (40.95%/21.60%), FNGS0239 (43.46%/19.67%), FNGS0217 (40.72%/22.01%), FNGS0256 (44.35%/20.94%), FNGS0280 (44.84%/18.52%), and FNGS0301 (46.69%/17.90%). Populations and parents were planted in Jiamusi (46°48′ N, 130°22′ E) in 2019. Each F4 family was derived via SSD, with OC/PC measured on bulk seeds. Plots were single-row (4.5 m × 0.5 m row spacing) without replication; border rows minimized edge effects, and standard agronomic practices ensured reliable phenotypic data.

4.2. Phenotype Data Collection and Analysis

Near infra-red (NIR) calibration models for OC/PC were built using 200 diverse soybean accessions (excluding 4404 F4 lines), with reference values from Soxhlet apparatus (Büchi Labortechnik AG, Flawil, Switzerland) for OC and Kjeldahl analyzer (FOSS Analytical, Hillerød, Denmark) for PC. Partial least squares (PLS) regression with full cross-validation was used, optimized by predicted sum of squares (PRESS). External validation (50 samples) showed: OC (root mean square error of calibration-RMSEC = 0.30%, root mean square error of prediction-RMSEP = 0.50%, coefficient of determination for thecalibration set-R2_cal = 0.99, coefficient of determination for the validation-R2_val = 0.97); and PC (RMSEC = 0.63%, RMSEP = 0.52%, R2_cal = 0.98, R2_val = 0.95). Spectral ranges (4100–4940, 5390–6690, 6900–7130, 7185–9000 cm−1) excluded water bands; standard normal variate (SNV) and detrending corrected scatter. Models were applied to F4 lines. OC/PC density maps (ggpubr R package, within R v3.6.0), descriptive statistics (psych “describe.by”), and Pearson’s r (psych “cor”) were analyzed in R (v3.6.0, https://www.r-project.org/, accessed 20 April 2020).

4.3. Genotyping Analysis

Genomic DNA was extracted via cetyltrimethylammonium bromide-CTAB method [43]. Genotyping used a 20K chip (MOLBREEDING Biotechnology Co., Ltd., Beijing, China), with 20,659 SNPs selected (missing rate ≤ 20%, heterozygosity ≤ 30%, allele frequency 5–95%; 50 kb sliding window, coding region preference). Of these, 19,954 SNPs were called; and filtering (missing rate > 10%, minor allele frequency-MAF < 0.05) retained 9942 high-quality SNPs (Table S6).

4.4. Phylogenetic Relationship and PCA

Phylogenetic analysis was performed by FastTree (v2.2, http://www.microbesonline.org/fasttree/, accessed 20 April 2020) [44], visualization ggtree (R package) [45] and PCA PLINK (v1.9, https://www.cog-genomics.org/plink/, accessed 20 April 2020) [46] were performed using 9942 SNPs to infer population genetic structure.

4.5. MTA Analysis

MTA was done by using mrMLM v3.6.0 package (https://cran.r-project.org/web/packages/mrMLM/, accessed 21 April 2020) [47] with six multi-locus methods. Critical p-values: 0.01 (except FASTmrEMMA, 0.005); LOD = 3.0 (p < 0.0002) for significant QTNs. The first six PCs (capturing genetic variation, Figure 1D) and kinship matrix controlled population structure. Consensus QTNs were detected by ≥2 methods, or single method (LOD ≥ 3.0, ≤50 kb from annotated genes). LD decay (≈450 kb) set 500 kb as the clustering threshold; redundant QTNs were merged, retaining the highest LOD markers, yielding 83 (OC) and 110 (PC) non-redundant QTNs.

4.6. Genomic Selection

GS models (rrBLUP, BayesA/B/C, BL, BRR) were implemented in R (v3.6.0) via BGLR [48] and rrBLUP [7]. Ten independent 10-fold cross-validations compared prediction accuracies, calculated as mean correlation of predicted/observed values per replicate.

5. Conclusions

In summary, this study presents a new strategy that integrating whole-genome marker association analysis with QTN-assisted GS to reduce genotyping costs as well as mitigates the negative genetic correlation between OC and PC in soybean. However, our findings are based on a single environment and a specific set of multi-parent populations derived from seven elite cultivars. For commercial applications of this strategy in the crop breeding, there is need for its validation across multiple environments and independent breeding populations. Nevertheless, our results provide a roadmap for future efforts to simultaneously improve negatively correlated traits in soybean and other crops.

Supplementary Materials

The following supporting information can be downloaded at https://www.mdpi.com/article/10.3390/plants15091296/s1, Figure S1. Schematic workflow diagram. Figure S2. QQ plot of the oil content and protein content. Figure S3. Distribution of QTN and SNP effect for genomics selection. Figure S4. Robustness analysis of GS models using full versus filtered QTN sets. Table S1. Six F4 populations derived from seven parents. Table S2. QTNs detected from MTA analysis. Table S3. Comparison of predictive accuracy between SNP and QTNS Markers. Table S4. Comparison of prediction accuracy between QTN set and filtered QTN. Table S5. Comparison of genomic prediction accuracy among three marker panels. Table S6. Summary of SNP quality control (QC) filtering.

Author Contributions

S.Y., X.F. and X.W. conceived the study and designed the experimental framework. G.L. and H.Z. performed the statistical analysis of phenotypic data, developed the genomic selection (GS) protocols, and drafted the manuscript. K.T. designed the field trials and analyzed phenotypic data. J.L. implemented the field trials and contributed to phenotypic data analysis. S.Y. and J.A.B. secured funding and oversaw project execution. All authors have read and agreed to the published version of the manuscript.

Funding

This study was supported by the Biological Breeding-National Science and Technology Major Project (2023ZD0403201) and the National Natural Science Foundation of China (32201825).

Data Availability Statement

The original contributions presented in this study are included in the article. Further inquiries can be directed to the corresponding authors.

Acknowledgments

We thank Li Candong from the Jiamusi Branch of Heilongjiang Academy of Agricultural Sciences for his assistance with the experiments, and we extend our gratitude to the institution for its invaluable support.

Conflicts of Interest

The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.

Abbreviations

The following abbreviations are used in this manuscript:
OCOil content
PCProtein content
QTNsQuantitative trait nucleotides
GSGenomic selection
TPTraining populations

References

  1. Escamilla, D.M.; Li, D.; Negus, K.L.; Kappelmann, K.L.; Kusmec, A.; Vanous, A.E.; Schnable, P.S.; Li, X.; Yu, J. Genomic selection: Essence, applications, and prospects. Plant Genome 2025, 18, e70053. [Google Scholar] [CrossRef]
  2. Alemu, A.; Åstrand, J.; Montesinos-López, O.A.; Isidro y Sánchez, J.; Fernández-Gónzalez, J.; Tadesse, W.; Vetukuri, R.R.; Carlsson, A.S.; Ceplitis, A.; Crossa, J.; et al. Genomic selection in plant breeding: Key factors shaping two decades of progress. Mol. Plant 2024, 17, 552–578. [Google Scholar] [CrossRef] [PubMed]
  3. Bernardo, R. Prediction of Maize single-cross performance using RFLPs and information from related hybrids. Crop Sci. 1994, 34, 20–25. [Google Scholar] [CrossRef]
  4. Christensen, O.F.; Lund, M.S. Genomic prediction when some animals are not genotyped. Genet. Sel. Evol. 2010, 42, 2. [Google Scholar] [CrossRef]
  5. Aguilar, I.; Misztal, I.; Johnson, D.L.; Legarra, A.; Tsuruta, S.; Lawlor, T.J. Hot topic: A unified approach to utilize phenotypic, full pedigree, and genomic information for genetic evaluation of Holstein final score. J. Dairy Sci. 2010, 93, 743–752. [Google Scholar] [CrossRef]
  6. Heffner, E.L.; Sorrells, M.E.; Jannink, J.-L. Genomic selection for crop improvement. Crop Sci. 2009, 49, 1–12. [Google Scholar] [CrossRef]
  7. Endelman, J.B. Ridge regression and other kernels for genomic selection with R package rrBLUP. Plant Genome 2011, 4, 250–255. [Google Scholar] [CrossRef]
  8. Habier, D.; Fernando, R.L.; Kizilkaya, K.; Garrick, D.J. Extension of the bayesian alphabet for genomic selection. BMC Bioinform. 2011, 12, 186. [Google Scholar] [CrossRef]
  9. Park, T.; Casella, G. The bayesian lasso. J. Am. Stat. Assoc. 2008, 103, 681–686. [Google Scholar] [CrossRef]
  10. Parveen, R.; Kumar, M.; Swapnil; Singh, D.; Shahani, M.; Imam, Z.; Sahoo, J.P. Understanding the genomic selection for crop improvement: Current progress and future prospects. Mol. Genet. Genom. 2023, 298, 813–821. [Google Scholar] [CrossRef]
  11. Kriaridou, C.; Tsairidou, S.; Houston, R.D.; Robledo, D. Genomic prediction using low density marker panels in aquaculture: Performance across species, traits, and genotyping platforms. Front. Genet. 2020, 11, 124. [Google Scholar] [CrossRef]
  12. e Sousa, M.B.; Galli, G.; Lyra, D.H.; Granato, Í.S.C.; Matias, F.I.; Alves, F.C.; Fritsche-Neto, R. Increasing accuracy and reducing costs of genomic prediction by marker selection. Euphytica 2019, 215, 18. [Google Scholar] [CrossRef]
  13. Zhang, A.; Wang, H.; Beyene, Y.; Semagn, K.; Liu, Y.; Cao, S.; Cui, Z.; Ruan, Y.; Burgueño, J.; San Vicente, F.; et al. Effect of trait heritability, training population size and marker density on genomic prediction accuracy estimation in 22 bi-parental tropical maize populations. Front. Plant Sci. 2017, 8, 01916. [Google Scholar] [CrossRef]
  14. Crossa, J.; Pérez-Rodríguez, P.; Cuevas, J.; Montesinos-López, O.; Jarquín, D.; de los Campos, G.; Burgueño, J.; González-Camacho, J.M.; Pérez-Elizalde, S.; Beyene, Y.; et al. Genomic selection in plant breeding: Methods, models, and perspectives. Trends Plant Sci. 2017, 22, 961–975. [Google Scholar] [CrossRef]
  15. Krishnappa, G.; Savadi, S.; Tyagi, B.S.; Singh, S.K.; Mamrutha, H.M.; Kumar, S.; Mishra, C.N.; Khan, H.; Gangadhara, K.; Uday, G.; et al. Integrated genomic selection for rapid improvement of crops. Genomics 2021, 113, 1070–1086. [Google Scholar] [CrossRef]
  16. Kaler, A.S.; Purcell, L.C.; Beissinger, T.; Gillman, J.D. Genomic prediction models for traits differing in heritability for soybean, rice, and maize. BMC Plant Biol. 2022, 22, 87. [Google Scholar] [CrossRef]
  17. Bandillo, N.B.; Jarquin, D.; Posadas, L.G.; Lorenz, A.J.; Graef, G.L. Genomic selection performs as effectively as phenotypic selection for increasing seed yield in soybean. Plant Genome 2023, 16, e20285. [Google Scholar] [CrossRef] [PubMed]
  18. Wilson, R.F. Seed composition. In Soybeans: Improvement, Production, and Uses; ASA/CSSA/SSSA: Madison, WI, USA, 2004; pp. 621–677. [Google Scholar]
  19. Duan, Z.; Li, Q.; Wang, H.; He, X.; Zhang, M. Genetic regulatory networks of soybean seed size, oil and protein contents. Front. Plant Sci. 2023, 14, 1160418. [Google Scholar] [CrossRef]
  20. Goettel, W.; Zhang, H.; Li, Y.; Qiao, Z.; Jiang, H.; Hou, D.; Song, Q.; Pantalone, V.R.; Song, B.-H.; Yu, D.; et al. POWR1 is a domestication gene pleiotropically regulating seed quality and yield in soybean. Nat. Commun. 2022, 13, 3051. [Google Scholar] [CrossRef] [PubMed]
  21. Zheng, H.; Feng, X.; Wang, L.; Shao, W.; Guo, S.; Zhao, D.; Li, J.; Yan, L.; Miao, L.; Sun, B.; et al. GmSop20 functions as a key coordinator of the oil-to-protein patio in soybean seeds. Adv. Sci. 2025, 12, e05181. [Google Scholar] [CrossRef]
  22. Brown, K.E.; Kelly, J.K. Antagonistic pleiotropy can maintain fitness variation in annual plants. J. Evol. Biol. 2018, 31, 46–56. [Google Scholar] [CrossRef] [PubMed]
  23. Shi, Q.; Mo, W.; Zheng, X.; Zhao, X.; Chen, X.; Zhang, L.; Qin, J.; Yang, Z.; Zuo, Z. Balancing act: Progress and prospects in breeding soybean varieties with high oil and seed protein content. Front. Plant Sci. 2025, 16, 1560845. [Google Scholar] [CrossRef] [PubMed]
  24. Sun, J.; Li, W.; Wei, X.; Shou, H.; Tran, L.P.; Feng, X.; Wang, S. Mechanistic roles of GmSWEET10a/b and GmSUT1 in the oil-protein balance in soybean mature seeds at transcriptional and metabolic levels. Plant J. 2025, 123, e70435. [Google Scholar] [CrossRef]
  25. Karikari, B.; Li, S.; Bhat, J.A.; Cao, Y.; Kong, J.; Yang, J.; Gai, J.; Zhao, T. Genome-wide detection of major and epistatic effect QTLs for seed protein and oil content in soybean under multiple environments using high-density bin map. Int. J. Mol. Sci. 2019, 20, 979. [Google Scholar] [CrossRef]
  26. Bu, M.; Zhang, Y.; Xu, W.; Li, Y.; Yu, H.; Zhang, Y.; Yang, S.; Bhat, J.A.; Feng, X. Genome-wide detection of superior haplotypes for seed oil and protein content in Northeast China soybean (Glycine max L.) germplasm. Front. Plant Sci. 2026, 17, 1767299. [Google Scholar] [CrossRef]
  27. Vuong, T.D.; He, G.; Hu, H.; Valliyodan, B.; Lee, D.; Bayer, P.E.; Schapaugh, W.T.; Hessel, R.; Edwards, D.; Nguyen, H.T. Identification of new genomic loci for seed protein and oil content in the soybean pangenome using genome-wide association and haplotype analyses. Theor. Appl. Genet. 2025, 138, 237. [Google Scholar] [CrossRef]
  28. Doszhanova, B.; Zatybekov, A.; Didorenko, S.; Fang, C.; Abugalieva, S.; Turuspekov, Y. Genome-wide association study of seed quality and yield traits in a soybean collection from southeast Kazakhstan. Agronomy 2024, 14, 2746. [Google Scholar] [CrossRef]
  29. Diers, B.W.; Specht, J.E.; Graef, G.L.; Song, Q.; Rainey, K.M.; Ramasubramanian, V.; Liu, X.; Myers, C.L.; Stupar, R.M.; An, Y.Q.; et al. Genetic architecture of protein and oil content in soybean seed and meal. Plant Genome 2023, 16, e20308. [Google Scholar] [CrossRef]
  30. Zhu, Z.; Wang, Y.; Liu, S.; Wang, S.; Li, J.; Fang, C.; Liu, Y.; Yang, X.; Tian, D.; Song, S.; et al. Genomic atlas of 8,105 accessions reveals stepwise domestication, global dissemination, and improvement trajectories in soybean. Cell 2025, 188, 6519–6535. [Google Scholar] [CrossRef]
  31. Delomas, T.A.; Hollenbeck, C.M.; Matt, J.L.; Thompson, N.F. Evaluating cost-effective genotyping strategies for genomic selection in oysters. Aquaculture 2023, 562, 738844. [Google Scholar] [CrossRef]
  32. Gorjanc, G.; Battagin, M.; Dumasy, J.F.; Antolin, R.; Gaynor, R.C.; Hickey, J.M. Prospects for cost-effective genomic selection via accurate within-family imputation. Crop Sci. 2017, 57, 216–228. [Google Scholar] [CrossRef]
  33. Budhlakoti, N.; Kushwaha, A.K.; Rai, A.; Chaturvedi, K.K.; Kumar, A.; Pradhan, A.K.; Kumar, U.; Kumar, R.R.; Juliana, P.; Mishra, D.C.; et al. Genomic selection: A tool for accelerating the efficiency of molecular breeding for development of climate-resilient crops. Front. Genet. 2022, 13, 832153. [Google Scholar] [CrossRef] [PubMed]
  34. Silva, F.F.; Jerez, E.A.Z.; de Resende, M.D.V.; Viana, J.M.S.; Azevedo, C.F.; Lopes, P.S.; Nascimento, M.; de Lima, R.O.; Guimarães, S.E.F. Bayesian model combining linkage and linkage disequilibrium analysis for low density-based genomic selection in animal breeding. J. Appl. Anim. Res. 2018, 46, 873–878. [Google Scholar] [CrossRef]
  35. Berro, I.; Lado, B.; Nalin, R.S.; Quincke, M.; Gutiérrez, L. Training population optimization for genomic selection. Plant Genome 2019, 12, 190028. [Google Scholar] [CrossRef]
  36. Teng, J.; Huang, S.; Chen, Z.; Gao, N.; Ye, S.; Diao, S.; Ding, X.; Yuan, X.; Zhang, H.; Li, J.; et al. Optimizing genomic prediction model given causal genes in a dairy cattle population. J. Dairy Sci. 2020, 103, 10299–10310. [Google Scholar] [CrossRef] [PubMed]
  37. Bhering, L.L.; Junqueira, V.S.; Peixoto, L.A.; Cruz, C.D.; Laviola, B.G. Comparison of methods used to identify superior individuals in genomic selection in plant breeding. Genet. Mol. Res. 2015, 14, 10888–10896. [Google Scholar] [CrossRef]
  38. Parsons, T.L.; Ralph, P.L. Large effects and the infinitesimal model. Theor. Popul. Biol. 2024, 156, 117–129. [Google Scholar] [CrossRef]
  39. Li, H.; Sun, J.; Zhang, Y.; Wang, N.; Li, T.; Dong, H.; Yang, M.; Xu, C.; Hu, L.; Liu, C.; et al. Soybean oil and protein: biosynthesis, regulation and strategies for genetic improvement. Plant Cell Environ. 2026, 15272. [Google Scholar] [CrossRef]
  40. Chebib, J.; Guillaume, F. Pleiotropy or linkage? Their relative contributions to the genetic correlation of quantitative traits and detection by multitrait GWA studies. Genetics 2021, 219, iyab159. [Google Scholar] [CrossRef]
  41. Michel, S.; Löschenberger, F.; Ametz, C.; Pachler, B.; Sparry, E.; Bürstmayr, H. Combining grain yield, protein content and protein quality by multi-trait genomic selection in bread wheat. Theor. Appl. Genet. 2019, 132, 2767–2780. [Google Scholar] [CrossRef]
  42. Michel, S.; Löschenberger, F.; Ametz, C.; Pachler, B.; Sparry, E.; Bürstmayr, H. Simultaneous selection for grain yield and protein content in genomics-assisted wheat breeding. Theor. Appl. Genet. 2019, 132, 1745–1760. [Google Scholar] [CrossRef]
  43. Milligan, B.G. Purification of chloroplast DNA using hexadecyltrimethylammonium bromide. Plant Mol. Biol. Rep. 1989, 7, 144–149. [Google Scholar] [CrossRef]
  44. Price, M.N.; Dehal, P.S.; Arkin, A.P. Fasttree 2—Approximately maximum-likelihood trees for large alignments. PLoS ONE 2010, 5, e9490. [Google Scholar] [CrossRef]
  45. Yu, G.; Smith, D.K.; Zhu, H.; Guan, Y.; Lam, T.T.-Y. ggtree: An R package for visualization and annotation of phylogenetic trees with their covariates and other associated data. Methods Ecol. Evol. 2017, 8, 28–36. [Google Scholar] [CrossRef]
  46. Purcell, S.; Neale, B.; Todd-Brown, K.; Thomas, L.; Ferreira, M.A.R.; Bender, D.; Maller, J.; Sklar, P.; de Bakker, P.I.W.; Daly, M.J.; et al. PLINK: A tool set for whole-genome association and population-based linkage analyses. Am. J. Hum. Genet. 2007, 81, 559–575. [Google Scholar] [CrossRef]
  47. Zhang, Y.-W.; Tamba, C.L.; Wen, Y.-J.; Li, P.; Ren, W.-L.; Ni, Y.-L.; Gao, J.; Zhang, Y.-M. mrMLM v4.0.2: An R platform for multi-locus genome-wide association studies. Genom. Proteom. Bioinform. 2020, 18, 481–487. [Google Scholar] [CrossRef] [PubMed]
  48. Pérez, P.; de los Campos, G. Genome-wide regression and prediction with the BGLR statistical package. Genetics 2014, 198, 483–495. [Google Scholar] [CrossRef] [PubMed]
Figure 1. Oil content (A) and protein content (B) distribution of six populations. The X-axis represents the oil (%) or protein content (%); density curves with different colors represent the distribution of the oil or protein content in different populations; vertical dashed lines represent the average trait values for different populations. Phylogenetic analysis (C) and principal component analysis (PCA) of six populations are shown (D). In the PCA scatterplot, different colors represent individuals from different populations, and the ellipses represent the standard error ranges of PCA1 and PCA2 for each population.
Figure 1. Oil content (A) and protein content (B) distribution of six populations. The X-axis represents the oil (%) or protein content (%); density curves with different colors represent the distribution of the oil or protein content in different populations; vertical dashed lines represent the average trait values for different populations. Phylogenetic analysis (C) and principal component analysis (PCA) of six populations are shown (D). In the PCA scatterplot, different colors represent individuals from different populations, and the ellipses represent the standard error ranges of PCA1 and PCA2 for each population.
Plants 15 01296 g001
Figure 2. Position distribution of 9942 SNP markers in the soybean genome, from the customized panel (A). Distribution of quantitative trait nucleotides (QTNs) of the two traits, namely OC (B) and PC (C), identified via MTA analysis across the soybean genome. The X-axis represents chromosomal coordinates, and the Y-axis represents the significance level of QTNs. In panel (B), yellow dots represent QTNs associated with oil content (OC); in panel (C), blue dots represent QTNs associated with protein content (PC).
Figure 2. Position distribution of 9942 SNP markers in the soybean genome, from the customized panel (A). Distribution of quantitative trait nucleotides (QTNs) of the two traits, namely OC (B) and PC (C), identified via MTA analysis across the soybean genome. The X-axis represents chromosomal coordinates, and the Y-axis represents the significance level of QTNs. In panel (B), yellow dots represent QTNs associated with oil content (OC); in panel (C), blue dots represent QTNs associated with protein content (PC).
Plants 15 01296 g002
Figure 3. GS prediction based on genome-wide SNP markers and QTNs for OC (A) and PC (B) using six methods. The height of the histogram represents the GS prediction accuracy, and the error bar represents the standard error of the prediction accuracy.
Figure 3. GS prediction based on genome-wide SNP markers and QTNs for OC (A) and PC (B) using six methods. The height of the histogram represents the GS prediction accuracy, and the error bar represents the standard error of the prediction accuracy.
Plants 15 01296 g003
Figure 4. Prediction accuracy of OC (A) and PC (B) with different training population sizes (accounting for 10–90% of the whole population) and comparison of prediction accuracy of genomic selection of oil content and protein content (using trait-specific QTNs, all QTNs and all SNPs).
Figure 4. Prediction accuracy of OC (A) and PC (B) with different training population sizes (accounting for 10–90% of the whole population) and comparison of prediction accuracy of genomic selection of oil content and protein content (using trait-specific QTNs, all QTNs and all SNPs).
Plants 15 01296 g004
Figure 5. Effects of subgroup screening using the GS model with different QTN sets. (A,B) True OC and PC distributions for the subgroups of the top 20% of individuals selected based on genomic selection using the two OC QTN sets. (C,D) True PC and OC distributions for the subgroups of the top 20% of individuals selected based on genomic selection using the two PC QTN sets. Comparison among groups in the figure was performed using a t-test; **** p < 0.0001, * p < 0.05, and ns represents no significant difference.
Figure 5. Effects of subgroup screening using the GS model with different QTN sets. (A,B) True OC and PC distributions for the subgroups of the top 20% of individuals selected based on genomic selection using the two OC QTN sets. (C,D) True PC and OC distributions for the subgroups of the top 20% of individuals selected based on genomic selection using the two PC QTN sets. Comparison among groups in the figure was performed using a t-test; **** p < 0.0001, * p < 0.05, and ns represents no significant difference.
Plants 15 01296 g005
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Li, G.; Zhou, H.; Akhter Bhat, J.; Tang, K.; Leng, J.; Feng, X.; Wang, X.; Yang, S. A Quantitative Trait Nucleotide-Based Genomic Selection Strategy for Seed Oil and Protein Content in Soybean. Plants 2026, 15, 1296. https://doi.org/10.3390/plants15091296

AMA Style

Li G, Zhou H, Akhter Bhat J, Tang K, Leng J, Feng X, Wang X, Yang S. A Quantitative Trait Nucleotide-Based Genomic Selection Strategy for Seed Oil and Protein Content in Soybean. Plants. 2026; 15(9):1296. https://doi.org/10.3390/plants15091296

Chicago/Turabian Style

Li, Guang, Huangkai Zhou, Javaid Akhter Bhat, Kuanqiang Tang, Jiantian Leng, Xianzhong Feng, Xiangfeng Wang, and Suxin Yang. 2026. "A Quantitative Trait Nucleotide-Based Genomic Selection Strategy for Seed Oil and Protein Content in Soybean" Plants 15, no. 9: 1296. https://doi.org/10.3390/plants15091296

APA Style

Li, G., Zhou, H., Akhter Bhat, J., Tang, K., Leng, J., Feng, X., Wang, X., & Yang, S. (2026). A Quantitative Trait Nucleotide-Based Genomic Selection Strategy for Seed Oil and Protein Content in Soybean. Plants, 15(9), 1296. https://doi.org/10.3390/plants15091296

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop