Next Article in Journal
Elevational Metabolic Reprogramming Optimizes Flavonoid Accumulation and Antioxidant Capacity in Chimonobambusa utilis Leaves
Previous Article in Journal
Dissecting the Phenotypic Regulation Characteristics of Lodging Resistance in Dry Direct Seeding Rice: Insights from Stem Mechanics and Structural Traits
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Development of a Medium-Density Genotyping Platform to Accelerate Genetic Gain in Fresh Edible Maize

1
CIMMYT-China Specialty Maize Research Center, Crop Breeding and Cultivation Research Institute, Shanghai Academy of Agricultural Sciences, Shanghai 201403, China
2
International Maize and Wheat Improvement Center (CIMMYT), El Batan, Texcoco 56237, Mexico
3
Shenyang City Key Laboratory of Maize Genomic Selection Breeding, College of Bioscience and Biotechnology, Shenyang Agricultural University, Shenyang 110866, China
*
Authors to whom correspondence should be addressed.
These authors contributed equally to this work.
Plants 2026, 15(9), 1288; https://doi.org/10.3390/plants15091288
Submission received: 21 March 2026 / Revised: 17 April 2026 / Accepted: 20 April 2026 / Published: 22 April 2026
(This article belongs to the Section Plant Genetics, Genomics and Biotechnology)

Abstract

Genotyping is a key step in molecular breeding. Due to its cost-effectiveness, accuracy, and flexibility, genotyping by target sequencing (GBTS) has become a preferred technology for medium-density genotyping. In this study, a new GBTS array for fresh edible maize was developed using resequencing data from 477 lines. The array contains 5759 SNPs evenly distributed across the maize genome, with average minor allele frequency (MAF) and polymorphism information content (PIC) values of 0.40 and 0.36, respectively. These SNPs are closely associated with 1566 functional genes. Cluster analysis of 198 maize lines based on the GBTS array was consistent with their pedigree relationships. Furthermore, 277 fresh waxy maize lines were genotyped and used for genomic selection analyses of hundred-kernel weight, kernel length, and kernel width. Comparative evaluation of different models indicated that Ridge Regression Best Linear Unbiased Prediction (rrBLUP) was the optimal model, with prediction accuracies of 0.33, 0.64, and 0.36, respectively. Additional analyses using different marker densities based on the rrBLUP model showed that prediction accuracy did not increase when the number of markers exceeded 2000, indicating that this array provides sufficient marker density for genetic analysis and genomic selection. Overall, this array provides a useful tool for genetic studies of fresh edible maize and facilitates the application of genomic selection in breeding programs.

1. Introduction

Maize (Zea mays L.) is one of the most important cereal crops worldwide and serves as a staple food, feed, and industrial raw material. Among its diverse types, fresh edible maize is widely consumed as a food and vegetable product and is favored by consumers worldwide. Fresh edible maize mainly includes sweet maize (Zea mays var. rugosa) and waxy maize (Zea mays var. ceratina) and sweet-waxy maize. Sweet maize appears to have originated in northern Mexico and was widely cultivated about 1000 years ago [1]. Sweet maize is a rich source of dietary fiber, folate, and provitamin A carotenoids, contributing to improved human nutrition and health. Sweet maize is characterized by mutations in key genes of the starch biosynthesis pathway (e.g., shrunken2 and su1), which disrupt starch accumulation in the endosperm and consequently increase the levels of soluble sugars in the kernels [2,3,4]. Waxy maize appears to have originated in southwest China several hundred years ago and is distinguished from other maize by a recessive mutation in the WAXY gene that reduces synthesis of amylose in the kernel [5,6]. Compared with field maize, breeding of fresh edible maize focuses more on improving eating quality and nutritional quality traits. However, traditional phenotypic selection is inefficient for the rapid improvement of these complex traits. In addition, transgenic breeding faces limitations due to low public acceptance. Therefore, molecular breeding provides a more efficient and feasible strategy for the genetic improvement of fresh edible maize.
The development of molecular markers plays a crucial role in advancing molecular breeding of fresh edible maize. The types of genetic variations suitable for developing molecular markers are primarily divided into two categories. One category comprises variations based on polymorphism length, including Insertion and Deletion (Indel) and Simple Sequence Repeats (SSR). The other category consists of Single-Nucleotide Polymorphisms (SNP), which involve single nucleotide substitutions or insertions and deletions. SNP markers have been widely applied in genetic research due to their extensive distribution across the genome, strong genetic stability, and suitability for automated detection. KASPar, array-based, and the next-generation sequencing (NGS) technology are the major genotyping platforms. KASPar is a single-plex SNP genotyping platform and highly suitable for low-density SNP genotyping and has been widely used in maize for genetic studies [7,8]. The array-based genotyping technology fixes the SNP identity probe on the array and is helpful for the cross-project comparison [9]. The array-based genotyping technology suffers from limited flexibility, as the addition of new SNP markers requires redesigning. NGS technologies have been widely used to identify genome-wide and high-density genotypes. However, large-scale sequencing of breeding populations remains costly, making it impractical for routine applications in breeding programs.
Genotyping by target sequencing (GBTS) is a method for sequencing and identifying genotypes at specified target sites using multiplex PCR or probe capture [10]. GBTS has been widely applied in plant and animal genetic research due to its advantages of high throughput, high accuracy, lower cost, and greater flexibility [11]. GBTS was first reported in maize with development of four arrays containing 1 K, 5 K, 10 K, and 20 K SNP markers [10]. Subsequently, GBTS has been successfully applied to major crops such as rice, wheat, and soybean [12,13,14,15].
Genomic selection (GS), first proposed by Meuwissen, refers to the use of both genotypic and phenotypic data from a training population to estimate the genetic effects of genome-wide markers [16]. These estimated effects are then applied to the genotypic data of individuals with unknown phenotypes to predict their genomic estimated breeding values (GEBVs), which can be used for breeding selection. As one of the core technologies in the Breeding 4.0 era, GS has been successfully applied to major cereal crops such as maize, rice, and wheat [11,17,18]. GBTS is considered one of the most suitable genotyping technologies for GS applications.
A large number of GS models have been developed, which can be broadly classified into three categories: linear models, Bayesian methods, and machine learning approaches [19]. Linear models, such as genomic best linear unbiased prediction (GBLUP) [20] and ridge regression best linear unbiased prediction (rrBLUP) [21], are widely used statistical approaches. These models estimate phenotypic values of individuals by constructing a genomic relationship matrix (G) based on genome-wide markers, offering advantages of stable prediction accuracy and high computational efficiency. Bayesian methods, including BayesA, BayesB, BayesC, Bayesian least absolute shrinkage and selection operator (Bayesian LASSO), and Bayesian Ridge Regression (BRR), assume that marker effects follow specific prior distributions and estimate their effects accordingly [22,23]. These models are known for their flexible model design and relatively high prediction accuracy. However, both linear models and their corresponding Bayesian approaches are primarily based on the assumption of additive marker effects, which limits their ability to capture complex non-linear relationships between genotype and phenotype. In contrast, machine-learning-based GS models, such as reproducing kernel Hilbert space (RKHS) [24], support vector machines (SVM) [25], and kernel ridge regression (KRR) [26], can effectively model non-linear relationships between genotypes and phenotypes, thereby improving prediction performance in complex traits. Nevertheless, these approaches often suffer from limited interpretability, making it difficult to quantify the contribution of individual markers.
Unlike field maize, which is primarily bred for grain yield and disease resistance, fresh edible maize breeding mainly targets eating quality and nutritional quality, resulting in a distinct genetic background compared with field maize. Although numerous GBTS arrays have been developed for maize [10,27,28], they were all designed based on the genetic background of field maize, thereby limiting the progress of molecular breeding in fresh edible maize. In this study, we designed a medium-density GBTS array from 477 diverse germplasm resources currently used in a fresh edible maize breeding program and evaluated its application in genetic analysis and GS for fresh edible maize. This study will be useful for accelerating genetic gain in fresh edible maize.

2. Results

2.1. Characteristics of Resequencing Data and Variants

In this study, whole-genome resequencing was performed on 477 maize inbred lines, generating a total of 6797.91 Gb of raw reads. The raw read size of each sample ranged from 9.48 to 22.97 Gb, with an average of 14.25 Gb. After quality control, 6733.29 Gb of high-quality clean reads were retained, accounting for 99.05% of the raw data. The clean reads size of each sample ranged from 9.38 to 22.85 Gb, averaging 14.12 Gb. Of these, the size of 442 samples (92.66% of the total) was more than 11 Gb of clean reads. Based on the estimated maize reference genome size of 2.2 Gb, this data size corresponds to a sequencing depth of ≥5×. Notably, although low-depth sequencing may reduce the sensitivity for detecting heterozygous SNPs, the large sample size and stringent filtering criteria effectively mitigate this limitation, thereby ensuring the overall reliability of SNP discovery. Detailed information on the whole-genome resequencing data is provided in Table S1.
Following variant calling, a total of 101,471,606 raw SNPs were identified. These variants were broadly distributed across all ten maize chromosomes, providing a rich source of genetic variation for the development of a medium-density fresh edible maize genotyping chip. In addition, these high-density genotype data provide valuable resources for genetic studies, such as genome-wide association studies. After quality control analysis, a total of 1,266,886 high-quality SNPs was obtained.

2.2. Characteristics of the Fresh Edible Maize 5K SNP Array

Following probe design and evaluation, a total of 5759 SNPs were retained from all high-quality SNPs for the medium-density genotyping platform (Table S2). These SNPs were uniformly distributed across the ten chromosomes of the maize genome (Figure 1A). The number of SNPs on each chromosome ranged from 365 on chromosome 7 to 856 on chromosome 5, with a mean of 575.90 SNPs per chromosome (Figure 2, Table 1). Additionally, the average distance between adjacent SNPs on each chromosome ranged from 261.57 Kb on chromosome 5 to 499.68 Kb on chromosome 7, with a mean of 365.75 Kb.
The SNP mutation type analysis indicated that 66.22% of the SNPs were transition (A/G and T/C), while 19.29% were A/C and T/G, and the remaining 14.48% were A/T and G/C (Figure 1B). The number of SNPs with A/G and T/C allelic types was significantly higher than that of other mutation types. The distribution of SNP mutation types in this chip is consistent with the transition bias observed in existing low- and medium-density genotyping platforms in maize [29,30], further supporting the rationality of SNP marker selection in this chip.
The MAF of these SNPs across the 477 fresh edible maize lines ranged from 0.29 to 0.50, with a mean of 0.40, whereas the PIC ranged from 0.33 to 0.38, with a mean of 0.36 (Figure 1C). The missing rate of these SNPs across the 477 fresh edible maize lines ranged from 0 to 1.81%, with a mean of 0.12%, and 5712 SNPs (99.18%) had a missing rate < 1%. Additionally, functional annotation showed that most SNPs were located in exonic regions (Figure 1D). Of the 5759 SNPs, 5445 were located in exonic regions, 277 in intronic regions, and only 37 in intergenic regions. Notably, the relatively high proportion of exonic SNPs can be attributed to the prioritization of functionally relevant variants during marker selection, as these loci are more likely to be associated with phenotypic variation and thus are advantageous for downstream genomic prediction. These 5759 SNPs were associated with 1566 genes. Specifically, 5719 SNPs were located in the gene regions of 1332 genes, 291 SNPs were within 2 kb upstream of 96 genes, and 566 SNPs were within 2 kb downstream of 172 genes. The gene function of all 1566 genes is provided in Table S3.
These results highlighted the high quality of the 5759 selected SNP markers, which showed consistent genotyping performance, low levels of missing data, and appropriate minor allele frequencies, making them suitable for genetic analyses and the development of a medium-density fresh edible maize genotyping chip.

2.3. The 5K SNP Array for Genetic Analysis

To test the developed SNP array, 198 randomly selected maize lines including fresh edible sweet maize lines, fresh edible waxy maize lines, and field maize lines were used for genotype detection (Table S4). The call rate of individual samples ranged from 90.54% to 98.45%, with a mean of 96.64%. Additionally, 164 samples (83.33%) exhibited a call rate greater than 95% (Figure 2A). The array exhibited a high sample call rate, indicating that it is robust and reliable for genotyping diverse fresh edible maize germplasm resources. After filtering, a total of 4829 high-quality SNPs were retained for further genetic analyses.
The result of principal component analysis (PCA) in all of the 198 maize lines is shown in Figure 2A,B. The first three principal components explained 54.48% of the genetic variations, and the genetic variant explained by PC1, PC2, and PC3 was 23.72%, 20.08%, and 10.69%, respectively. A small number of sweet maize inbred lines and waxy maize inbred lines could not be clearly separated from grain maize, likely due to germplasm introgression between fresh edible maize and grain maize populations during the breeding process. Overall, a clear genetic divergence was observed among the sweet maize inbred lines, waxy maize inbred lines, and field maize inbred lines.
Further phylogenetic analyses based on the genetic distance matrix clearly separated the 198 maize lines into three groups (Figure 2C). Group 1 consisted entirely of sweet maize inbred lines, Group 2 was mainly composed of waxy maize inbred lines, and Group 3 was predominantly made up of field maize inbred lines. The clustering result was highly consistent with their pedigree classifications. Notably, four sweet maize inbred lines and ten field maize lines were assigned to Group 2, and one waxy maize line was assigned to Group 3. These results indicate that, within the selected validation population, the improvement of waxy maize primarily involved the use of grain and sweet maize germplasm, which is consistent with the actual breeding practices of our team.
The optimal K was selected using the “elbow method” based on the cross-validation (CV) error. The CV error decreased substantially from K = 1 to 3, after which the rate of decline diminished and the error fluctuated without further improvement for K ≥ 4 (Table S5). The statistical indication of K = 3 as the elbow point is consistent with prior biological knowledge, as the 198 maize inbred lines are derived from three known breeding populations. Consequently, K = 3 was determined as the optimal number of ancestral populations. In the sweet maize group, the proportion of Ancestry 3 was significantly higher than that of the other two ancestral components. In contrast, Ancestry 1 predominated in the waxy maize group, whereas Ancestry 2 was more abundant than the other components in the field maize group (Figure 2D). These results indicate that Ancestry 2 is the primary factor underlying the genetic differentiation between fresh edible maize and field maize.
Collectively, these results indicate that the 5K SNP array for fresh edible maize can be used for genetic analysis. Additionally, the array could also be applied to field maize.

2.4. The 5K SNP Array for GS

To evaluate the predictive performance of the 5K SNP array, eight different statistical models were used to predict three agronomic traits within the fresh edible waxy maize population comprising 277 inbred lines. Results indicated that prediction accuracy varied across traits (Figure 3). For 100-kernel weight, no significant differences were detected among the eight statistical models (p > 0.05). However, for kernel length and kernel width, significant differences in prediction accuracy were observed among models (p < 0.05). Notably, rrBLUP consistently achieved the highest prediction accuracy across traits, significantly outperforming the other seven models, highlighting its robustness and suitability for genomic prediction in this maize population.
Different marker densities were evaluated to assess their effects on the prediction accuracy of kernel-related traits in fresh edible waxy maize using the rrBLUP model. As the number of markers increased, prediction accuracy initially improved but reached a plateau when approximately 2000 markers were included. At this marker density, the average prediction accuracies for hundred-kernel weight, kernel length, and kernel width were 0.32, 0.45, and 0.35, respectively. Notably, no further significant increase in prediction accuracy was observed for any of the three traits when additional markers were included (Figure 4). These results indicate that a marker set of approximately 2000 SNPs is sufficient to achieve stable genomic prediction for kernel-related traits in fresh edible waxy maize using the present SNP array.
GS analysis in fresh edible maize indicated that the array is capable of delivering high-quality genotype data in a timely manner, providing a valuable tool for accelerating the improvement of fresh edible maize.

3. Discussion

In maize, a series of GBTS arrays have been developed and used in genetic studies, QTL mapping, and GS [10,27,28,29]. However, most GBTS arrays developed for maize are based on the genetic background of field maize. Since the breeding objectives of fresh edible maize differ from those of field maize, its genetic background may also differ considerably [31]. As one of the core strategies in the era of “Breeding 4.0,” GS requires large-scale genotyping for the development and selection of pure lines, making cost control a critical consideration [32,33,34]. In maize, most agronomic traits are complex quantitative traits controlled by numerous genes with small effects [35]. Therefore, low-density genotyping approaches (e.g., InDel, SSR) are often insufficient to achieve satisfactory prediction performance. In contrast, high-density genotyping platforms (e.g., whole-genome resequencing, GBS) provide comprehensive genomic information but are associated with high costs when applied to large breeding populations. Consequently, medium-density marker platforms represent an optimal compromise between cost efficiency and prediction performance, making them particularly suitable for large-scale GS applications in breeding programs. In the present study, a total of 5759 SNP markers were selected for development of a GBTS array for fresh edible maize. These SNPs are distributed across all ten maize chromosomes, which enhances the power of genetic analysis and genomic selection [34]. These markers exhibited relatively high levels of polymorphism, with mean MAF and PIC values of 0.40 and 0.36, respectively. High-polymorphism markers are essential for the effective implementation of GS, as they enhance the ability to capture genome-wide genetic variation and improve the accuracy of marker effect estimation [36]. The elevated MAF and PIC values observed here indicate strong allelic diversity and high discriminatory power within the fresh edible maize population. Importantly, all 5759 SNPs are biallelic loci, which facilitates their direct conversion into KASP markers [8]. This characteristic provides substantial flexibility for practical breeding applications, particularly in low-density genotyping scenarios such as DNA fingerprint construction, germplasm identification, and MAS. In terms of functional annotation, 5445 SNPs were located within exonic regions, involving 1556 annotated functional genes. SNPs situated in coding regions are more likely to be directly associated with functional variation or tightly linked to causal loci, thereby increasing their biological relevance [37,38,39,40]. Collectively, this SNP panel demonstrates strong genetic diversity, broad genome coverage, and functional significance, making it suitable for GS breeding, genetic diversity analysis, QTL mapping, and fingerprinting applications.
We applied the developed SNP array to genetic analyses of 198 maize inbred lines. The results demonstrated that the array performed robustly in PCA, phylogenetic tree construction, and population structure analysis, effectively capturing genetic relationships and population stratification. These findings suggest that the developed SNP array has the potential to be applied in the classification of heterotic groups in maize breeding, although further validation is required. Although the fresh edible waxy maize population used in this study was primarily composed of Chinese germplasm, the panel of 198 lines also included 76 field maize inbred lines originating from diverse global sources. This diversity suggests that the developed SNP array may have broader applicability beyond Chinese germplasm and could be useful across different ecological regions. However, further validation using more diverse populations is still required. We further applied the GBTS-based SNP array to genome-wide selection for kernel-related traits in 277 fresh edible waxy maize inbred lines and systematically compared eight statistical models. For HKW, no significant differences in prediction accuracy were detected among the models, suggesting that alternative assumptions regarding marker effect distributions did not substantially influence predictive performance for this trait. In contrast, for KL and KW, the rrBLUP model showed significantly higher prediction accuracy than the other seven models. These results indicate that KL and KW are typical highly polygenic traits primarily controlled by numerous loci with small additive effects, whereas the genetic architecture of HKW may be relatively more complex. The superior performance of rrBLUP, which assumes homogeneous marker variances and applies uniform shrinkage, suggests that additive genetic effects dominate the variation in kernel morphological traits in this population. Therefore, rrBLUP appears to be a robust and efficient model for genomic prediction in fresh waxy maize breeding programs.
To further evaluate the impact of marker density on predictive ability, different subsets of SNP markers were tested using the rrBLUP model. Prediction accuracy increased with marker number at lower densities but reached a plateau when approximately 2000 markers were included. Beyond this threshold, no significant improvement in predictive performance was observed. This plateau effect indicates that approximately 2000 well-distributed SNP markers are sufficient to capture most of the additive genetic variance in this population. This result is consistent with previous GS studies in maize [41]. Therefore, the medium-density array seems sufficient for the implementation of genomic selection in fresh waxy maize breeding, offering a cost-effective and scalable strategy for routine genomic selection programs.
The moderate prediction accuracies observed for hundred-kernel weight, kernel length, and kernel width are consistent with the quantitative and polygenic nature of these traits, which are typically controlled by many loci with small effects. Under such a genetic structure, the maximum achievable prediction accuracy using linear genomic prediction models may be inherently constrained. In addition, these traits are sensitive to environmental variation and genotype-by-environment interactions, and limited multi-environment phenotypic data may further reduce effective prediction accuracy. Therefore, the prediction accuracies reported in this study likely reflect both biological complexity and practical limitations of the available training data. Nevertheless, the developed GBTS5K platform still provides sufficient predictive ability for relative ranking and early-stage selection of breeding materials, demonstrating its practical value for genomic selection in fresh edible maize. Future improvements may be achieved by expanding the training population, incorporating multi-environment phenotypic data, and evaluating models that better capture non-additive genetic effects and genotype-by-environment interactions.

4. Materials and Methods

4.1. Plant Materials, Sampling, and Resequencing

477 diverse edible maize inbred lines were collected from the Maize Center of the Shanghai Academy of Agricultural Sciences for the development of a medium-density genotyping platform, including 170 sweet maize inbred lines, 286 waxy maize inbred lines, and 21 sweet-waxy maize inbred lines. The inbred lines were cultivated at the Zhuanghang Experimental Station of the Shanghai Academy of Agricultural Sciences. Genomic DNA was extracted from young leaf tissues of the inbred lines using the CTAB method. The integrity of the extracted DNA was assessed by 0.8% agarose gel electrophoresis, while the concentration and purity were measured using a NanoDrop spectrophotometer. Samples that passed the quality control were used for subsequent resequencing. The Illumina sequencing libraries were constructed with a size of 350 bp, and whole-genome resequencing was performed based on Illumina sequencing platform with paired 150 bp at NovoGene company (Beijing, China). Each sample was sequenced at a preset depth of more than 5×. The raw data from the sequencing platform were filtered using the FASTP (version: 0.23.2) software [42], retaining reads with a Phred quality score ≥ 20 (Q20). The clean data were mapped to the maize reference genome B73v4 (https://download.maizegdb.org/Zm-B73-REFERENCE-GRAMENE-4.0/Zm-B73-REFERENCE-GRAMENE-4.0.fa.gz, accessed on 1 May 2025) using bwa (version: 0.7.17) software with default parameters [43,44]. The variant calling was conducted using Freebayes (version: 0.9.21) software with default parameters [45].

4.2. Selection of SNP for Development of the Chip Array

The raw SNPs derived from the resequencing data were initially filtered using bcftools (version: 1.6) according to the following criteria: sequencing depth of between 2 and 50, genotyping quality ≥ 20, sample missing rate ≤ 0.05, biallelic loci only, and MAF ≥ 0.2. To account for LD and reduce SNP redundancy, SNPs were further pruned using PLINK (version: 1.9) with the parameter of “--indep-pairwise 50 5 0.2”, retaining a set of approximately independent SNPs for downstream analyses. High-quality SNPs located within genes and their flanking 2 kb regions were selected for GBTS probe design. The probes were designed using GenoBaits Probe Designer with a length of 110 nt and a GC content ranging from 30% to 70% [10]. The probe sequences were aligned to the reference genome B73v4 using BLAST (version: 2.12.0) to remove non-specific sequences. A two-cross-covered-probe strategy was employed for each SNP site. A random set of 6000 probe sequences was outsourced to MOLBREEDING Biotechnology Co., Ltd. (Shijiazhuang, Hebei, China) for probe synthesis, sample–probe hybridization, and Illumina sequencing.
MAF and missing rate were calculated using VCFtools (version: 0.1.16). PIC was calculated using the following formula:
P I C = 1 p 2 + q 2 2 p 2 q 2
where p and q are the population frequencies of the allele1 and the allele2.

4.3. Genotyping of 198 Maize Inbred Lines

To evaluate the developed GBTS5K chip, a panel of 198 maize inbred lines was genotyped. This validation population included 50 sweet maize inbred lines, 72 waxy maize inbred lines, and 76 field maize inbred lines. All 198 lines were obtained from the Maize Center of the Shanghai Academy of Agricultural Sciences. The plants were grown in a greenhouse until the five-leaf stage, and leaf tissues were sampled and flash-frozen in liquid nitrogen, and DNA was extracted via the CTAB method. The extracted DNA samples were genotyped using the GBTS5K chip following the procedure described in previous studies [46]. In this step, genomic DNA samples were randomly fragmented into 300–350 bp fragments, followed by hybridization capture using probes from the GBTS5K chip. The captured fragments were then used for library construction and sequenced on an Illumina platform with paired-end 150 bp. Raw sequencing data were subjected to quality control using fastp, followed by alignment to the reference genome using Bwa, and SNP calling was performed using GATK4 Best Practices pipeline with the specified genomic region. The genotype of 198 lines was filtered with quality of genotype ≥ 20, sample missing rate ≤ 0.05 and MAF ≥ 0.05. The high-quality genotype data was used to perform the principal component analysis (PCA) using PLINK (version: 1.9) with the parameter “--pca 10” [47]. The variance explained by each principal component was estimated by dividing its eigenvalue by the sum of the eigenvalues of all principal components. The pairwise genetic distances among the 198 maize inbred lines were calculated using TASSEL (version: 5.2.90), and a phylogenetic tree was subsequently constructed based on the genetic distance matrix using the neighbor-joining (NJ) method [48]. Population structure analyses of 198 maize lines was performed using Admixture (version: 1.3.0). The optimal K was determined using the “elbow method,” defined as the point where the decrease in cross-validation (CV) error markedly slowed. The results of PCA and population structure were visualized using the ggplot2 package in R (version: 4.5.2), while the phylogenetic tree was visualized with the ggtree package.

4.4. Genomic Prediction

Phenotypic data for 100-kernel weight, kernel length, and kernel width from 277 fresh edible waxy maize inbred lines, obtained from our previous study, were used for GS [49]. The population were genotyped using the array developed in this study. GS was performed with eight statistical models, including rrBLUP, GBLUP, BayesA, BayesB, BayesC, Bayesian LASSO, BRR, and RKH. These statistical models were performed using the two R packages, rrBLUP (https://CRAN.R-project.org/package=rrBLUP, accessed on 1 May 2025) with the default parameters and BGLR (https://CRAN.R-project.org/package=BGLR, accessed on 1 May 2025) with the parameters of “iterations = 12,000, burn-in = 3000”. The prediction accuracy was assessed using the five-fold cross-validation method with 100 replicates. The dataset was randomly divided into five subsets, with four subsets used as the training population and the remaining subset as the validation population in each iteration. This process was repeated 100 times to ensure robust and stable estimates of prediction accuracy. The Pearson correlation coefficient between the observed phenotypic values and the genomic estimated breeding values (GEBVs) in each cross-validation replicate served as a measure of prediction accuracy. Random subsets consisting of 100, 300, 500, 1000, 2000, 3000, 4000, and 5000 markers were sampled from the array to perform GS, aiming to assess whether the marker density is adequate for GS in fresh edible maize.

Supplementary Materials

The following supporting information can be downloaded at: https://www.mdpi.com/article/10.3390/plants15091288/s1, Table S1: The resequencing information of 477 fresh edible maize germplasm; Table S2: The information of 5759 SNPs for the GBTS array; Table S3: The information of gene related to the 5759 SNP markers; Table S4: The genotyping data of 198 maize lines; Table S5: Cross-validation (CV) errors for K = 1 to 10 in ADMIXTURE analysis.

Author Contributions

Y.G. and H.Z. designed the research; J.Q. and D.Y. performed data collection and analysis with help from W.G. and Y.Z.; K.L., H.W. and P.S. conducted field management and provided assistance; J.Q. wrote the manuscript; F.S.V., A.Z. and X.Z. were involved in the discussion and revision of the manuscript. All authors have read and agreed to the published version of the manuscript.

Funding

This research was funded by the Shanghai Agricultural Science and Technology Innovation Program (Grant No. T2025317) and AI-Empowered Agriculture Project of Shanghai Academy of Agricultural Sciences (Grant No. AI-B-2025-001).

Data Availability Statement

The original contributions presented in this study are included in the article. Further inquiries can be directed to the corresponding authors.

Acknowledgments

The authors sincerely thank the CIMMYT-China Special Maize Research Center (CCSMRC) for providing germplasm resources and continuous support.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Hendry, G.W. Archaeological Evidence Concerning the Origin of Sweet Maize 1. Agron. J. 1930, 22, 508–514. [Google Scholar] [CrossRef]
  2. Rahman, A.; Wong, K.-S.; Jane, J.-L.; Myers, A.M.; James, M.G. Characterization of SU1 Isoamylase, a Determinant of Storage Starch Structure in Maize. Plant Physiol. 1998, 117, 425–435. [Google Scholar] [CrossRef]
  3. Tsai, C.-Y.; Nelson, O.E. Starch-Deficient Maize Mutant Lacking Adenosine Diphosphate Glucose Pyrophosphorylase Activity. Science 1966, 151, 341–343. [Google Scholar] [CrossRef]
  4. Tracy, W.F. History, Genetics, and Breeding of Supersweet (Shrunken2) Sweet Corn. In Plant Breeding Reviews; Janick, J., Ed.; John Wiley & Sons, Inc.: Oxford, UK, 1996; pp. 189–236. ISBN 9780470650073. [Google Scholar]
  5. Wessler, S.R.; Varagona, M.J. Molecular Basis of Mutations at the Waxy Locus of Maize: Correlation with the Fine Structure Genetic Map. Proc. Natl. Acad. Sci. USA 1985, 82, 4177–4181. [Google Scholar] [CrossRef]
  6. Tian, M.; Tan, G.; Liu, Y.; Rong, T.; Huang, Y. Origin and Evolution of Chinese Waxy Maize: Evidence from the Globulin-1 Gene. Genet. Resour. Crop Evol. 2009, 56, 247–255. [Google Scholar] [CrossRef]
  7. Nair, S.K.; Babu, R.; Magorokosho, C.; Mahuku, G.; Semagn, K.; Beyene, Y.; Das, B.; Makumbi, D.; Kumar, P.L.; Olsen, M.; et al. Fine Mapping of Msv1, a Major QTL for Resistance to Maize Streak Virus Leads to Development of Production Markers for Breeding Pipelines. Theor. Appl. Genet. 2015, 128, 1839–1854. [Google Scholar] [CrossRef] [PubMed]
  8. Semagn, K.; Babu, R.; Hearne, S.; Olsen, M. Single Nucleotide Polymorphism Genotyping Using Kompetitive Allele Specific PCR (KASP): Overview of the Technology and Its Application in Crop Improvement. Mol. Breed. 2014, 33, 1–14. [Google Scholar] [CrossRef]
  9. Rasheed, A.; Hao, Y.; Xia, X.; Khan, A.; Xu, Y.; Varshney, R.K.; He, Z. Crop Breeding Chips and Genotyping Platforms: Progress, Challenges, and Perspectives. Mol. Plant 2017, 10, 1047–1064. [Google Scholar] [CrossRef] [PubMed]
  10. Guo, Z.; Wang, H.; Tao, J.; Ren, Y.; Xu, C.; Wu, K.; Zou, C.; Zhang, J.; Xu, Y. Development of Multiple SNP Marker Panels Affordable to Breeders through Genotyping by Target Sequencing (GBTS) in Maize. Mol. Breed. 2019, 39, 37. [Google Scholar] [CrossRef]
  11. Wang, C.; Zhang, D.; Ma, Y.; Zhao, Y.; Liu, P.; Li, X. WheatGP, a Genomic Prediction Method Based on CNN and LSTM. Briefings Bioinform. 2025, 26, bbaf191. [Google Scholar] [CrossRef]
  12. Liu, S.; Xiang, M.; Wang, X.; Li, J.; Cheng, X.; Li, H.; Singh, R.P.; Bhavani, S.; Huang, S.; Zheng, W.; et al. Development and Application of the GenoBaits WheatSNP16K Array to Accelerate Wheat Genetic Research and Breeding. Plant Commun. 2025, 6, 101138. [Google Scholar] [CrossRef]
  13. Lee, C.; Cheon, K.-S.; Shin, Y.; Oh, H.; Jeong, Y.-M.; Jang, H.; Park, Y.-C.; Kim, K.-Y.; Cho, H.-C.; Won, Y.-J.; et al. Development and Application of a Target Capture Sequencing SNP-Genotyping Platform in Rice. Genes 2022, 13, 794. [Google Scholar] [CrossRef]
  14. Liu, Y.; Liu, S.; Zhang, Z.; Ni, L.; Chen, X.; Ge, Y.; Zhou, G.; Tian, Z. GenoBaits Soy40K: A Highly Flexible and Low-Cost SNP Array for Soybean Studies. Sci. China Life Sci. 2022, 65, 1898–1901. [Google Scholar] [CrossRef] [PubMed]
  15. Yang, Q.; Zhang, J.; Shi, X.; Chen, L.; Qin, J.; Zhang, M.; Yang, C.; Song, Q.; Yan, L. Development of SNP Marker Panels for Genotyping by Target Sequencing (GBTS) and Its Application in Soybean. Mol. Breed. 2023, 43, 26. [Google Scholar] [CrossRef] [PubMed]
  16. Meuwissen, T.H.E.; Hayes, B.J.; Goddard, M.E. Prediction of Total Genetic Value Using Genome-Wide Dense Marker Maps. Genetics 2001, 157, 1819–1829. [Google Scholar] [CrossRef] [PubMed]
  17. Cui, Y.; Li, R.; Li, G.; Zhang, F.; Zhu, T.; Zhang, Q.; Ali, J.; Li, Z.; Xu, S. Hybrid Breeding of Rice via Genomic Selection. Plant Biotechnol. J. 2020, 18, 57–67. [Google Scholar] [CrossRef]
  18. Alemu, A.; Åstrand, J.; Montesinos-López, O.A.; Sánchez, J.I.Y.; Fernández-Gónzalez, J.; Tadesse, W.; Vetukuri, R.R.; Carlsson, A.S.; Ceplitis, A.; Crossa, J.; et al. Genomic Selection in Plant Breeding: Key Factors Shaping Two Decades of Progress. Mol. Plant 2024, 17, 552–578. [Google Scholar] [CrossRef] [PubMed]
  19. Wang, X.; Xu, Y.; Xu, Y.; Xu, C. Research Progress in Genomic Selection for Crop Breeding. Biotechnol. Bull. 2023, 40, 1–13. [Google Scholar]
  20. VanRaden, P. Efficient Methods to Compute Genomic Predictions. J. Dairy Sci. 2008, 91, 4414–4423. [Google Scholar] [CrossRef]
  21. Piepho, H.P. Ridge Regression and Extensions for Genome-wide Selection in Maize. Crop Sci. 2009, 49, 1165–1176. [Google Scholar] [CrossRef]
  22. Pérez, P.; de los Campos, G. Genome-Wide Regression and Prediction with the BGLR Statistical Package. Genetics 2014, 198, 483–495. [Google Scholar] [CrossRef] [PubMed]
  23. Habier, D.; Fernando, R.L.; Kizilkaya, K.; Garrick, D.J. Extension of the Bayesian Alphabet for Genomic Selection. BMC Bioinform. 2011, 12, 186. [Google Scholar] [CrossRef]
  24. Campos, G.D.L.; Gianola, D.; Rosa, G.J.M.; Weigel, K.A.; Crossa, J. Semi-Parametric Genomic-Enabled Prediction of Genetic Values Using Reproducing Kernel Hilbert Spaces Methods. Genet. Res. 2010, 92, 295–308. [Google Scholar] [CrossRef]
  25. Kasnavi, S.A.; Afshar, M.A.; Shariati, M.M.; Kashan, N.E.J.; Honarvar, M. Performance Evaluation of Support Vector Machine (SVM)-Based Predictors in Genomic Selection. Indian J. Anim. Sci. 2017, 87, 1226–1231. [Google Scholar] [CrossRef]
  26. Exterkate, P.; Groenen, P.J.; Heij, C.; van Dijk, D. Nonlinear Forecasting with Many Predictors Using Kernel Ridge Regression. Int. J. Forecast. 2016, 32, 736–753. [Google Scholar] [CrossRef]
  27. Guo, Z.; Yang, Q.; Huang, F.; Zheng, H.; Sang, Z.; Xu, Y.; Zhang, C.; Wu, K.; Tao, J.; Prasanna, B.M.; et al. Development of High-Resolution Multiple-SNP Arrays for Genetic Analyses and Molecular Breeding through Genotyping by Target Sequencing and Liquid Chip. Plant Commun. 2021, 2, 100230. [Google Scholar] [CrossRef]
  28. Ma, J.; Cao, Y.; Wang, Y.; Ding, Y. Development of the Maize 5.5K Loci Panel for Genomic Prediction through Genotyping by Target Sequencing. Front. Plant Sci. 2022, 13, 972791. [Google Scholar] [CrossRef]
  29. Wang, Z.; Zhang, H.; Ye, W.; Han, Y.; Li, H.; Zhou, Z.; Li, C.; Zhang, X.; Zhang, J.; Chen, J.; et al. Development of a FER0.4K SNP Array for Genomic Predication of Fusarium Ear Rot Resistance in Maize. Crop J. 2025, 13, 996–1002. [Google Scholar] [CrossRef]
  30. Qu, J.; Chassaigne-Ricciulli, A.A.; Fu, F.; Yu, H.; Dreher, K.; Nair, S.K.; Gowda, M.; Beyene, Y.; Makumbi, D.; Dhliwayo, T.; et al. Low-Density Reference Fingerprinting SNP Dataset of CIMMYT Maize Lines for Quality Control and Genetic Diversity Analyses. Plants 2022, 11, 3092. [Google Scholar] [CrossRef]
  31. Li, C.; Li, Z.; Lu, B.; Shi, Y.; Xiao, S.; Dong, H.; Zhang, R.; Liu, H.; Jiao, Y.; Xu, L.; et al. Large-Scale Metabolomic Landscape of Edible Maize Reveals Convergent Changes in Metabolite Differentiation and Facilitates Its Breeding Improvement. Mol. Plant 2025, 18, 619–638. [Google Scholar] [CrossRef]
  32. Voss-Fels, K.P.; Cooper, M.; Hayes, B.J. Accelerating Crop Genetic Gains with Genomic Selection. Theor. Appl. Genet. 2019, 132, 669–686. [Google Scholar] [CrossRef] [PubMed]
  33. Crossa, J.; Pérez-Rodríguez, P.; Cuevas, J.; Montesinos-López, O.; Jarquín, D.; de los Campos, G.; Burgueño, J.; González-Camacho, J.M.; Pérez-Elizalde, S.; Beyene, Y.; et al. Genomic Selection in Plant Breeding: Methods, Models, and Perspectives. Trends Plant Sci. 2017, 22, 961–975. [Google Scholar] [CrossRef] [PubMed]
  34. Heffner, E.L.; Sorrells, M.E.; Jannink, J. Genomic Selection for Crop Improvement. Crop Sci. 2009, 49, 1–12. [Google Scholar] [CrossRef]
  35. Wallace, J.G.; Bradbury, P.J.; Zhang, N.; Gibon, Y.; Stitt, M.; Buckler, E.S. Association Mapping across Numerous Traits Reveals Patterns of Functional Variation in Maize. PLoS Genet. 2014, 10, e1004845. [Google Scholar] [CrossRef]
  36. Chang, L.-Y.; Toghiani, S.; Aggrey, S.E.; Rekaya, R. Increasing Accuracy of Genomic Selection in Presence of High Density Marker Panels through the Prioritization of Relevant Polymorphisms. BMC Genet. 2019, 20, 21. [Google Scholar] [CrossRef] [PubMed]
  37. Huang, X.; Han, B. Natural Variations and Genome-Wide Association Studies in Crop Plants. Annu. Rev. Plant Biol. 2014, 65, 531–551. [Google Scholar] [CrossRef]
  38. Bhoite, R.; Smith, R.; Bansal, U.; Dowla, M.; Bariana, H.; Sharma, D. Exome-Based New Allele-Specific PCR Markers and Transferability for Sodicity Tolerance in Bread Wheat (Triticum aestivum L.). Plant Direct 2023, 7, e520. [Google Scholar] [CrossRef]
  39. Watanabe, K.; Stringer, S.; Frei, O.; Mirkov, M.U.; de Leeuw, C.; Polderman, T.J.C.; van der Sluis, S.; Andreassen, O.A.; Neale, B.M.; Posthuma, D. A Global Overview of Pleiotropy and Genetic Architecture in Complex Traits. Nat. Genet. 2019, 51, 1339–1348. [Google Scholar] [CrossRef]
  40. Vij, S.; Tyagi, A.K. Emerging Trends in the Functional Genomics of the Abiotic Stress Response in Crop Plants: Review Article. Plant Biotechnol. J. 2007, 5, 361–380. [Google Scholar] [CrossRef]
  41. Song, J.; Liu, Y.; Guo, R.; Pacheco, A.; Muñoz-Zavala, C.; Song, W.; Wang, H.; Cao, S.; Hu, G.; Zheng, H.; et al. Exploiting Genomic Tools for Genetic Dissection and Improving the Resistance to Fusarium Stalk Rot in Tropical Maize. Theor. Appl. Genet. 2024, 137, 109. [Google Scholar] [CrossRef]
  42. Chen, S.; Zhou, Y.; Chen, Y.; Gu, J. Fastp: An Ultra-Fast All-in-One FASTQ Preprocessor. Bioinformatics 2018, 34, i884–i890. [Google Scholar] [CrossRef]
  43. Schnable, P.S.; Ware, D.; Fulton, R.S.; Stein, J.C.; Wei, F.; Pasternak, S.; Liang, C.; Zhang, J.; Fulton, L.; Graves, T.A.; et al. The B73 Maize Genome: Complexity, Diversity, and Dynamics. Science 2009, 326, 1112–1115. [Google Scholar] [CrossRef]
  44. Li, H.; Durbin, R. Fast and Accurate Short Read Alignment with Burrows—Wheeler Transform. Bioinformatics 2009, 25, 1754–1760. [Google Scholar] [CrossRef]
  45. Garrison, E.; Marth, G. Haplotype-Based Variant Detection from Short-Read Sequencing. arXiv 2012, arXiv:1207.3907. [Google Scholar]
  46. Xu, Y.; Yang, Q.N.; Zheng, H.J.; Xu, Y.F.; Sang, Z.Q.; Guo, Z.F.; Peng, H.; Zhang, C.; Lan, H.F.; Wang, Y.B.; et al. Genotyping by Target Sequencing (GBTS) and Its Applications. Sci. Agric. Sin. 2020, 53, 2983–3004. [Google Scholar] [CrossRef]
  47. Purcell, S.; Neale, B.; Todd-Brown, K.; Thomas, L.; Ferreira, M.A.R.; Bender, D.; Maller, J.; Sklar, P.; de Bakker, P.I.W.; Daly, M.J.; et al. PLINK: A Tool Set for Whole-Genome Association and Population-Based Linkage Analyses. Am. J. Hum. Genet. 2007, 81, 559–575. [Google Scholar] [CrossRef] [PubMed]
  48. Bradbury, P.J.; Zhang, Z.; Kroon, D.E.; Casstevens, T.M.; Ramdoss, Y.; Buckler, E.S. TASSEL: Software for Association Mapping of Complex Traits in Diverse Samples. Bioinformatics 2007, 23, 2633–2635. [Google Scholar] [CrossRef] [PubMed]
  49. Qu, J.; Yu, D.; Gu, W.; Bin Khalid, M.H.; Kuang, H.; Dang, D.; Wang, H.; Prasanna, B.; Zhang, X.; Zhang, A.; et al. Genetic Architecture of Kernel-Related Traits in Sweet and Waxy Maize Revealed by Genome-Wide Association Analysis. Front. Genet. 2024, 15, 1431043. [Google Scholar] [CrossRef]
Figure 1. Characteristics of the fresh edible maize 5K SNP array. (A) The distribution of 5k markers on ten maize chromosomes. (B) Types of variation among the 5k SNPs in the array. (C) Boxplot of MAF and PIC across the 477 fresh edible maize population for 5K SNP array. (D) The genomic distribution of SNPs across functional regions.
Figure 1. Characteristics of the fresh edible maize 5K SNP array. (A) The distribution of 5k markers on ten maize chromosomes. (B) Types of variation among the 5k SNPs in the array. (C) Boxplot of MAF and PIC across the 477 fresh edible maize population for 5K SNP array. (D) The genomic distribution of SNPs across functional regions.
Plants 15 01288 g001
Figure 2. The principal component analysis, NJ tree, and population structure in 198 maize lines. (A) indicates PC1 versus PC2. (B) indicates PC1 versus PC3. (C) indicates the circular NJ tree in the population. (D) indicates the population structure in the population (K = 3).
Figure 2. The principal component analysis, NJ tree, and population structure in 198 maize lines. (A) indicates PC1 versus PC2. (B) indicates PC1 versus PC3. (C) indicates the circular NJ tree in the population. (D) indicates the population structure in the population (K = 3).
Plants 15 01288 g002
Figure 3. Prediction accuracy of different statistical models based on 5759 SNP markers. Error bars represent the standard deviation (SD) across 100 cross-validation replicates. Different letters (a, b, and c) indicate significant differences among models based on one-way ANOVA followed by Tukey’s HSD test (p < 0.05).
Figure 3. Prediction accuracy of different statistical models based on 5759 SNP markers. Error bars represent the standard deviation (SD) across 100 cross-validation replicates. Different letters (a, b, and c) indicate significant differences among models based on one-way ANOVA followed by Tukey’s HSD test (p < 0.05).
Plants 15 01288 g003
Figure 4. Prediction accuracy across different marker densities using rrBLUP model. Error bars represent the standard deviation (SD) across 100 cross-validation replicates. Different letters (a, b, c, and d) indicate significant differences among models based on one-way ANOVA followed by Tukey’s HSD test (p < 0.05).
Figure 4. Prediction accuracy across different marker densities using rrBLUP model. Error bars represent the standard deviation (SD) across 100 cross-validation replicates. Different letters (a, b, c, and d) indicate significant differences among models based on one-way ANOVA followed by Tukey’s HSD test (p < 0.05).
Plants 15 01288 g004
Table 1. Distribution of 5759 SNPs across ten maize chromosomes.
Table 1. Distribution of 5759 SNPs across ten maize chromosomes.
ChromosomeLength
(Mb)
Number of SNPsDensity of SNP
(SNPs/Mb)
Average Distance
(Kb)
Chr1307.04754.002.46407.22
Chr2244.44735.003.01332.57
Chr3235.67625.002.65377.07
Chr4246.99576.002.33428.81
Chr5223.90856.003.82261.57
Chr6174.03524.003.01332.12
Chr7182.38365.002.00499.68
Chr8181.12380.002.10476.64
Chr9159.77454.002.84351.92
Chr10150.98490.003.25308.13
Mean210.63575.902.73365.75
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Qu, J.; Yu, D.; Gu, W.; Zhao, Y.; Li, K.; Wang, H.; Sun, P.; Vicente, F.S.; Zhang, X.; Zhang, A.; et al. Development of a Medium-Density Genotyping Platform to Accelerate Genetic Gain in Fresh Edible Maize. Plants 2026, 15, 1288. https://doi.org/10.3390/plants15091288

AMA Style

Qu J, Yu D, Gu W, Zhao Y, Li K, Wang H, Sun P, Vicente FS, Zhang X, Zhang A, et al. Development of a Medium-Density Genotyping Platform to Accelerate Genetic Gain in Fresh Edible Maize. Plants. 2026; 15(9):1288. https://doi.org/10.3390/plants15091288

Chicago/Turabian Style

Qu, Jingtao, Diansi Yu, Wei Gu, Yingjie Zhao, Kai Li, Hui Wang, Pingdong Sun, Felix San Vicente, Xuecai Zhang, Ao Zhang, and et al. 2026. "Development of a Medium-Density Genotyping Platform to Accelerate Genetic Gain in Fresh Edible Maize" Plants 15, no. 9: 1288. https://doi.org/10.3390/plants15091288

APA Style

Qu, J., Yu, D., Gu, W., Zhao, Y., Li, K., Wang, H., Sun, P., Vicente, F. S., Zhang, X., Zhang, A., Zheng, H., & Guan, Y. (2026). Development of a Medium-Density Genotyping Platform to Accelerate Genetic Gain in Fresh Edible Maize. Plants, 15(9), 1288. https://doi.org/10.3390/plants15091288

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop