DArTseq SNP-based genetic diversity and population structure of spider plant (Cleome gynandra L.) accession populations from Zimbabwe
Description
Understanding the genetic diversity and structure of the spider plant (Cleome gynandra L.) is essential for its potential improvement and conservation, although limited research has been reported in this area. In the present study, we characterized a set of 43 spider plant accessions collected from three agroecological zones (AEZ) in Zimbabwe, categorized by varying rainfall levels: high, moderate, and low. A total of 5,607 single nucleotide polymorphism (SNP) markers were identified using Diversity Array Technology sequencing (DArTseq), which facilitated the analysis of the population structure and genetic diversity of the spider plant. The SNP markers exhibited an average polymorphic information content (PIC) of 0.28, indicating highly informativeness. Population structure analysis, conducted through admixture and principal component analysis, consistently revealed three gene pools among the spider plant accessions, with average observed (Ho) and expected heterozygosity (He) values of 0.20 and 0.21, respectively. The mean Ho was found to be 0.15, while the He was 0.18 for the agroecological zone-based sub-populations. Notably, the structure-based clustering did not align with the AEZ. Analysis of molecular variance based on structure clustering indicated that 49% of the variation observed within samples, with an average genetic distance of 0.398, and 38% of the variance was among populations, reflecting high genetic differentiation (PhiPT = 0.38). The three gene pools and individual accession variation enhance our understanding of the genetic diversity in the 43 spider plant accessions from Zimbabwe and provide critical information for breeding initiatives, further genetic applications, and the conservation of these valuable genetic resources.
Files
Steps to reproduce
Data filtration was conducted utilizing the dartR package within the R software environment (Gruber et al., 2018). This process involved the elimination of monomorphic markers while retaining those that exhibited a call rate greater than 50% and a minor allele frequency exceeding 5%. Subsequently, data imputation was performed based on the Hardy-Weinberg equilibrium (Graffelman et al., 2015). To assess the characteristics of the filtered markers, metrics including polymorphic information content (PIC) and minor allele frequency and proportion of mutation types, including transversion (Tv) and transition (Ts), responsible for the observed polymorphism were calculated using the dartR package in R. Two classical approaches were used for the estimation of population genetic structure: ADMIXTURE and principal component analysis (PCA), implemented in the R package “LEA” . The ADMIXTURE analysis was conducted to estimate individual ancestry proportions. The analysis involved running the ADMIXTURE model for a range of K values (1-10) to identify the optimal number of clusters. Hundred replications were performed and the cross-validation procedure implemented was used to determine the best-fitting K, optimum number of ancestral populations and the best run (Frichot & François, 2015). The Q-matrices from the best run of the ADMIXTURE analysis were visualized to illustrate the proportion of ancestry from each inferred population for each individual using the STRUCTURE PLOT web application (Ramasamy et al., 2014). The first three principal components were extracted and visualized to summarize the major patterns of genetic variation using the CurlyWhirly software version 1.19.09.04. The PCA plots were overlaid with the admixture proportions and the AEZ to assess how well the admixture-based and AEZ-based clusters corresponded to the genetic variation observed in the PCA. The genetic diversity indices, including observed heterozygosity (Ho), expected heterozygosity (He), and the inbreeding coefficient (Fis), were computed for each population derived from the ADMIXTURE analysis and the AEZ classification using the adegenet package in R (Jombart & Ahmed, 2011). To further investigate the partitioning of the total genetic variance into components due to differences among populations, within populations and those within samples, we performed the Analysis of Molecular Variance (AMOVA) and generated the average PhiPT values (analogue of FST, fixation index) with the poppr package in R (Kamvar & Grunwald, 2021). Detailed population pairwise FST values were determined using the hierfstat package in R (Goudet, 2005). Euclidean-based genetic distances were calculated and Neighbor-Joining phylogenetic tree was constructed using the ape package in R (Paradis et al., 2004). The phylogenetic tree was exported in the “Newick” format and subsequently annotated using the iTOL v4 online tool (Letunic & Bork, 2021)