DArTseq_SNP_markers_common_bean_Peru
Description
This dataset contains the SNP markers of Phaseolus vulgaris landraces collected from the Peruvian Amazon and forms part of the study entitled "Population Structure and Diversity of Common Bean (Phaseolus vulgaris L.) Landraces in the Peruvian Amazon". Clustering analysis based on the neighbor‐joining algorithm resolved the evaluated accessions into two well‐defined major groups, reflecting their underlying genetic relationships. The first cluster comprised 284 accessions assigned to the Andean gene pool, whereas the second included 363 accessions corresponding to the Mesoamerican gene pool. This clear bifurcation underscores the strong genetic differentiation between the two domestication lineages. The first group comprised primarily accessions collected from the Peruvian Amazon, together with reference accessions included as an outgroup of Andean origin from the Peruvian highlands, the coastal region, and CIAT. The second group consisted mainly of Peruvian Amazon accessions, along with outgroup materials of Mesoamerican origin, including CIAT accessions and a reference sample from Mexico. This clear bipartite structure indicates that the first group corresponds to the Andean gene pool, whereas the second group represents the Mesoamerican gene pool, consistent with the recognized primary domestication centers of common bean. Principal Coordinate Analysis (PCoA) further supported this genetic differentiation, with the first three axes explaining 62.62%, 8.97%, and 2.77% of the total molecular variation, respectively. The first axis clearly separated the Andean and Mesoamerican gene pools, while the second and third axes captured additional within‐group variation. These results confirm the strong genetic structure within the collection and highlight the coexistence of both major gene pools among common bean landraces cultivated in the Peruvian Amazon. Population structure analysis provided additional evidence for this pattern. The Evanno method identified the strongest signal of genetic structure at K = 2, in agreement with the cluster analysis and PCoA results. These two genetic clusters correspond to the Andean and Mesoamerican gene pools and likely reflect contrasting domestication histories, evolutionary trajectories, and adaptation processes associated with their geographic origins. At a higher level of resolution (K = 4), two subgroups with low admixture were detected within each major group, revealing finer‐scale genetic differentiation. Notably, these subgroups showed a strong association with farmers’ common names for the cultivars, providing valuable insight into the historical diffusion, local adaptation, and cropping practices of common bean in the Amazonian region. Of the Amazonia accessions, 56.1 % clustered with the Mesoamerican pool and 43.9 % with the Andean gene pools.
Files
Steps to reproduce
Molecular markers discovery was performed using an analytical pipeline (DArTsoft14) developed by DArT company. This software identifies two types of markers: SilicoDArT (presence/absence) and SNP markers. Marker scoring and reproducibility were assessed using technical replicates, and only markers with high call rates and reproducibility were retained. Both markers were aligned to the Phaseolus vulgaris YP4 reference genome to identify the chromosome positions. Initially, we received 80,079 DArTSeq-derived SNP markers from SAGA-CIMMYT, which were polymorphic across common bean genotypes analyzed. Markers with unknown position were first removed from the analysis. Marker parameters such as call rate and minimum allele frequency (MAF) were calculated using dartR package in R version 4.5.1. Accordingly, SNP markers with ≥ 50% call rate and MAF of ≥ 0.5% were retained for further analysis. In this study, genotypes with missing data above ≥ 30% were removed from the analysis. From the initial set of 665 common bean accessions, 647 presented less than 30% missing data and were retained for downstream analyses. SNP markers were filtered to retain only those with a call rate of at least 50% and a minor allele frequency (MAF) of 0.5% or higher, yielding a filtered dataset of 24,283 high-quality SNPs. Alignment to the common bean YP4 reference genome revealed that 23,050 SNPs (94.9%) were successfully mapped to chromosome positions.
Institutions
- Universidad Nacional Agraria La MolinaLima Province, Lima