Assessing the genotype-by-year effect on training set composition: scripts and dataset
Description
This dataset supports the study “Assessing the genotype-by-year effect on training set composition and alternatives to mitigate its impact on genomic selection accuracy.” The research investigates how genotype-by-year (G×Y) interactions influence genomic prediction (GP) accuracy and whether including overlapping checks and progenies in the training set can mitigate these effects. We hypothesized that limited overlap of entries across years would fail to correct for G×Y effects, reducing selection response. To test this, we conducted stochastic simulations using AlphaSimR, based on the structure of the LSU AgCenter rice breeding program. A total of 36 scenarios were simulated, combining four levels of G×Y interaction (0%, 25%, 50%, 75%) with three levels of overlap (0%, 5%, 10%) for both progenies and checks. The dataset includes: • Input parameters and configuration files for all scenarios • R scripts used for simulation and analysis • Output data containing genetic parameters: additive variance, population mean, best line performance, and prediction accuracy Results showed that stronger G×Y interactions consistently reduced GP accuracy and selection gains. Including up to 10% of checks and progenies in the training set did not significantly improve predictive performance. These findings indicate that such overlap is insufficient to account for temporal variation, especially in breeding programs with limited connectivity between cycles. The dataset enables full reproducibility and can support further research on training set optimization, modeling G×Y interactions, and evaluating selection strategies.
Files
Steps to reproduce
To reproduce the results of the study “Assessing the genotype-by-year effect on training set composition and alternatives to mitigate its impact on genomic selection accuracy,” please follow the steps below: 1. Run the simulations Navigate to the Scripts directory and run the R script entitled GxY_final.R. This script performs stochastic simulations for the 36 scenarios described in the paper, which combine different levels of genotype-by-year (G×Y) interaction (0%, 25%, 50%, and 75%) and overlap proportions (0%, 5%, and 10%) for both progenies and checks. The script uses the R package AlphaSimR and generates simulated phenotypic and genotypic data for each scenario. 2. Generate the graphical outputs After completing the simulations, use the script graphs.R, also located in the Scripts directory. This script processes the output from the simulations and generates the main plots shown in the manuscript: • Population Mean (PM) • Additive Variance (Va) • Best Line Performance • Prediction Accuracy (Ac) Each plot is generated for all combinations of G×Y interaction and overlap proportions. 3. Access the results The simulated results used in the publication are available in the Results directory. This includes all output files already generated from the GxY_final.R script and organized by scenario. If you do not wish to rerun the simulations, you can directly access these files to replicate the analysis and figures. Additional notes: • All scripts are written in R and require the packages: AlphaSimR, ggplot2, dplyr, and readr. • The simulations may take some time to complete depending on your system performance. • The structure of output folders follows a naming convention that reflects the G×Y level and the proportion of checks and progenies. This pipeline ensures full reproducibility of the analyses and figures presented in the manuscript and can be adapted to other breeding scenarios or species.
Institutions
- Louisiana State University
- North Carolina State University
- Universidade Federal de Lavras