SURE-Pipe: A pipeline to compare genomes and extract Shared and Unique Regions

Published: 12 August 2026| Version 1 | DOI: 10.17632/vr9ps58vyw.1
Contributors:
,
,

Description

This Mendeley Data repository contains the archived SURE-Pipe software package together with the benchmarking datasets used for performance evaluation and validation. It serves as the citable archival release accompanying the SURE-Pipe publication and includes the original source code, documentation, example workflows, and datasets required to reproduce the analyses. The original development repository is available on GitHub (https://github.com/BPaul-bioinfoLAB/SURE-Pipe), while this Mendeley Data release provides a stable, versioned archive for long-term reproducibility. The SURE-Pipe software identifies shared and unique genomic regions from pairwise and groupwise comparative genomics datasets. The archived release contains the scripts, workflow components, and supporting files corresponding to the version used in the study. Benchmarking was performed using simulated and real-world comparative genomics datasets. Simulated genomes were generated using a custom in-house simulator based on the Stan framework, modelling coalescent divergence, point mutations, and structural rearrangements between target and neighbouring clades. Each simulation consisted of five target and five neighbouring genomes with >97% ANI. Shared and unique genomic regions (200–5000 bp) were embedded into the target genomes to establish benchmarking ground truth, while inversions and duplications were introduced to mimic realistic genomic rearrangements. Simulations covered genome sizes from 100 kb to 75 Mb. Accuracy was assessed using 150 independent simulations with 4 Mb genomes, comparing SURE-Pipe v1.1 with KEC v1.1 and FUR v4.3 using sensitivity, specificity, precision, accuracy, and Matthews correlation coefficient (MCC). Scalability was evaluated through genome-size and genome-number benchmarking. Genome-size experiments used datasets ranging from 0.1 Mb to 640 Mb, while genome-number experiments fixed genome size at 5 Mb and increased dataset sizes from 5 target:20 neighbouring genomes to 160 target:640 neighbouring genomes. Runtime and peak memory usage were measured using /usr/bin/time on a Pop!_OS 22.04 LTS workstation equipped with an Intel Xeon E-2124G CPU (4 cores, 3.40 GHz) and 16 GB RAM. Biological validation included pairwise comparisons of six genome pairs representing viral, bacterial, and fungal genomes, and groupwise comparisons across 24 closely related Bacillus species using four genomes per species and neighbouring reference genomes from 42 Bacillus species. Together, the archived software and datasets enable full reproduction of the computational benchmarking and biological validation presented in the associated publication. Due to file size limitation we added remining data in following repository : Thomas, Infant; Paul, Bobby (2026), “SURE-Pipe: Supplementary Benchmarking Datasets-Scalability (Part 2)”, Mendeley Data, V1, doi: 10.17632/225k9hvbwj.1

Files

Institutions

Categories

Comparative Genomics, Genome

Funders

Licence