iTARGET - Transposon sequencing read data

Published: 8 May 2025| Version 1 | DOI: 10.17632/wg4h59sr7c.1
Contributors:
JaeSeong Hwang,
,

Description

Transposon sequencing read data from research article titled 'Integrated Tn-seq and MAGE Assisted Rapid Genome Engineering Targeting in Escherichia coli' Data is provided as an excel file. The file composed of two data sheet as follows: Table 1. The number and fraction of Tn-seq reads generated from each population. Table 2. The number and fraction of Tn-seq reads generated from each population, highlighting peaks with no reads in the control population and a read fraction of more than 1% in the enriched population.

Files

Steps to reproduce

Transposon sequencing data generation and processing Two amplified Tn-seq libraries (control and enriched populations) were pooled and sequenced using a NextSeq 550 High Output Kit v2.5 (75 cycles; Illumina, San Diego, CA, USA). Of the reads generated by paired-end sequencing, only read1 (forward read) was of interest and used for analysis. Read1 contains the genomic sequences adjacent to the transposon insertion site, a 6 bp unique index sequence specific to each library at the 5' end, and partial ITR sequences. As the reads from both libraries were mixed, we first demultiplexed the raw data using a 6 bp unique index sequence (CAGATC and ACTTGA for the control and enriched populations, respectively) with Cutadapt version 2.8, using the following options: -a ctrl=CAGATC -a tet=ACTTGA. Cutadapt was used for subsequent read processing. Theoretically, the sequence and length of the ITR in the reads should be identical. However, in a significant proportion of reads, 1–2 bases were deleted from both ends of the ITR sequences. To improve the efficiency of read alignment and ensure a sufficient number of reads for subsequent analysis, we filtered reads containing more than 10 consecutive ITR bases from the demultiplexed raw data (Cutadapt option: -a ITR=TGGATGATAA -O 10). The remaining ITR sequences at the 3' end of the reads were trimmed, leaving only the genomic DNA sequences (Cutadapt option: -u -6). Given the genome size (~5 Mbps) and the potential for very short reads to align at multiple genomic locations, reads containing genomic DNA sequences longer than 12 bp were filtered out from the ITR-removed reads (Cutadapt option: -m 12). The final processed reads were aligned to the genome of E. coli BL21(DE3) (GenBank accession number: NC_012971.2) using Bowtie version 1.2.2 with the default settings. Genomic annotations and Tn-seq signals were visualized for direct comparison using MetaScope and Circos version 0.69.8. The raw FASTQ and processed (GFF file) data generated in this study have been deposited in the Gene Expression Omnibus (GEO) under Series GSE279638 (Reviewer’s token: cfotuggsjhcvhyx).

Institutions

Categories

Transposon, Target Identification, Next Generation Sequencing, Amplicon Sequencing

Funders

Licence