Low read counts from Nanopore 18S rRNA amplicon-based sequencing

Published: 18 June 2025| Version 1 | DOI: 10.17632/95ymwxcj9v.1
Contributors:
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,

Description

This dataset includes taxonomic assignments derived from low read count samples (<20 reads) from Nanopore 18S rRNA amplicon-based sequencing, identified using Kraken2 and subsequently validated via BLAST against the NCBI nucleotide (nr) database. Despite the limited read depth, these assignments were retained for exploratory purposes, particularly to detect rare or potentially novel taxa and their possible areas of ocurrence, particularly in understudied areas. BLAST validation was performed using MEGABLAST under stringent criteria (≥95% identity, ≥98% query coverage, E = 0, bit score ≥50, and a ≥32-point difference between the top hit and the best alternative). Of the 49 assignments, 75.5% (37) were supported by BLAST validation at the genus level, suggesting a reasonable level of correspondence between methods. These data offer insight into the hidden diversity of hemoparasites and support future surveys in understudied areas.

Files

Steps to reproduce

The sequence data are provided in FASTQ format, which includes both nucleotide sequences and base-level quality scores (Phred scores). These files enable the reconstruction of consensus sequences when desired, allowing for consideration of base-call confidence. However, due to the low read depth (<20 reads per taxon), these data should not be used for population-level inference, variant calling, or quantitative analysis. Their intended use is for exploratory, pairwise comparisons and qualitative taxonomic validation using tools such as BLAST. Any biological or ecological interpretations drawn from these reads should be approached with caution and framed within the exploratory scope of the study.

Categories

DNA Sequencing, Next Generation Sequencing

Funders

Licence