Shotgun metagenomic dataset from a traditional Bario paddy field soil in the Kelabit Highlands, Sarawak, Malaysia
Description
Bario rice is a traditional aromatic rice variety cultivated once annually in the Kelabit Highlands of Sarawak, Malaysia, at approximately 3,500 feet above sea level. Its cultivation represents a distinctive low-input paddy agroecosystem that differs from conventionally managed rice production, where external inputs such as compound fertilizers and herbicides are commonly used to support crop productivity and weed management. In contrast, Bario rice is cultivated only once a year using traditional farming practices, with limited mechanization and no routine application of synthetic fertilizers or herbicides. Following the rice-growing season, the fields remain uncropped for much of the year, during which rice straw are retained in the field and the fields are traditionally used for water buffalo grazing. The combination of annual rice cultivation, residue decomposition, and water buffalo activity creates a distinctive soil environment in which organic matter and nutrients are returned to the field and subsequently transformed by soil microorganisms. These characteristics make Bario paddy soil a valuable agroecosystem for investigating microbial communities and functional processes under minimal external chemical inputs. Metagenomic DNA was extracted and subjected to shotgun sequencing, generating 2 × 150 bp paired-end reads using the Illumina HiSeq 2000 platform. The paired-end sequencing reads have been deposited in the NCBI Sequence Read Archive (SRA) under accession number SRA174292. A total of approximately 94.1 million paired-end Illumina reads were assembled de novo using MEGAHIT, generating 1,083,643 contigs with a total assembly size of 774.3 Mb and a maximum contig length of 107.7 kb. After filtering contigs shorter than 1 kb, 161,096 contigs representing 323.5 Mb were retained for downstream analyses. The assembly had a GC content of 53.83%, an N50 of 1,161 bp, and a largest contig length of 107,689 bp. For the present reanalysis, assembled metagenomic contigs were filtered to retain sequences longer than 1 kb prior to taxonomic and functional characterization. Genome annotation of assembled contigs was conducted using Prokka. Microbial taxonomic composition was predicted using Kraken2. Based on the taxonomic profile, the Bario paddy soil microbiome appears to be dominated by diverse soil-associated bacteria. A large proportion of reads remain unclassified, suggesting substantial unexplored microbial diversity in this traditional highland paddy ecosystem. From an agricultural perspective, the most interesting finding is the strong presence of nitrogen-cycling microorganisms. This Mendeley Data contained the assembly contigs and the annotations of the open reading frames. This dataset can be used as a resource for investigating the microbial diversity associated with Bario paddy soil and identifying genes involved in nutrient cycling.
Files
Steps to reproduce
1.Download the raw sequencing reads. The paired-end shotgun metagenomic sequencing reads are publicly available in the NCBI Sequence Read Archive (SRA) under accession number SRA174292. Users can retrieve the raw forward and reverse reads from the SRA for downstream analysis. 2. Download the assembled contigs. The assembled contigs generated from the metagenomic sequencing data are provided in this Mendeley Data repository. Both the complete assembly and a filtered dataset containing contigs >1,000 bp are available for download. Users may use these files directly for downstream analyses or reproduce the assembly from the raw SRA reads. 3. Reproduce the taxonomic profiling. An overview of the taxonomic composition of the metagenomic dataset is provided as the shared Kraken output file. Users can use this file to examine the taxonomic assignments and overall microbial composition of the dataset. 4. Access the predicted genes and functional annotations. The predicted open reading frames (ORFs) and their corresponding functional annotations are provided in a compressed ZIP folder. Download and extract the ZIP folder to access the individual files. These files can be used for downstream analyses of gene content and functional potential.