PeaDetect Dataset: Curated Audio Recordings of Indian Peafowl (Pavo cristatus) Vocalizations for Bioacoustic Monitoring

Published: 8 July 2026| Version 1 | DOI: 10.17632/skryf4tkff.1
Contributors:
, Asanka Sayakkara,
, Titus Jayarathna, Roshan Ragel

Description

This dataset was curated mainly to use as an enabling technology for mitigation strategies for the human-peafowl conflict that exists in most regions in Sri Lanka and India. The absence of natural predators has contributed to a significant increase in the peafowl population, exacerbating challenges for farmers. Peafowls are sometimes considered agricultural pests due to their tendency to feed on and damage crops. The vocalizations are from the Indian Peafowl (Pavo cristatus), a species native to the Indian subcontinent and especially abundant in India and Sri Lanka. Peafowl belong to the genus Pavo in the family Phasianidae. "Peafowl" refers to both males (peacocks) and females (peahens). Users of this dataset are required to cite the following publications: R. Puvanendran, A. P. Sayakkara, S. S. Seneviratne, T. Jayarathna, and R. G. Ragel, "PeaDetect Dataset: Curated Audio Recordings of Indian Peafowl (Pavo cristatus) Vocalizations for Bioacoustic Monitoring," Data in Brief, Art. no. 113038, 2026. doi: 10.1016/j.dib.2026.113038

Files

Steps to reproduce

A- Full_Dataset The Full_Dataset/ directory contains a total of 2,950 audio clips, consisting of presence and absence classes. The presence class includes 1,475 audio clips containing vocalizations of the Indian peafowl (Pavo cristatus). The absence class includes 1,475 audio clips containing no peafowl vocalizations, comprising other bird species calls and environmental noises. Each audio file is stored in WAV format with a sampling rate of 44.1 kHz, 16-bit depth, and stereo channels. All clips have a fixed duration of 5 seconds. B-Metadata Files: The metadata/ directory contains two CSV files. 1) The metadata.csv file contains 2,950 rows and 6 columns, with each row corresponding to one audio clip. Table 1 describes the columns in this file. Table 2 shows few samples from metadata.csv. 2) The file sources.csv provides full attribution for each original presence recording, including catalog number / source ID, recordist name, recording date, length, time, country, location, coordinates, and quality rating (A through E). C-Documentation: The documentation/ directory contains three files: • README.md: A Markdown file providing a quick overview of the dataset, its structure, and instructions for use. • collection_protocol.pdf: A PDF document detailing the data collection methodology, including search criteria, filtering rules, and processing steps. • annotation_guidelines.pdf: A PDF document describing the protocol used for manual annotation, including decision rules for challenging cases. D- Ecological Metadata Language To ensure full compliance with FAIR (Findable, Accessible, Interoperable, Reusable) data principles and professional ecological data standards, dataset metadata was documented using Ecological Metadata Language (EML) E- Additional Information: This includes the information about the geographical distribution and the taxonomy of the Species (Indian peafowl)

Institutions

Categories

Conservation Agriculture, Bioacoustics, Computational Engineering

Licence