Data from Castro et al[Taxonomic invisibility and knowledge shortfalls in terrestrial molluscs of the Caatinga Dominion, a seasonally dry tropical region]
Description
This dataset supports the study “Taxonomic invisibility and knowledge shortfalls in terrestrial molluscs of the Caatinga Dominion, a seasonally dry tropical region”. We analyse whether the currently available occurrence data for Caatinga terrestrial molluscs are insufficient to represent the known and expected spatial distribution of the group, and whether many species become analytically “invisible” when minimum occurrence-record thresholds are applied. The dataset contains curated occurrence records and R scripts used to assess observed richness, expected richness, sampling completeness, spatial sampling gaps and species visibility under increasing occurrence thresholds. Occurrence records were compiled from scientific collections, GBIF, and the published literature, then taxonomically standardised and spatially filtered to retain valid records of terrestrial molluscs in the Caatinga Dominion, north-eastern Brazil. The data show that 155 terrestrial mollusc species, belonging to 62 genera and 25 families, are currently documented from 944 occurrence records inside the Caatinga Dominion. However, these records are highly sparse and spatially uneven. Only a small fraction of 10 km grid cells contain occurrence records, whereas most cells are included in the expected-occupancy area but lack any formal record. The outputs also show very low sampling completeness, low overlap between observed and expected richness, and a rapid decline in the number of species eligible for spatial analyses as the minimum occurrence threshold increases. Expected richness should be interpreted as a standardised, data-driven hypothesis of potential occupancy. It was derived from simple geometric occupancy envelopes, using buffers for poorly recorded species and concave hulls for species with enough records to construct polygons. These outputs are intended to support reproducibility, transparency and future studies on biodiversity knowledge shortfalls, spatial sampling gaps and conservation planning for terrestrial molluscs in tropical dry regions.
Files
Steps to reproduce
The workflow can be reproduced using the R script provided with this dataset. The main input file is the curated occurrence table. A spatial polygon delimiting the Caatinga Dominion is also available. First, occurrence records were compiled from Brazilian scientific collections, GBIF and published literature. GBIF records were filtered to retain only collection-based records, excluding human observations, machine observations, living specimens and records without evidence of deposited material. Records lacking species-level identification, valid coordinates or reliable taxonomic status were excluded. Coordinates were checked for common spatial errors, including zero coordinates, marine points and obvious outliers. Duplicate records were collapsed based on species, coordinates, and available collection metadata. Second, taxonomy was standardised using the updated checklist of Brazilian terrestrial gastropods and associated synonymy. Only valid species names were retained. Each record was then classified as occurring inside or outside the Caatinga boundary polygon. Observed richness and record-based summaries were calculated only from records inside the Caatinga Dominion, whereas records outside the Caatinga were used, when available, to help construct species-specific expected-occupancy envelopes and reduce boundary artefacts. Third, a 10 km grid was created and clipped to the Caatinga Dominion. Observed richness was calculated by counting the number of species recorded in each grid cell. Expected richness was estimated by constructing geometric occupancy envelopes for each species. Species with fewer than three unique Caatinga records were represented by 5 km buffers around each Caatinga record. Species with three or more unique Caatinga records were represented by concave-hull polygons using a ratio parameter of 0.35. All envelopes were clipped to the Caatinga Dominion and rasterised to the 10 km grid. Fourth, expected richness was calculated by summing binary expected-presence layers across species. Sampling completeness was calculated as the observed richness divided by the expected richness per cell. Additional spatial metrics were calculated, including cells with records, cells with expected occurrence, cells expected to be occupied but lacking records, cells with no expected occurrence and no records, expected-richness percentile classes and observed–expected overlap using the Jaccard index. Finally, species visibility was assessed by applying minimum occurrence-record thresholds of t = 3, 10, 20, 30 and 38 records. A species was considered eligible at a given threshold when it had at least that number of unique Caatinga records. The script exports raster files, CSV/XLSX summary tables, supplementary tables and figures used in the manuscript.