From local patches to stable representations: Multiscale analysis and robustness assessment of visual similarity in polished Portuguese granites

Published: 14 July 2026| Version 2 | DOI: 10.17632/8scf593pn6.2
Contributor:
Jose Antonio Valido García

Description

This repository contains the image dataset used in the study of self-supervised visual representation and similarity analysis of polished ornamental granites from Portugal. The dataset comprises 100 images representing ten commercial granite varieties, with ten specimens per variety. Each deposited image corresponds to a central crop of 3072 × 2048 px extracted from a specimen digitized under controlled acquisition conditions at 600 dpi. These cropped images constituted the source dataset from which the patch-based representations used for visual feature extraction and subsequent analysis were generated. In the associated study, each image was divided into a regular 12 × 8 grid of 96 non-overlapping patches of 256 × 256 px, resulting in 9600 patches overall. The repository does not include the derived patches, which can be reproduced directly from the deposited images following the procedure provided in the “Steps to reproduce” section.

Files

Steps to reproduce

1. Use the 100 cropped specimen images provided in the repository, corresponding to ten commercial polished granite varieties with ten specimens per variety. Each deposited image has dimensions of 3072 × 2048 px and represents the central region selected from the original 600 dpi acquisition. 2. Divide each image into a regular 12 × 8 grid of non-overlapping square patches of 256 × 256 px. No resizing, overlap, interpolation, image enhancement, filtering, or colour correction should be applied during patch generation. 3. This partitioning procedure produces 96 patches per specimen, 960 patches per granite variety, and 9600 patches in total. 4. Convert each patch to RGB and process it using the pretrained facebook/dinov2-base model. Disable resizing and centre cropping in the DINOv2 image processor so that each patch is introduced at its original 256 × 256 px resolution. Apply only the rescaling and normalization operations defined by the pretrained model. 5. Extract the 768-dimensional feature vector associated with the [CLS] token of the final transformer block for each patch. 6. Obtain one specimen-level representation by calculating the arithmetic mean of the 96 patch embeddings belonging to each specimen. No additional normalization or standardization should be applied to the DINOv2 embeddings in the main analysis. 7. For the multiscale analysis, randomly select 1, 2, 4, 8, 16, 32, 64, or 96 patches without replacement within each specimen and average their embeddings. Repeat the procedure 20 times for each scale below 96 patches, using a fixed random seed of 42. The 96-patch condition uses all available patches and therefore produces a single deterministic result. 8. Perform the hierarchical, robustness, methodological-sensitivity, and conventional-descriptor analyses using the procedures and parameter settings reported in the associated publication.

Institutions

Categories

Computer Science, Geology, Materials Science, Data Acquisition

Funders

  • FCT – Foundation for Science and Technology under the strategic projects
    Grant ID: UID/00073/2025, UID/PRR/00073/2025, and UID/PRR2/00073/2025
  • Government of the Canary Islands (Agencia Canaria de Investigación, Innovación y Sociedad de la Información)
    Grant ID: Catalina Ruiz

Licence