Frozen Vision Transformer and Convolutional Representations for Endoscopic Anatomical Image Classification
Description
This dataset contains frozen embeddings extracted from 14 pretrained CNN and Vision Transformer (ViT) backbones on the GastroHUN endoscopic image classification dataset, together with the full set of experimental results reported in the associated paper: 10-fold nested cross-validation scores, Friedman/Nemenyi statistical test outputs, held-out test predictions, per-class clinical metrics (sensitivity, specificity, AUC), and confusion matrices. The embeddings let others reproduce every downstream statistic and figure in the paper (classifier training, CV, statistical testing, tables and figures) without re-running feature extraction or needing GPU access only the SVM stage (fast, CPU-only, scikit-learn) is required.
Files
Steps to reproduce
https://github.com/EloneSampaio/gastrohun-frozen-cnn-vit-benchmark
Institutions
- Universidade Federal do PampaRio Grande do Sul, Bagé