Frozen Vision Transformer and Convolutional Representations for Endoscopic Anatomical Image Classification

Published: 4 August 2026| Version 2 | DOI: 10.17632/47k67b9gh7.2
Contributors:
,
,
,

Description

This dataset contains frozen embeddings extracted from 14 pretrained CNN and Vision Transformer (ViT) backbones on the GastroHUN endoscopic image classification dataset, together with the full set of experimental results reported in the associated paper: 10-fold nested cross-validation scores, Friedman/Nemenyi statistical test outputs, held-out test predictions, per-class clinical metrics (sensitivity, specificity, AUC), and confusion matrices. The embeddings let others reproduce every downstream statistic and figure in the paper (classifier training, CV, statistical testing, tables and figures) without re-running feature extraction or needing GPU access only the SVM stage (fast, CPU-only, scikit-learn) is required.

Files

Steps to reproduce

https://github.com/EloneSampaio/gastrohun-frozen-cnn-vit-benchmark

Institutions

Categories

Artificial Intelligence, Computer Vision, Gastroenterology, Machine Learning, Endoscopy, Image Classification, Diagnostic Imaging, Computer-Aided Diagnosis, Medical Image Processing, Convolutional Neural Network, Deep Learning

Licence