HugSelect Dataset and Replication Package
Description
The HugSelect Dataset and Replication Package accompanies the study “HugSelect: An Explainable Multi-Criteria Decision-Support Framework for Foundation Model Selection.” HugSelect treats foundation model selection as an explicit and auditable software-component selection problem by combining repository metadata, functional capabilities, community-perceived quality attributes, and multi-criteria decision-making (MCDM). The package contains the data and software artifacts used to construct and evaluate a curated knowledge base of 71,274 Hugging Face models. The processing pipelines integrate model metadata, model-card and README content, and community discussions into structured information describing model families, modalities, tasks, functional capabilities, operational characteristics, and ISO/IEC 25010-inspired perceived quality attributes. The repository is organized into six components: Raw Information contains the collected model metadata, descriptions, and supporting community-derived information. Processing Pipelines contains intermediate and final artifacts generated during metadata preparation, functional-feature extraction and clustering, model-family identification, sentiment analysis, and quality-attribute mapping. Pipeline Validation provides manually inspected and reference datasets used to evaluate the automated extraction pipelines for functional features, sentiment analysis, and quality-attribute mapping. Case Study contains 44 literature-derived foundation-model selection scenarios, recommendation outputs from HugSelect and LLM-based baselines, model- and family-level evaluation results, and ablation-analysis artifacts. User Study contains empirical data from the exploratory practitioner evaluation of HugSelect, including Technology Acceptance Model (TAM)-based assessments of usefulness, usability, and transparency. HugSelectCode contains the source code implementing the HugSelect framework, including data processing, knowledge-base preparation, requirement processing, candidate matching, WSM/SAW-based multi-criteria ranking, and generation of explainable recommendations. The package supports reproduction of the empirical evaluation reported in the accompanying paper and can be reused for research on foundation model selection, AI engineering, recommender systems, repository mining, software analytics, information retrieval, MCDM, and explainable decision support. By publishing the data, validation artifacts, evaluation results, and source code together, this repository provides an inspectable, reproducible, extensible, and reusable research artifact for evidence-driven foundation model selection.
Files
Institutions
- Wageningen University & ResearchGelderland, Wageningen
- Utrecht UniversityUtrecht, Utrecht
- Shiraz UniversityFars, Shiraz