JournalFit-NLP derived evaluation data for evidence-resolved scholarly journal recommendation

Published: 30 July 2026| Version 1 | DOI: 10.17632/cyb4fwbvyw.1
Contributor:

Description

This Version 2 dataset contains legally shareable derived evaluation data and reproducibility files from the JournalFit-NLP scholarly journal recommendation project. It includes 152,735 evidence-resolved active submission venue identifiers, frozen title-hash splits, query-level ranks and ranking metrics for six retrieval systems, cross-encoder and local-LLM diagnostic outputs, aggregate and stratified tables, bootstrap confidence intervals, paired randomization tests, protocol and environment records, data dictionaries, checksums, figures, and public validation scripts. The release does not redistribute licensed article abstracts, licensed venue-profile text, vector caches, or QLoRA adapter weights; the excluded artifacts and available hashes are documented in restricted_manifest.json.

Files

Steps to reproduce

1. Extract the dataset package. 2. Create the minimal environment from 05_reproducibility/environment_public.yml or requirements_public.txt. 3. Run python 05_reproducibility/validate_dataset.py. 4. Run python 05_reproducibility/recompute_aggregate_metrics.py. 5. Run python 05_reproducibility/recompute_significance.py. 6. Run python 05_reproducibility/make_data_figure.py. The public scripts validate and recompute the derived results; end-to-end model training requires restricted licensed text and model artifacts listed in restricted_manifest.json.

Institutions

Categories

Computer Science, Information Science, Recommendation System

Licence