JournalFit NLP Data Repository Package

Published: 17 July 2026| Version 1 | DOI: 10.17632/4d6d79wwmd.1
Contributor:

Description

The repository contains processed and analyzed data files generated during the JournalFit-NLP computational evaluation. CSV files provide tabular summaries of dataset composition, retrieval ablation, reranker blending, latency, quantization fidelity, resource monitoring, hardware/software environment, and expanded-pool evaluation metrics. JSON files contain machine-readable metric outputs for baseline, fine-tuned, expanded-pool, and reranking experiments. YAML files document model training and evaluation configurations. PNG and PDF figures visualize the architecture, development flow, retrieval ablations, reranker blending, reranker failure diagnosis, training/thermal monitoring, quantization trade-off, and CPU latency profiles. No private manuscripts, identifiable user records, social media data, or animal data are included.

Files

Steps to reproduce

Download and extract the repository package. Review README.txt, DATA_DESCRIPTION.txt, and file_manifest.csv to understand the dataset structure. Verify file integrity using checksums_sha256.txt. Open the YAML files for experiment settings, JSON files for aggregate metrics, and CSV files for tabular results. The reported tables and figures can be reproduced by loading the CSV/JSON files in Python 3.10+ using pandas and matplotlib. Raw third-party bibliographic metadata and user-uploaded manuscripts are not redistributed because of licensing, privacy, and consent constraints; the repository provides the aggregate metrics, configuration files, figure source data, and supplementary outputs required to reproduce the reported results.

Institutions

Categories

Computer Science, Artificial Intelligence

Licence