Benchmark Data for “From Formulation to Verification: A Reliability Benchmark and Error-Governance Framework for Large Language Models in Operations Research Decision Support”

Published: 11 June 2026| Version 1 | DOI: 10.17632/4st9wck76t.1
Contributor:
Burak Bora

Description

This dataset supports the manuscript entitled “From Formulation to Verification: A Reliability Benchmark and Error-Governance Framework for Large Language Models in Operations Research Decision Support”. The repository contains: • A benchmark of 120 operations research problems covering seven problem families: linear programming, integer/binary programming, transportation and assignment, scheduling, inventory models, multi-criteria decision-making, and sensitivity/diagnostic reasoning. • Gold-standard solutions, scoring rubrics, and benchmark documentation. • Model-generated responses collected from five large language model conditions under two prompt conditions, yielding 1,200 evaluated responses. • Final scored datasets, error-coding results, inter-rater reliability files, and statistical analysis outputs. • Prompt templates, raw model outputs, and analysis scripts used to reproduce the reported findings. The dataset was created to evaluate the reliability of large language models in operations research decision-support tasks. It supports the analyses reported in the manuscript, including model-performance comparisons, error typology development, reliability assessment, and the proposed decision-governance framework. All files are provided for research transparency, reproducibility, and future benchmark development in AI-assisted operations research and decision-support systems.

Files

Categories

Artificial Intelligence, Operations Research, Mathematical Modeling, Benchmarking

Licence