LLM NetOps/AIOps Six-Pillar Evidence Audit Dataset

Published: 1 September 2026| Version 3 | DOI: 10.17632/2p5ppxzy4s.3
Contributors:
Muhammad Bilal,
,
,
, Schahram Dustdar

Description

This dataset accompanies the survey “Large Language Models for Agentic NetOps and AIOps: Architectures, Evaluation, and Safety” (arXiv:2605.12729, https://arxiv.org/abs/2605.12729). It provides the record-level evidence audit used to support the survey’s analysis of LLM-enabled NetOps and AIOps across operational capability and autonomy, architecture and tool grounding, assurance and operational control, evaluation and benchmarking, security and adversarial robustness, and human factors, governance, and standards. The dataset records evidential status, analytical pillar, operational domain, task, autonomy role where applicable, tool surface, evaluation setting, safety controls, inclusion rationale, manuscript location, and permissible claim boundary. Separate files provide the direct LLM-facing subset, operational-system coding, search-update records, coverage-validation records, transparent exclusions, contextual comparison surveys, and PRISMA-style search and corpus accounting. The release includes machine-readable CSV and JSON files, a formatted workbook, a data dictionary, a matching machine-readable bibliography, citation metadata, a CC BY 4.0 licence, and integrity checksums. It contains no source full text, reviewer correspondence, private contact details, or internal compilation material.

Files

Steps to reproduce

1. Open `Evidence_Base_190_Public.csv`. Treat each row as one claim-linked evidence record. 2. Group records by `Primary analytical pillar`. Count each record once only. Do not include `Secondary analytical pillars`. Divide each count by 190. Expected counts are 53, 21, 37, 32, 27, and 20 in the order used by `Pillar_Distributions.csv`. 3. Group `Evidential status`. Expected totals are 95 Domain-operational technical, 33 Direct LLM-facing, 53 Cross-domain transferred, and 9 Normative. Compare these with `Corpus_Summary.json`. 4. Filter `Direct LLM-facing subset = Yes`, or use `Direct_LLM_33.csv`. The result must contain 33 records. 5. Filter `Direct operational-system subset = Yes`, or use `Operational_Systems_20.csv`. Only these 20 records may be assigned autonomy rungs. Do not assign autonomy rungs to frameworks, benchmarks, safety evaluations, transferred evidence, or normative sources. 6. Group operational-system records by `Autonomy rung`, `Evaluation setting`, `Write capability`, or other reported fields as required. Expected autonomy totals are R1=1, R2=8, R3=10, R4=1. Interpret all fields within each record's `Permissible claim boundary`. 7. Compare primary-pillar counts and shares with `Pillar_Distributions.csv`, and corpus-level totals with `Corpus_Summary.json`. 8. Use `Update_Search_Log.csv` to inspect the frozen targeted search routes through 21 August 2026. Use `Search_Update_18.csv` for the final retained 18-record update view. Use `Record_Substitutions_2.csv` to reconcile the two 28 August substitutions. `Coverage_Validation_8.csv` remains a separate later coverage-validation layer. 9. Use `PRISMA_Accounting.csv` to check aggregate identification, deduplication, screening, eligibility, and retained-record arithmetic. 10. Check record-level classifications against the source citation plus `Specific inclusion rationale`, `First substantive manuscript location`, and `Permissible claim boundary`.

Institutions

Categories

Artificial Intelligence, Computer Network, Computer Communications, Network Operation

Licence