Benchmark dataset for safety assessment under blast loading

Published: 25 December 2025| Version 1 | DOI: 10.17632/j53z3y7bx8.1
Contributor:
lu Wang

Description

This dataset provides a benchmark comprising 50 curated assessment tasks designed to evaluate the performance of large language model (LLM)-based multi-agent frameworks. Each item represents an independent task with both simple and complex task descriptions, along with corresponding ground-truth references.

Files

Categories

Safety, Blast Wave

Licence