Benchmark dataset for safety assessment under blast loading
Published: 25 December 2025| Version 1 | DOI: 10.17632/j53z3y7bx8.1
Contributor:
lu WangDescription
This dataset provides a benchmark comprising 50 curated assessment tasks designed to evaluate the performance of large language model (LLM)-based multi-agent frameworks. Each item represents an independent task with both simple and complex task descriptions, along with corresponding ground-truth references.
Files
Categories
Safety, Blast Wave