TieOutBench v1 - finance LLM evaluation suite, gold cases, and graded model runs
Published: 14 September 2026| Version 2 | DOI: 10.17632/m5wfrnhdsp.2
Contributor:
Dmitry KrutousDescription
This dataset is the exact repository snapshot behind the technical report "TieOutBench: Measuring whether an LLM's finance work ties out - before deciding whether to trust it" (SSRN abstract 7243025, August 2026). It is the state of the public repository at git commit b965426 (2026-08-12), the version the report's numbers were produced from.
Files
Categories
Computer Science, Accounting, Finance, Artificial Intelligence, Natural Language Processing