Dataset of Evidence-linked tribology data for cobalt chromium molybdenum implant alloys extracted by language models
Description
The dataset presented here consolidates experimentally measured tribological and corrosion data for cobalt–chromium–molybdenum (CoCrMo) alloy and its candidate alternative bearing materials for artificial joint applications, compiled through a systematic literature survey. Manual annotation.zip contains excel files that are manual annotated. Gold-standard pairs.xlsx file is the final dataset after manul checking. processed Files.zip contains parsed and extracted data in json and xlsx format. curated data.csv is the curated dataset after removing duplicates.
Files
Steps to reproduce
All data extraction code is available at https://github.com/MarrytheToilet/SPED. The repository contains the multi-agent extraction pipeline (schema handling, extraction, merging, review and evidence verification) and the analysis scripts that regenerate every statistic and figure in this Descriptor, covering corpus ingestion, gold-standard alignment and scoring, evidence localization, the schema ablation and the confidence model. Document parsing used MinerU. Language-model inference used deepseek-v4-flash at temperature 0.1 with a 65,536-token output budget through an OpenAI-compatible interface. Sequence alignment used the RapidFuzz library. Package versions are pinned in the repository.
Institutions
- University of Science and Technology BeijingBeijing, Beijing
Categories
Funders
- National Key R&D Program of ChinaGrant ID: 2024YFB3817500