spor-lora

Published: 8 July 2026| Version 1 | DOI: 10.17632/wmgrnb7dyh.1
Contributor:
Yihang Guo

Description

This dataset is constructed for structured cybersecurity knowledge generation and parameter-efficient fine-tuning of large language models. It covers four major domains: vulnerability analysis, code and software security, security operations and intelligence, and compliance, cryptography, and emerging security topics. Each sample contains task instructions, technical context, and structured reference responses, with annotations designed to support both domain knowledge learning and category-specific structure modeling. The dataset was developed from authoritative public cybersecurity resources and processed through manual collection, quality scoring, review, and revision procedures. It can support research on cybersecurity instruction tuning, structured text generation, domain adaptation, retrieval-augmented generation, and related large language model applications.

Files

Categories

Cybersecurity, Prompt-based Fine-Tuning LLM

Licence