ILS-Bench: Investor Language-to-Suitability Benchmark
Description
This dataset provides a synthetic benchmark for evaluating suitability-aware robo-advisory systems that convert investor language into portfolio and escalation decisions. The dataset contains 400 AI-assisted synthetic investor narratives designed to represent realistic client statements involving investment objectives, time horizons, risk tolerance, risk capacity, liquidity needs, behavioral signals, contradictions, and potential suitability concerns. No real investor data, personally identifiable information, or confidential financial records are included. Each narrative is accompanied by structured suitability labels, including risk tolerance, risk capacity, liquidity need, suitability risk, recommended portfolio class, and human-escalation decision. Portfolio classes include capital preservation, conservative, balanced, growth, aggressive growth, and human review. The dataset is intended for technical evaluation of robo-adviser architectures, including static questionnaire baselines, rule-based NLP systems, LLM-only advisory models, and hybrid NLP–digital twin systems with suitability rules and explainable audit trails. The benchmark was reviewed by a purposive expert-validation panel composed of four independent financial-domain experts: a retired portfolio manager, a senior trading professional, a FinTech executive, and a FinTech academic. Experts reviewed the proposed labels and indicated whether each classification was professionally reasonable. When they disagreed, they revised the relevant labels; comments were optional. The repository includes the synthetic cases, expert-validation records, consensus labels, codebook, and summary statistics. The dataset is designed for research on robo-advisers, financial digital twins, natural language processing, suitability assessment, explainable AI, and responsible AI governance in wealth management. It should be interpreted as a synthetic benchmark for methodological and system-evaluation purposes, not as evidence of real investor behavior.
Files
Steps to reproduce
The dataset can be reproduced through the following procedure: First, define a suitability codebook for robo-advisory classification, including risk tolerance, risk capacity, liquidity need, suitability risk, portfolio class, and human-escalation decision. Portfolio classes should include capital preservation, conservative, balanced, growth, aggressive growth, and human review. Second, generate synthetic investor narratives using a generative AI tool. Each narrative should be written as a short client statement and should include realistic combinations of investment objectives, time horizon, income stability, debt or emergency-reserve constraints, liquidity needs, risk preferences, behavioral reactions to losses, and possible contradictions between desired return and actual financial capacity. No real investor data or personally identifiable information should be used. Third, assign preliminary labels to each narrative according to the predefined codebook. Each case should be labeled for risk tolerance, risk capacity, liquidity need, suitability risk, recommended portfolio class, and escalation decision. The preliminary labels represent the initial benchmark classification. Fourth, submit the labeled cases to independent financial-domain experts for validation. Experts should review the proposed labels and indicate whether each classification is professionally reasonable using Agree, Disagree, or Unsure. When experts disagree, they should revise the relevant final label. Qualitative comments may be collected but should remain optional to reduce reviewer burden. Fifth, consolidate the expert responses into a long-format validation file, with one row per case per expert. Calculate agreement counts, disagreement counts, and uncertainty counts. Produce majority-consensus labels for each case across the expert panel. Cases with no clear majority or strong disagreement should be flagged for author adjudication or retained as ambiguous benchmark cases. Sixth, create the final repository file containing the synthetic cases, expert-validation records, consensus labels, codebook, expert-profile summary, and validation summary statistics. The resulting dataset can then be used to evaluate alternative robo-advisory architectures, including static questionnaire baselines, rule-based NLP models, LLM-only advisory models, and hybrid NLP–digital twin systems with suitability rules and human-escalation controls.
Institutions
- Ca' Foscari University of VeniceVeneto, Venice