Large language model-based artificial intelligence for improving personal finance in higher technical institutes

Published: 18 August 2026| Version 1 | DOI: 10.17632/2ryx4gsr56.1
Contributors:
Gary Galvez,

Description

This dataset contains individual pretest and posttest measurements from 77 students of a technical higher education institute in the La Libertad region, Peru, who participated in a single-group quasi-experimental study on the use of a locally executed large language model for personal finance management, conducted during the 2025-II academic term. Measurements were obtained through the CFPE-IMFP questionnaire, a 39-item instrument organised in three sections: financial knowledge (K, 15 multiple-choice items), financial behaviour (C, 16 frequency-scale items) and financial outcome (R, 8 indicators of economic situation). Each section is transformed to a standardised 0-100 scale, and the Personal Finance Improvement Index is computed as IMFP = 0.30K + 0.40C + 0.30R, weighting behaviour most heavily given its mediating role between declarative knowledge and observable economic outcomes. The posttest was administered after 7 to 14 days of system use. The dataset includes raw item-level responses for both measurement points, the derived standardised scores, and the scoring keys required to recompute those scores from scratch. Participants are identified by sequential codes (EST-001 to EST-077) that match across all files, enabling the paired analyses reported in the associated article. No personal identifiers are included and all participants provided written informed consent. Note on scoring: the maximum score for section R is 20, not 24, because four of the eight indicators are scored 0-2 rather than 0-3. Using 24 as the denominator will not reproduce the published results. These data come from a design without a control condition, with a short intervention period and self-reported measures. Effect sizes should be read as preliminary evidence of feasibility and acceptability rather than as unbiased estimates of a causal effect.

Files

Steps to reproduce

Load resultados-imfp.csv, which contains the standardised K, C, R and IMFP scores for both measurement points. For each dimension, compute the mean and standard deviation at pretest and posttest, the paired differences, the Wilcoxon signed-rank test and the paired t-test. This reproduces every value in Table 3 of the associated article: IMFP rises from 45.46 (6.58) to 65.22 (6.96), with a mean difference of 19.76 (3.55), W = 0, t(76) = 48.78 and Cohen's d = 5.56. To recompute the standardised scores from raw responses, load resultados-pretest.csv and resultados-posttest.csv. Score each K item against clave-k.csv and divide the number of correct answers by 15. Sum C1 to C16 and divide by 48. Sum R1 to R8 and divide by 20, following the per-item scales in clave-r.csv. Multiply each result by 100 and apply the weighted formula. The README.md file includes a ten-line Python script that performs both steps.

Institutions

Categories

Higher Education, Technology, Financial Analysis, Applied Machine Learning

Licence