BALS-D3QN Training and Evaluation Data

Published: 12 July 2026| Version 1 | DOI: 10.17632/t645xgpp6y.1
Contributors:
Konstantinos Voulgaridis, Thomas Lagkas

Description

The present datasets present the outcomes occurred during the research and development of the BALS-D3QN mechanism. BALS-D3QN is a balanced and adaptive scheduler based on the D3QN algorithm, designed to decide when the battery data of a DPP should be updated and resubmitted to an external DLT network, through the Age of Incorrect Information (AoII) metric. Also, BALS-D3QN acts as the decision mechanism of our previous ABS-TD3 solution, by reutilizing the latest observed submission latency of the Shimmer DLT to enable latency-exploratory updates. The proposed framework is evaluated using simulated battery trajectories from products listed within the EPREL database of the European Commission (EC), and compared against key baselines, namely consecutive updates, random policies, and periodic decisions, reflecting different types of behaviors. The folder "Battery" consists the battery trajectories that were utilized within the scope of or project. For our solution, a total of 100 battery trajectories were simulated through the PyBaMM tool and utilized for the required training process. Also, 7 additional battery trajectories were simulated as part of the evaluation process of our proposed solution. In both cases, the trajectories consist the data representing their lifecycle up to their 80% State of Health The folder "Evaluation Results" consists of the outcomes that occur based on the functionality of our BALS-D3QN solution, and the selected baselines: Spamming, Random, and Periodic policies. In all cases, and to ensure transparency, we provide all the available generated data, specifically, detailed per-battery trace, global evaluation data, and per-battery evaluation data. The folder "Latencies" consists of the latencies that were collected during the interaction process with the Shimmer DLT, and specifically by using the ABS-TD3 mechanism. Through the submission of a static DPP version, a total of 500 latencies were collected as part of the training process of BALS-D3QN, as well as a total of 1300 latencies as part of its evaluation process. Finally, the folder "Normalization Comparisons" consists of the results that occurred during the selection process of suitable normalization approach for the BALS-D3QN solution. Specifically, the folder presents the data that occurred from the Running Max, Running Mean, and Running Min-Max approaches, including per-seed results, last episode traces, loss history and training summaries. The results suggest that BALS-D3QN successfully minimizes the congestion impact occurring from the aggressive policies, minimizes the AoII compared to sparse periodic decisions, and showcases better adaptive behavior compared to periodic cases of similar ranges.

Files

Categories

Training, Deep Reinforcement Learning

Licence