AdviceNet: A Bengali Advice Dataset of Human Advice for Natural Language Processing
Description
This work is accepted at ICCIT-2025. This dataset was created to address the lack of labeled Bengali datasets for advice and intent classification. The texts reflect commonly used Bengali advice expressions from everyday communication and were manually annotated into Suggestive, Directive, and Opinion-based classes.
Files
Steps to reproduce
This dataset was created by collecting Bangla advice-related text from publicly available online sources such as social media platforms, discussion forums, and blogs. The data was manually reviewed to ensure relevance and quality. After collection, the text data was cleaned by removing duplicates, special characters, and irrelevant content. The dataset was then preprocessed using Python, including tokenization and normalization of Bangla text. Each entry was labeled based on the type of advice (e.g., general advice, emotional support, motivational guidance) through manual annotation.
Institutions
- Leading UniversitySylhet Division, Sylhet