Bangla Wikipedia dataset

Name: Bangla Wikipedia dataset
Creator: Aisha Khatun
Published: 2020-10-04T06:03:57.305Z
Keywords: Natural Language Processing, Bengali Language

Khatun, Aisha; Rahman, Anisur; Islam, Md. Saiful

doi:10.17632/3ph3n78fp7.4

Bangla Wikipedia dataset

Published: 4 October 2020| Version 4 | DOI: 10.17632/3ph3n78fp7.4

Contributors:

Aisha Khatun, Anisur Rahman, Md. Saiful Islam

Description

A subset of the Bangla version of the Wikipedia text. To create the Wikipedia dataset, we collected the Bangla wiki-dump of 10th June, 2019. The files are then merged and each article is selected as a sample text. All HTML tags were removed and the title of the page was stripped from the beginning of the text. This dataset contains 70377 samples with a total number of words being 18229481. The entire dataset has 1289249 unique words, which is 7% of the total vocabulary.

Files

Institutions

Shahjalal University of Science and Technology

Bangla Wikipedia dataset

Description

Files

Institutions

Categories

Licence