Preprocessed Corpus Articles

Name: Preprocessed Corpus Articles
Creator: Bryar Hassan
Published: 2024-06-06T07:43:24.870Z
Keywords: Corpus Linguistics

Hassan, Bryar; Rashid, Tarik A.

doi:10.17632/fkjv82s5r4.2

Preprocessed Corpus Articles

Published: 6 June 2024| Version 2 | DOI: 10.17632/fkjv82s5r4.2

Contributors:

Bryar Hassan, Tarik A. Rashid

Description

It is somewhat helpful to use well-understood and commonly used standard datasets so that the findings can be quickly evaluated. Nevertheless, most of the corpora datasets are prepared for NLP tasks. Wikipedia is a good source of well-organised written text corpora with a wide range of expertise. These articles are available online freely and conveniently. Based on the number of English articles that exist in Wikipedia, we have randomly chosen several articles to be used as corpus datasets in this experiment. The sample size is determined from the confidence level, margin of error, population proportion, and population size. That means 385 articles are needed to achieve a 95 percent confidence level that the actual value is within ± 5% of the calculated value.

Preprocessed Corpus Articles

Description

Files

Categories

Licence