Published: 16 May 2023| Version 1 | DOI: 10.17632/fkjv82s5r4.1
Bryar Hassan,


It is somewhat helpful to use well-understood and commonly used standard datasets so that the findings can be quickly evaluated. Nevertheless, most of the corpora datasets are prepared for NLP tasks. Wikipedia is a good source of well-organised written text corpora with a wide range of expertise. These articles are available online freely and conveniently. Based on the number of English articles that exist in Wikipedia, we have randomly chosen several articles to be used as corpus datasets in this experiment. The sample size is determined from the confidence level, margin of error, population proportion, and population size. That means 385 articles are needed to achieve a 95 percent confidence level that the actual value is within ± 5% of the calculated value.



Corpus Linguistics