Thumbnail
Access Restriction
Open

Author Guerrero, Fabio G.
Source arXiv.org
Content type Text
File Format PDF
Date of Submission 2009-11-11
Language English
Subject Domain (in DDC) Computer science, information & general works
Subject Keyword Computer Science - Computation and Language ♦ cs
Abstract A simple method for finding the entropy and redundancy of a reasonable long sample of English text by direct computer processing and from first principles according to Shannon theory is presented. As an example, results on the entropy of the English language have been obtained based on a total of 20.3 million characters of written English, considering symbols from one to five hundred characters in length. Besides a more realistic value of the entropy of English, a new perspective on some classic entropy-related concepts is presented. This method can also be extended to other Latin languages. Some implications for practical applications such as plagiarism-detection software, and the minimum number of words that should be used in social Internet network messaging, are discussed.
Educational Use Research
Learning Resource Type Article


Open content in new tab

   Open content in new tab