Computer Science

Multi-Language Statistical Classification of Natural and Computer-Generated Texts

Kingston University

Not stated

Location
London, United Kingdom, United Kingdom
Funding
Self-Funded PhD Students Only
Application deadline
Year-round applications

About the project

About the Project This project concerns the statistics of written documents, whose properties are analogous to those of other type-token systems such as city populations, personal incomes and galactic superclusters. Large-scale studies on the Project Gutenberg corpus reveal wide statistical variations between written documents [1][2]. Though many texts can be modelled with existing statistical theory, others are anomalous and deserve deeper examination. Significant differences emerge between different languages, so that languages of diverse groups (e.g. Finnish and English) could potentially be classified by their statistics alone. It would be instructive to observe the differences between AI and natural text for different languages and investigate how statistical detection of the former could operate between languages. We envisage the following program of study: 1. The compilation of a large body of literature from multiple corpora in which all significant language groups are evenly represented. 2. The identification and separation of statistical outliers within this group. 3. The generation of a corresponding corpus of AI-generated texts, of a similar size and covering the same language groups as the natural corpus. 4. The characterization of items within each corpora in terms of existing type-token models. 5. The subsequent ranking of these models by their ability to distinguish between languages and between AI and natural texts. Assuming the results are positive, this could lead to the creation of a language classifier. If the results are less successful, they may nevertheless guide the development of better models, applicable to a wider range of documents and languages.

Research areas

ComputerScienceProbabilityStatisticsMulti-LanguageStatisticalClassificationofNaturalandComputer-GeneratedTexts