Multi-Language Statistical Classification of Natural and Computer-Generated Texts
Not stated
- Funding
- Self-Funded PhD Students Only
- Application deadline
- Year-round applications
About the project
About the Project This project concerns the statistics of written documents, whose properties are analogous to those of other type-token systems such as city populations, personal incomes and galactic superclusters. Large-scale studies on the Project Gutenberg corpus reveal wide statistical variations between written documents [1][2]. Though many texts can be modelled with existing statistical theory, others are anomalous and deserve deeper examination. Significant differences emerge between different languages, so that languages of diverse groups (e.g. Finnish and English) could potentially be classified by their statistics alone. It would be instructive to observe the differences between AI and natural text for different languages and investigate how statistical detection of the former could operate between languages. We envisage the following program of study: 1. The compilation of a large body of literature from multiple corpora in which all significant language groups are evenly represented. 2. The identification and separation of statistical outliers within this group. 3. The generation of a corresponding corpus of AI-generated texts, of a similar size and covering the same language groups as the natural corpus. 4. The characterization of items within each corpora in terms of existing type-token models. 5. The subsequent ranking of these models by their ability to distinguish between languages and between AI and natural texts. Assuming the results are positive, this could lead to the creation of a language classifier. If the results are less successful, they may nevertheless guide the development of better models, applicable to a wider range of documents and languages.