Analysis of English text genre classification based on dependency types
-
Yaqin Wang
Abstract
The present study aims to explore whether dependency type can be used as a distinctive text vector for classifying English genres. Three classification methods, namely principal component analysis, hierarchical clustering, and random forest were employed to investigate the clustering effect. Results show that dependency type is an effective measure in distinguishing text genres, especially between spoken genre and written genre.
Abstract
The present study aims to explore whether dependency type can be used as a distinctive text vector for classifying English genres. Three classification methods, namely principal component analysis, hierarchical clustering, and random forest were employed to investigate the clustering effect. Results show that dependency type is an effective measure in distinguishing text genres, especially between spoken genre and written genre.
Chapters in this book
- Prelim pages i
- Table of contents v
- Introduction 1
- Part I. Theory and models 7
- On the impact of the initial phrase length on the position of enclitics in Old Czech 9
- Term distance, frequency and collocations 21
- A method for the comparison of general sequences via type-token ratio 37
- Quantitative analysis of syllable properties in Croatian, Serbian, Russian, and Ukrainian 55
- N -grams of grammatical functions and their significant order in the Japanese clause 69
- Linking the dependents 93
- Grammar efficiency and the One-Meaning–One-Form Principle 109
- Distribution and characteristics of commonly used words across different texts in Japanese 121
- Part II. Empirical studies 135
- The perils of big data 137
- From distinguishability to informativity 145
- A Modern Greek readability tool 163
- Phonological properties as predictors of text success 177
- Calculating the victory chances 195
- Topological mapping for visualisation of high-dimensional historical linguistic data 209
- Book genre and author’s gender recognition based on titles 225
- Quantitative analysis of bibliographic corpora 239
- Analysis of English text genre classification based on dependency types 257
- In memory of Gabriel Altmann 271
- Index 277
Chapters in this book
- Prelim pages i
- Table of contents v
- Introduction 1
- Part I. Theory and models 7
- On the impact of the initial phrase length on the position of enclitics in Old Czech 9
- Term distance, frequency and collocations 21
- A method for the comparison of general sequences via type-token ratio 37
- Quantitative analysis of syllable properties in Croatian, Serbian, Russian, and Ukrainian 55
- N -grams of grammatical functions and their significant order in the Japanese clause 69
- Linking the dependents 93
- Grammar efficiency and the One-Meaning–One-Form Principle 109
- Distribution and characteristics of commonly used words across different texts in Japanese 121
- Part II. Empirical studies 135
- The perils of big data 137
- From distinguishability to informativity 145
- A Modern Greek readability tool 163
- Phonological properties as predictors of text success 177
- Calculating the victory chances 195
- Topological mapping for visualisation of high-dimensional historical linguistic data 209
- Book genre and author’s gender recognition based on titles 225
- Quantitative analysis of bibliographic corpora 239
- Analysis of English text genre classification based on dependency types 257
- In memory of Gabriel Altmann 271
- Index 277