The ALeSKo learner corpus: Design – annotation – quantitative analyses

Heike Zinsmeister; Margit Breckle

Chapter

The ALeSKo learner corpus

Design – annotation – quantitative analyses

Heike Zinsmeister and Margit Breckle

Published by

View more publications by John Benjamins Publishing Company

To Publisher Page

This chapter is in the book Multilingual Corpora and Multilingual Corpus Analysis

Abstract

The ALesKo learner corpus is a small-scale comparable corpus consisting of two subcorpora: annotated essays by advanced Chinese learners of German and comparable essays by German native speakers. The motivation for its compilation was the investigation of discourse-related phenomena such as local coherence in second-language acquisition of German. After introducing how the texts were compiled and annotated, the article focuses on quantitative studies at the token level. We discuss problems of tokenisation and part-of-speech tagging and compare the inventory of the two subcorpora in terms of frequently used N-grams and lexical richness, among other aspects. We conclude the article by describing possible applications of the study in foreign language acquisition research and language teaching.

You are currently not able to access this content.

Abstract

You are currently not able to access this content.

Chapters in this book

Prelim pages i
Table of contents v
Introduction xi
Section 1. Learner and attrition corpora
The LeaP corpus 3
Technological and methodological challenges in creating, annotating and sharing a learner corpus of spoken German 25
Creation and analysis of a reading comprehension exercise corpus 47
The ALeSKo learner corpus 71
Corpora of spoken Spanish by simultaneous and successive German-Spanish bilingual and Spanish monolingual children 97
Monolingual and bilingual phonoprosodic corpora of child German and child Spanish 107
Pragmatic corpus analysis, exemplified by Turkish-German bilingual and monolingual data 123
Corpus of Polish spoken in Germany 153
The HABLA-corpus (German-French and German-Italian) 163
Section 2. Language contact corpora
The Hamburg Corpus of Argentinean Spanish (HaCASpa) 183
Ad hoc contact phenomena or established features of a contact variety? 199
Phonoprosodic corpus of spoken Catalan (PhonCAT) 215
Researching the intelligibility of a (German) dialect 231
Annotating ambiguity 245
Section 3. Interpreting corpora
Sharing community interpreting corpora 275
CoSi – A Corpus of Consecutive and Simultaneous Interpreting 295
The corpus “Interpreting in Hospitals” 305
Section 4. Comparable and parallel corpora
The GeWiss corpus 319
Korpus C4 339
Treebanks in translation studies 347
Section 5. Corpus tools
Multilingual phonological corpus analysis 365
Finding the balance between strict defaults and total openness 383
General index 401
Corpora index 405
Language index 407

https://doi.org/10.1075/hsm.14.06zin

Chapters in this book

Prelim pages i
Table of contents v
Introduction xi
Section 1. Learner and attrition corpora
The LeaP corpus 3
Technological and methodological challenges in creating, annotating and sharing a learner corpus of spoken German 25
Creation and analysis of a reading comprehension exercise corpus 47
The ALeSKo learner corpus 71
Corpora of spoken Spanish by simultaneous and successive German-Spanish bilingual and Spanish monolingual children 97
Monolingual and bilingual phonoprosodic corpora of child German and child Spanish 107
Pragmatic corpus analysis, exemplified by Turkish-German bilingual and monolingual data 123
Corpus of Polish spoken in Germany 153
The HABLA-corpus (German-French and German-Italian) 163
Section 2. Language contact corpora
The Hamburg Corpus of Argentinean Spanish (HaCASpa) 183
Ad hoc contact phenomena or established features of a contact variety? 199
Phonoprosodic corpus of spoken Catalan (PhonCAT) 215
Researching the intelligibility of a (German) dialect 231
Annotating ambiguity 245
Section 3. Interpreting corpora
Sharing community interpreting corpora 275
CoSi – A Corpus of Consecutive and Simultaneous Interpreting 295
The corpus “Interpreting in Hospitals” 305
Section 4. Comparable and parallel corpora
The GeWiss corpus 319
Korpus C4 339
Treebanks in translation studies 347
Section 5. Corpus tools
Multilingual phonological corpus analysis 365
Finding the balance between strict defaults and total openness 383
General index 401
Corpora index 405
Language index 407

The ALeSKo learner corpus

Abstract

Chapter PDF View

Abstract

Chapters in this book

Chapters in this book

Chapters in this book