ArticleName: Genre markup of the Tomsk dialect corpus: from concept to implementation Authors: Zemicheva S. S. Tomsk State University, Tomsk, Russian Federation In the section Linguistics
Abstract: The relevance of the study is due to the fact that it has been conducted at the intersection of two scientific fields: corpus linguistics and communicative dialectology. The paper presents a comparative analysis of corpus practice based on the material of spoken language. Also, consideration is given to the process and results of creating a discursively annotated corpus of dialect speech with a size of more than 2 million tokens. Discursive markup implies the labeling of three parameters: topic, type, and genre of the text. The novelty of this research project is related to the fact that, for the first time, a large array of dialectal texts has been marked up according to intentional orientation: not only folklore but also speech genres have been annotated. The value of the new source is provided by the combination of archived data with current materials. A methodological advantage of the corpus is the possibility of combining qualitative and quantitative analysis. 