Bite-sized
Day 1: Thursday, 12 September
-Future Data Services 3 and Corpus-assisted discourse studies
Session convener: Bettina Moltrecht, Maria Leedham, University College London, The Open University
Session 3.A: Harmony: A natural language processing approach to data discovery and harmonisation.
In this workshop we will introduce you to our Harmony tool and platform. Harmony, the winner of the Wellcome Trust Data Prize, uses natural language processing technology to compare and match text content. Harmony was originally developed to facilitate fast and easy harmonisation of questionnaires commonly used in social science and mental health research. As part of the ESRC's Future Data Services programme, we have now extended Harmony's functionality to help researchers discover data from existing data catalogues, including our partner platforms the Catalogue of Mental Health Measures and Explorer by the UK Longitudinal Linkage Collaboration. In this session you will have the chance to try Harmony and we will show you how to use it to discover data from existing UK longitudinal studies and how to import the information back into Harmony to harmonise your measures and get your data research ready. We are inviting you to become part of our Harmony community and share your vision and ideas for Harmony's future developments with us. We look forward to welcoming you to our session.
Session 3.B:Corpus-assisted discourse studies: An introduction to a mixed methodology.
Corpus linguistics denotes the investigation of an electronic collection of texts (a ‘corpus’) using specialist software. It is the fastest-growing methodology within applied linguistics and is widely-used across the arts, humanities and social sciences. Rather than being used in isolation, corpus linguistics is frequently combined with discourse analysis and is known as CADS: corpus-assisted discourse studies. This combination harnesses the quantitative power of large-scale analysis using big datasets with the qualitative strengths of detailed textual reading and the human ability to spot patterns. This session aims to demonstrate the advantages of CADS and to inspire participants to consider how it could be used in their own field of research.
I will demonstrate CADS techniques such as extracting, sorting and categorising concordance lines, and tools such as plot dispersion and semantic tagging (cf. sentiment analysis). We will explore how an initial quantitative investigation such as keyword searches can provide a ‘way in’ to more detailed qualitative coding of data, illustrating searches through my recent research on Young Adult fiction.
The session will briefly consider future directions of CADS, particularly around the development of new tools such as visualisation software in the age of AI. Participants will leave the session with an understanding of what corpus linguistics can offer their area of research, and with online resources to find out more about the free software currently available. No prior knowledge or experience of corpus linguistics is needed.
