What Teachers Say When they Write or Talk about Discourse Analysis
WHAT TEACHERS SAY WHEN THEY WRITE ABOUT DISCOURSE ANALYSIS
Why a corpus?
As Hunston (2002: 3) has said, “a corpus by itself can do nothing at all, being nothing other than a store of used language”. However, if organized under the right conditions, a corpus has significant advantages over other kinds of data. To begin with, a corpus is made up of real language rather than idealized examples. In addition, a corpus provides evidence of what is typical, as well as of what is exceptional. In order to substantiate intuition and claims, this evidence can actually be quantified, but only if the corpus is machine-readable. Once a computer can read a text, it allows a number of easily manageable software tools to identify, sort, count and group words in a particular text and across texts, making it easy (even for the most junior of researchers) to visualize lexical frequency and usual or unusual phraseology and collocations, a task that would be practically impossible if done manually. Therefore, digitized texts may be approached from a variety of entry-points and the results are easily stored and retrieved.
The first step in the data collection was to ask the student-teachers to write an evaluation of the two DA modules. Their appraisal was to be done in English and e-mailed as an attachment to one of the course tutors. It was hoped that the essays were the result of reflection outside class hours, and thus had been edited before submission. Out of the 13 teachers regularly attending the course, 11 submitted their evaluation in English as required but 2 wrote in Portuguese. Unfortunately, these latter answers had to be set aside because the software can only yield results if it reads strings of the same language The eleven attachments were subsequently saved as text documents; misspellings and typos were first checked in order to minimize interference with the computer program. Grammatical mistakes were kept as in the original.
The next step was to identify the most frequent words or clusters of words in the data, so as to verify whether there were signals of recurrent patterns of attitudinal language. To this end, we used WordSmith Tools, a software for lexical analysis (Scott, 1986).