Specialized corpora — principled electronic collections of texts assembled to represent a specific genre, discipline, register, or writer population — have become a central methodological resource in academic writing research. This article reviews the applications of such corpora across the field, distinguishing them from large reference corpora and surveying the major purpose-built collections that have shaped the discipline, including the British Academic Written English (BAWE) corpus, the Michigan Corpus of Upper-level Student Papers (MICUSP), learner corpora, and small self-compiled (“do-it-yourself”) corpora. It then examines five interlocking domains of application: the description of formulaic language and lexical bundles; the analysis of stance, engagement, and metadiscourse; the identification of genre structure and rhetorical moves; the derivation of academic vocabulary and lexico-grammatical profiles; and the mapping of disciplinary and developmental variation. The article further considers pedagogical applications, in particular data-driven learning, and closes with a discussion of methodological limitations concerning representativeness, corpus design, and practitioner access. The evidence indicates that specialized corpora offer a uniquely fine-grained, context-sensitive window onto academic discourse that general reference corpora cannot provide.
Copyrights © 2026