This systematic review synthesized empirical evidence on large language models (LLMs) and generative AI for English-language learning, assessment, and academic writing, with implications for English for specific purposes (ESP) in sport sciences. Following PRISMA 2020, a structured Scopus TITLE-ABS-KEY search identified 871 records; 50 peer-reviewed journal studies published between 2024 and 2026 met the eligibility criteria. Studies were synthesized using a three-stage thematic approach based on coding, descriptive themes, and analytical themes. Methodological reporting and applicability were appraised using a transparent four-domain FICO rubric. Screening was conducted by one reviewer; therefore, inter-rater reliability statistics were not calculable retrospectively. Four analytical clusters were identified: language-skills learning (8 studies), assessment and evaluation (8), academic writing and feedback (30), and adoption, perception, and learner agency (4). Evidence generally favoured structured, teacher-mediated GenAI use, while LLM assessment showed high reliability in some rubric-based tasks but variable validity. GenAI was most consistently useful for feedback, revision, and self-regulated learning. Because designs and outcomes were heterogeneous, no pooled effect size was calculated. Future research should prioritise discipline-specific ESP studies, longitudinal designs, transparent validation, and responsible AI integration.
Copyrights © 2026