Yahya Aulia Abdillah
S2 Teknik Informatika, Universitas Amikom Yogyakarta, Indonesia

Published : 1 Documents Claim Missing Document
Claim Missing Document
Check
Articles

Found 1 Documents
Search

Harnessing Big Data for Corpus Linguistics: Redefining Language Patterns and Usage in the Digital Age Yahya Aulia Abdillah
Prosiding SENALA (Seminar Nasional Linguistik Indonesia) Vol. 1 (2024): Linguistik Indonesia dalam Lanskap Teknologi Digital
Publisher : Prodi Linguistik Indonesia UPN "Veteran" Jawa Timur

Show Abstract | Download Original | Original Source | Check in Google Scholar

Abstract

Corpus linguistics has long relied on the systematic collection and analysis of large text datasets to uncover patterns of language use. In the era of Big Data, this discipline undergoes a significant transformation, as the availability of massive digital corpora fundamentally changes the scope, methods, and applications of linguistic research. This study explores how Big Data reshapes corpus linguistics in terms of scale, representativeness, and analytical possibilities. Using examples from large-scale corpora derived from social media, online news, and digital archives, the paper demonstrates how linguistic patterns can now be analyzed with greater precision and across diverse contexts. The methodological section introduces computational approaches, such as natural language processing (NLP) tools and machine learning algorithms, that enhance corpus analysis. The results highlight novel findings in lexical variation, discourse structures, and language change over time, made possible by Big Data analytics. The discussion critically evaluates the advantages and challenges of this transformation, including issues of data quality, ethics, and accessibility. The conclusion suggests that corpus linguistics, when integrated with Big Data methodologies, not only advances linguistic theory but also has practical implications for education, policy, and digital communication.