Jejak digital: Jurnal Ilmiah Multidisiplin
Vol. 2 No. 4 (2026): JUNI-JULI

Karakteristik Yuridis Pelanggaran Hak Cipta dalam Penggunaan Data untuk Pelatihan Large Language Model (LLM) Generatif

Asep Supriyadi (Universitas Dirgantara Marsekal Suryadarma)
Aturkian Laia (Universitas Dirgantara Marsekal Suryadarma)



Article Info

Publish Date
16 Jul 2026

Abstract

The rapid development of generative artificial intelligence, particularly Large Language Models (LLMs), has created new copyright issues concerning the use of copyrighted works as training data without the authors’ permission. This study aims to examine the legal characteristics of copyrighted data used in LLM training, identify potential copyright infringements under Indonesian law, and analyze the regulatory challenges surrounding generative AI. The research employs a normative legal method using statutory and conceptual approaches, based on Law Number 28 of 2014 on Copyright, the Electronic Information and Transactions Law, and relevant academic literature. The findings indicate that data scraping, reproduction, and data processing for AI training may infringe the exclusive rights of copyright holders because such activities do not fall within the scope of fair use. The absence of explicit regulation on text and data mining creates legal uncertainty. Therefore, Indonesia should establish specific copyright exceptions, collective licensing mechanisms, and fair compensation to balance AI innovation with copyright protection.

Copyrights © 2026






Journal Info

Abbrev

jejakdigital

Publisher

Subject

Economics, Econometrics & Finance Education Languange, Linguistic, Communication & Media Social Sciences Other

Description

Jurnal Ilmiah Multidisiplin adalah jurnal elektronik dan cetak Open Access Journal yang diterbitkan oleh Indo Publishing setiap 6 kali dalam setahun menyediakan forum untuk mempublikasikan artikel penelitian asli, artikel review dari kontributor, dan berita teknologi baru mencangkup multidisiplin ...