A growing number of Islamic content creators on TikTok now separate two message channels within a single video: the visual channel carries entertainment material such as comedy clips, memes, gameplay recordings, or aesthetic footage, while the audio channel carries excerpts of religious sermons. This hybrid format remains understudied because the existing literature on digital da’wah generally examines preachers who themselves perform humorously, rather than the detachment of humor from the religious source. This study aims to map the typology of visual–audio pairings in TikTok da’wah content, to explain how visual humor lowers the psychological resistance of Generation Z audiences, and to identify the risk of decontextualizing sermon excerpts. The study employs a descriptive qualitative approach using the netnographic method. Data were collected through non-participant observation of 570 videos from four Indonesian-language TikTok da’wah accounts between January and April 2026, supplemented by public metrics and audience comments, and analyzed through data condensation, data display, and conclusion verification. The findings identify four types of visual–audio pairing: consolation–admonition, absurd–reflective, gameplay–narrative, and aesthetic–devotional, distributed very unevenly across the corpus: gameplay–narrative accounts for 87.72 percent (500 videos), consolation–admonition together with its absurd–reflective variant for 7.02 percent (40 videos), and aesthetic–devotional for 5.26 percent (30 videos). Visual humor operates through the peripheral route: it diverts the cognitive resources ordinarily used to construct counterarguments while generating positive affect that eases message acceptance. This effect, however, is stronger for recall than for attitude change, so claims about the effectiveness of the format should be stated proportionally. The study further finds a decontextualization risk when sermons are cut into 20–40 second fragments stripped of conditions, context, and attribution; 82 percent of the corpus identifies the preacher in no form whatsoever. It concludes that the two-channel format is an effective rhetorical strategy for breaching scroll culture, yet demands citation ethics and adequate Islamic literacy from creators.
Copyrights © 2026