TikTok has emerged as one of the dominant audiovisual platforms among Generation Z, characterized by short-form videos that integrate visual aesthetics, audio elements, and algorithm-driven personalization. This study examines the role of voice-over narration as an audiovisual communication strategy in fashion-related TikTok content. Specifically, the research analyzes four vocal dimensions—intonation, narrative structure, vocal personality, and emotional expression—and explores how these elements influence message clarity, narrative coherence, audience engagement, and the construction of creator identity. The study employs a descriptive qualitative approach through content observation of fashion creator Bella Clarissa and in-depth interviews with six Generation Z TikTok users and one fashion content creator. The findings demonstrate that voice-over narration functions as a narrative guide that mediates the relationship between visual presentation and audience interpretation, thereby enhancing the coherence and fidelity of fashion storytelling in short-form video content. Intonation and emotional cues contribute significantly to the construction of mood and affective engagement, while vocal personality strengthens perceptions of authenticity and reinforces creator branding and identity. The study further considers gender-based perspectives on emotional processing, particularly the tendency of female Generation Z audiences to exhibit greater sensitivity toward emotional and tonal nuances in vocal delivery. This perspective provides deeper insight into how voice-over narration shapes emotional resonance and audience connection within digital fashion communication. The study concludes that voice-over narration should not be understood merely as a supplementary audio feature, but rather as a central communicative strategy that enhances narrative comprehension, emotional engagement, and aesthetic appreciation in TikTok-based fashion content. These findings contribute to broader discussions on audiovisual communication, digital storytelling, and identity construction in contemporary social media environments. TikTok menjadi salah satu platform audiovisual paling berpengaruh di kalangan Gen Z dengan format video pendek yang menekankan perpaduan visual kreatif, musik, dan suara. Penelitian ini mengkaji fungsi voice over sebagai elemen audiovisual dalam konten fashion TikTok, dengan fokus pada empat komponen utama: intonasi, narasi, personalitas suara, dan ekspresi emosi. Penelitian ini menggunakan metode kualitatif deskriptif melalui observasi konten kreator fashion Bella Clarissa (@bellaclrs_) serta wawancara mendalam dengan enam mahasiswa Gen Z dan seorang fashion content creator. Hasil penelitian menunjukkan bahwa voice over berperan sebagai pengikat narasi yang membantu audiens memahami konteks visual, memperkuat makna, serta menciptakan kohesi antara suara dan gambar. Intonasi menentukan daya tarik awal konten, narasi membantu proses interpretasi, personalitas suara membangun identitas kreator, sementara emosi vokal menambah kedalaman pengalaman menonton. Penelitian ini juga mempertimbangkan karakteristik audiens perempuan Gen Z yang menurut teori Gender Differences in Emotional Processing cenderung lebih teliti dalam menangkap detail vokal, sehingga memberikan pemaknaan yang lebih mendalam terhadap penggunaan voice over. Penelitian ini menyimpulkan bahwa voice over bukan hanya pelengkap visual, tetapi merupakan strategi komunikasi penting dalam membentuk kejelasan pesan, identitas kreator, serta resonansi emosional dalam konten fashion digital.