Siva Sathya S
Department of Computer Science, School of Engineering and Technology, Pondicherry University, Puducherry, India

Published : 1 Documents Claim Missing Document
Claim Missing Document
Check
Articles

Found 1 Documents
Search

Structured Nursing Handover Report Generation from Clinical Speech using Fine-Tuned XLSR-53 and T5: A Benchmarking Study Sasikala D; Siva Sathya S; Niranjan Kumar D; Vignesh S
Journal of Electronics, Electromedical Engineering, and Medical Informatics Vol 8 No 4 (2026): October
Publisher : Department of Electromedical Engineering, POLTEKKES KEMENKES SURABAYA

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.35882/jeeemi.v8i4.1778

Abstract

Accurate nursing handovers are critical for patient safety, as miscommunication during shift transitions leads to irreversible clinical errors. This work proposed an end-to-end pipeline that converts unstructured clinical nursing speech into standardized handover reports using a fine-tuned XLSR-53 acoustic model and T5-base text-to-text transformer. An Australian English clinical corpus of 200 synthetic nursing handover recordings from the CSIRO data access portal was utilised in this work. This benchmarking study was conducted within the CSIRO synthetic Australian English nursing handover corpus and does not represent a broad cross-domain clinical ASR benchmark. A domain-specific benchmarking study across seven state-of-the-art ASR architectures (Whisper Tiny/Base/Small, Wav2Vec2 Base/Large, HuBERT Large, XLSR-53) was conducted using this corpus. The experimental results further revealed XLSR-53 as the optimal architecture for clinical nursing speech recognition. A partial layer-freeze strategy was adopted in XLSR-53 by freezing the first 12 of 24 encoder layers, empirically validated through an ablation study with five freeze configurations (L=0, 6, 12, 18, 24). XLSR-53 preserves cross-lingual phonetic representations while enabling clinical vocabulary adaptation. A clinically motivated evaluation framework using curated medical vocabulary terms computes Medical Precision, Recall, and F1-Score along with standard WER, CER, and PER to assess reliability in clinical term recognition. Benchmarking against Google Health AI's MedASR zero-shot revealed that the proposed system XLSR-53 (L=12) achieved 17.15% WER against MedASR's 28.87% (p<0.001) with a Medical F1-Score of 0.98 and ROUGE-L of 0.92. Although results were obtained on synthetic Australian English speech, performance under real clinical conditions with background noise, overlapping speakers, and spontaneous interruptions requires further validation.