Sriwijaya Journal of Radiology and Imaging Research
Vol. 4 No. 1 (2026): Sriwijaya Journal of Radiology and Imaging Research

Vision Transformer Reconstruction for Super-Resolution and Artifact Reduction in 1.5-Tesla Fast-Spin-Echo Pelvic MRI: A Diagnostic-Accuracy Study

Muhammad Rusli (Department of Radiology, Al Bayan Medical Center, Kendari, Indonesia)
Febria Suryani (Department of Health Sciences, CMHC Research Center, Palembang, Indonesia)
Desiree Montesinos (Department of Women and Child Welfare, Lira State Hospital, Lira, Uganda)



Article Info

Publish Date
07 Aug 2026

Abstract

Introduction: Fast-spin-echo (FSE) MRI is the reference modality for pelvic evaluation, but high-resolution acquisition is slow and motion-prone. Deep-learning reconstruction may recover diagnostic quality from short, motion-tolerant acquisitions, yet most evidence relies on image-similarity indices rather than radiologist performance. We validated a Vision-Transformer (ViT) reconstruction for simultaneous super-resolution and motion-artifact reduction in 1.5-Tesla T2-weighted FSE pelvic MRI. Methods: In this retrospective diagnostic-accuracy study (STARD 2015) at a tertiary hospital in Palembang, Indonesia, 450 examinations (development n=360; test n=90) were analysed. A U-shaped shifted-window ViT reconstructed high-resolution images from retrospectively degraded low-resolution/motion-corrupted inputs. Two blinded radiologists scored each test case under native low-resolution, UNet- and ViT-reconstructed conditions against a histopathology/expert-consensus reference standard. Sensitivity, specificity, predictive values, AUC, likelihood ratios (95% CI), inter-reader kappa, DeLong and McNemar tests, and multivariable logistic regression were computed. Results: Target prevalence was 53.3%. ViT reconstruction achieved sensitivity 93.8% (95% CI 83.2–97.9), specificity 88.1% (75.0–94.8), AUC 0.943 (0.894–0.992), LR+ 7.87 and LR− 0.071, versus AUC 0.881 (UNet) and 0.751 (low-resolution); ViT vs low-resolution DeLong p<0.001, McNemar p<0.001. Inter-reader agreement rose from kappa 0.49 to 0.87. ViT gave the best fidelity (PSNR 34.82 dB; SSIM 0.941; p<0.001 vs UNet) at 0.15 s/slice. Sub-centimetre lesions (OR 3.84, p=0.006) and severe motion (OR 2.97, p=0.029) independently predicted error. Conclusion: A shifted-window Vision Transformer recovered diagnostic-quality pelvic FSE MRI from short, motion-tolerant acquisitions, significantly improving radiologist lesion detection and inter-reader agreement over convolutional reconstruction. The real-time, PACS-compatible pipeline is promising for high-throughput pelvic MRI and warrants prospective validation.

Copyrights © 2026






Journal Info

Abbrev

sjrir

Publisher

Subject

Dentistry Health Professions Medicine & Pharmacology Neuroscience Physics

Description

Focus Sriwijaya Journal of Radiology and Imaging Research (SJRIR) focused on the development of medical sciences especially radiology & imaging research for human well-being. Scope Sriwijaya Journal of Radiology and Imaging Research (SJRIR) publishes articles which encompass all aspects of basic ...