Automated visual inspection of welded joints increasingly uses supervised deep-learning models, particularly YOLO object detectors. Although effective for defect localization, these approaches require domain-specific annotated data and provide limited engineering-oriented textual explanations. This study compared supervised YOLO-based pipelines with zero-shot Vision-Language Models (VLMs) for image-based weld-defect severity assessment. A public Welding Defect–Object Detection dataset containing 1,983 images was used. A four-level ordinal severity rubric was developed as an AWS D1.1-informed engineering heuristic rather than a direct implementation of AWS acceptance criteria because the source images lacked consistent physical scale calibration. Two supervised detectors, YOLOv8n and YOLO26n, and two zero-shot VLMs, NVIDIA Nemotron and Meta Llama 3.2 Vision-90B, were evaluated. All pipelines processed 401 validation-and-test images, while primary comparison against independently rated human references used an adjudicated subset of 100 images. Human inter-rater agreement was high (Cohen’s κ = 0.839; weighted κ = 0.929). Nemotron achieved the highest accuracy (55.0%), followed by YOLOv8n (50.0%), YOLO26n (48.0%), and Llama Vision-90B (38.0%), and the highest Level-4 recall (0.871). Both YOLO pipelines showed zero recall for Level 2. Although unadjusted McNemar testing found p = 0.030 for Nemotron versus Llama Vision-90B, no pairwise comparison remained significant after Holm correction. VLM explanation quality was moderately associated with classification correctness (r = 0.533; p = 0.0001). None of the methods was sufficiently reliable for standalone engineering acceptance decisions, but their complementary failure patterns support VLMs as an auxiliary review layer within human-supervised weld inspection.
Copyrights © 2026