Clear semantic definition of laparoscopic video images is an important prerequisite for the development of context-sensitive computer-assisted surgery systems. Although the CholecSeg8k dataset contains pixel-wise annotations of 13 anatomical structures, making real-time predictions while retaining localization accuracy of the object boundaries is a difficult task in surgical informatics. The present study tries to solve these problems through the use of the lightweight one-stage YOLOv8m-seg neural network model along with an automatic pre-processing pipeline, which uses contour-filtering techniques to transform color-coded masks into normalized polygonal annotations. After the fine-tuning procedure, which included multi-scale geometric transformations such as random spatial rotations and scaling, the system performance was validated on a specific validation set. Mask mean average precision (mAP50) was measured at 0.955 and 0.793 (mAP50-95). The class-based metrics had values above 0.98 for the classes’ liver, abdominal wall, and grasper, while amorphous tissues such as blood and connective tissue were the most difficult to classify due to their higher morphological ambiguity. In the computational analysis, an average inference time of 13.2 milliseconds was recorded on a consumer-level NVIDIA RTX 3060 GPU, thus achieving a rate of 75 frames per second. From the results recorded, it is evident that anchor-free single-stage frameworks perform best in balancing latency and segmentation precision as opposed to multi-stage frameworks
Copyrights © 2026