House-level visual socioeconomic information can provide a more detailed spatial understanding of residential areas and can support applied urban and commercial analysis. However, official socioeconomic data are often available only at broader administrative levels and may not capture variation among small residential clusters or individual houses. This study compares five instance segmentation models for detecting, segmenting, and classifying individual houses into lower, middle, and upper visual socioeconomic classes using high-resolution satellite imagery. The evaluated models are Mask R-CNN, Cascade Mask R-CNN, YOLO11n-seg, SOLOv2, and Mask2Former. The dataset consists of 1,209 training images with 7,052 annotations and 213 validation images with 1,353 annotations from residential areas in Banten and DKI Jakarta. Annotation reliability was assessed using 100 annotation pairs, producing a quadratic-weighted Cohen’s kappa of 0.8077. Model performance was evaluated using COCO metrics for both segmentation masks and bounding boxes. The results show that Cascade Mask R-CNN achieved the highest observed overall validation performance among the tested configurations. Under the current experimental setting, it produced the strongest combination of object-localization and mask-segmentation metrics. These findings show that comparing multiple instance segmentation models can help identify a more suitable method for house-level visual socioeconomic classification. Unlike previous studies that generally perform area-level socioeconomic estimation or building extraction alone, this study compares multiple instance-segmentation approaches for expert-defined socioeconomic classification at the individual-house level. The resulting output can serve as a supplementary visual socioeconomic layer that complements demographic, accessibility, and commercial data in applied spatial and market-development analyses.