This paper evaluates a cross-dataset framework for parcel segmentation, next-day last-mile operational-intensity forecasting, and capacity-oriented error analysis. The operational study uses 4,514,661 delivery tasks and 6,136,147 pickup tasks from LaDe-D and LaDe-P across five cities (May–October 2022), while the visual study uses 2,197 Package Segmentation images with 7,643 annotated package instances. A segmentation model estimates package count, foreground area, and instance-area dispersion to construct an additive workload descriptor. Because the images are not paired with operational records, city-day visual variables are represented as cross-dataset distributional proxies derived from empirical-rank mapping. Segmentation baselines are evaluated independently of forecasting. The random-forest pixel classifier achieves the best segmentation performance (IoU 0.5323, Dice 0.6450), outperforming YOLO11n-seg (IoU 0.3439, Dice 0.4591). Operational intensity is represented by the first principal component of delivery orders, active couriers, areas of interest, regions, and area types, explaining 89.19% of training variance. On a 230 city-day chronological holdout, ElasticNet Operational achieves the best forecasting accuracy (MAE 2.0521). Within XGBoost models, Vision-Prior records MAE 2.4988, slightly outperforming Delivery+Pickup (MAE 2.5154) and a shuffled-prior control (MAE 2.5973). However, the 0.0166 MAE improvement is not statistically significant under a paired moving-block bootstrap (95% CI: −0.0558 to 0.0319). A separate analysis of 6,112 Amazon routes and 1,457,175 packages similarly shows only a 0.24% MAE reduction from parcel-mix features. Overall, the proposed proxy provides only limited gains, while the operational ElasticNet remains the strongest baseline. Time- and location-paired visual observations are needed to establish practical operational benefits from computer vision.