Separating upgoing (reflection) energy from downgoing (source-side and multiple) energy is a critical preprocessing step in reflection and vertical seismic profile (VSP) seismology, yet classical frequency–wavenumber, median, and Radon-based filters degrade sharply under lateral velocity variation, topography, and spatial aliasing. This paper reports a systematic, same-dataset comparison of seven wavefield-separation algorithms: a simple convolutional network (CNN), U-Net, bidirectional long short-term memory (BiLSTM), Transformer, ResNet, an MLP with PCA dimensionality reduction, and the classical f–k filter trained and evaluated on 201 synthetic acoustic shot gathers and stress-tested on an independent 240-gather cross-dataset. BiLSTM achieved the best in-distribution performance (correlation = 0.9822, SNR = 14.52 dB) and the smallest relative degradation (41.2%) under domain shift, while U-Net was the strongest convolutional architecture (correlation = 0.7600) and the classical f–k filter performed worst (correlation = 0.3003, SNR = −0.37 dB). All models lost substantial accuracy on the cross-dataset, confirming that domain shift not architectural capacity is the principal barrier to field deployment. The study contributes a reproducible, consistently evaluated benchmark; a rigorous cross-dataset generalization test rarely reported in the literature; and quantitative evidence that recurrent and attention-based sequence models outperform convolutional counterparts for 1-D trace-wise wavefield separation. The findings motivate transfer learning, physics-informed regularization, and larger, more diverse training sets as the next steps toward field-ready deployment.