Speech Emotion Recognition (SER) has achieved remarkable performance under controlled, clean-audio conditions; however, its robustness in noise-laden, real-world environments remains insufficiently characterized. This study investigates the performance degradation of a Convolutional Neural Network (CNN)-based SER system on the RAVDESS dataset when subjected to synthetic noise at various Signal-to-Noise Ratio (SNR) levels (−5, 0, 5, 10, and 15 dB). We compare two widely used feature representations: Mel-Frequency Cepstral Coefficients (MFCC) and Log-Mel Spectrogram. Both models were trained exclusively on clean audio and evaluated under twelve noise conditions using Additive White Gaussian Noise (AWGN) and babble noise. Experimental results on the RAVDESS dataset (2,880 samples, 8 emotion classes) reveal a distinct asymmetry: the MFCC-based CNN achieves a 48.84% clean-audio accuracy with a maximum degradation of 27.08 percentage points (pp). Conversely, the Log-Mel Spectrogram model achieves a higher clean baseline of 66.67% but suffers a severe drop of up to 47.69 pp under noise, approaching the random baseline of 12.5%. These findings demonstrate that MFCC features offer superior robustness to additive noise due to implicit spectral smoothing via mel filterbanks and Discrete Cosine Transform (DCT), despite exhibiting lower clean-audio discriminability. This research highlights a fundamental trade-off between feature discriminability and noise robustness in uncontrolled acoustic environments.