Polyphonic orchestral recordings pose significant challenges in music information retrieval (MIR) due to their overlapping frequency ranges and timbral similarities among instrument families, which complicate multi-label instrument classification. Prior studies have explored the integration of convolutional neural networks (CNN)-based feature extraction with classical machine learning (ML) classifiers, often on monophonic or simpler datasets like IRMAS. But the integration of deep learning (DL) feature extraction, Autoencoder (AE)-based dimensionality reduction, and ML classifiers for polyphonic orchestral instrument recognition remains underexplored. This study proposes a hybrid framework utilizing a pre-trained Inception V3 CNN for feature extraction from mel-spectrograms, followed by an optional 50% dimensionality reduction via AE, and finally, classification with support vector machines (SVM) or extreme gradient boosting (XGBoost). Experiments were run on two polyphonic datasets, OpenMIC-2018 and Orchset. The results demonstrate that non-AE configurations generally outperform AE variants. These results extend prior studies such using polyphonic datasets. The results highlight the practical value of hybrid CNN-ML pipelines and the trade-offs of feature compression in MIR.
Copyrights © 2026