As interconnected devices proliferate, secure and efficient pairing methods are critical. Environmental acoustic signals offer a promising solution, but their effectiveness depends on robust audio features that perform well across varying conditions. This study investigates optimal audio features for fingerprinting, focusing on synchronized audio in time-frequency domains. Six diverse datasets were collected across controlled environments to simulate real-world scenarios. Thirteen audio features were extracted and analyzed for robustness across distances, devices, scenes, and sample lengths. Cosine similarity assessed consistency, while the Youden index determined thresholds. The Mel Spectrogram, particularly with 5-second samples, achieved an AUC of 0.8758 and a J-statistic of 0.7278. Augmenting it with Tonnetz and spectral bandwidth yielded the highest performance (AUC: 0.9346, J-statistic: 0.7955, Accuracy: 0.8320, recall: 0.9648), demonstrating the potential of combining robust base features with complementary acoustic characteristics for reliable device pairing.
Copyrights © 2026