Sammanfattning

Sound event localization and detection (SELD) is the task of identifying sound events and estimating their direction of arrival simultaneously. This study investigates whether explicit spatial feature fusion improves joint classification and localization performance compared to implicit representations in a stereo based system, examining three spatial cues: interaural phase difference (IPD), interaural level difference (ILD), and mid-side representations across single-source and dual-source reverberant conditions. A convolutional recurrent neural network is trained on a synthetically generated dataset derived from the Google Speech Commands corpus, using geometric simulation for the single-source case and Pyroomacoustics for the dual-source case. In the single-source condition, differences between configurations are small, with mid-side representations producing the best SELD score and IPD the lowest direction-of-arrival (DOA) error. In the dual-source reverberant condition, the differences are larger. Without IPD, no configuration achieves meaningful localization, with DOA errors around 44◦and sensitive F-scores, which measure correct class prediction within 20◦of the true azimuth angle. Adding IPD reduces DOA error to 13.66◦ and raises the location-sensitive F-score to 0.342. The full feature set (L, R, Mid-side, IPD, ILD) achieves the best result, with a DOA error of 13.33◦ and a SELD score of 0.362. The result analysis shows that IPD is essential for localization in reverberant multi-source conditions and that implicit representations alone are not sufficient. ILD contributes negligibly at a 2 cm microphone spacing, while mid-side features provide the most consistent benefit to classification.

Utforska vidare

Liknande uppsatser

Uppsatser med liknande ämnen och nyckelord.