BAST: Binaural Audio Spectrogram Transformer for Binaural Sound Localization

Read original: arXiv:2207.03927 - Published 8/9/2024 by Sheng Kuang, Jie Shi, Kiki van der Heijden, Siamak Mehrkanoon

🤔

Overview

Sound localization is crucial for human auditory perception, especially in reverberant environments.
Convolutional Neural Networks (CNNs) have been used to model the binaural human auditory pathway, but they struggle to capture global acoustic features.
To address this, the paper proposes a novel Binaural Audio Spectrogram Transformer (BAST) model to predict sound azimuth in anechoic and reverberant environments.
The BAST model is explored in two modes: BAST-SP with shared parameters and BAST-NSP with non-shared parameters.
The BAST model with subtraction interaural integration and hybrid loss achieves state-of-the-art performance, surpassing CNN-based models.

Plain English Explanation

The paper focuses on the challenge of sound localization in real-world environments, where sound waves bounce off surfaces and create reverberation. Accurate sound localization is essential for human hearing and various applications, such as audio telepresence.

Traditionally, Convolutional Neural Networks (CNNs) have been used to model the human auditory system and predict the direction of sound sources. However, CNNs have difficulty capturing the overall acoustic features that are important for sound localization.

To address this, the researchers proposed a new Binaural Audio Spectrogram Transformer (BAST) model. The BAST model takes in binaural audio (sound from both ears) and uses a transformer architecture to predict the direction of the sound source, even in reverberant environments.

The BAST model was tested in two different configurations: BAST-SP with shared parameters and BAST-NSP with non-shared parameters. The researchers found that the BAST model, with its unique features like subtraction interaural integration and hybrid loss, significantly outperformed traditional CNN-based models in sound localization accuracy.

The paper also explores the BAST model's performance in different acoustic environments, demonstrating its ability to generalize and its potential for real-world applications.

Technical Explanation

The paper proposes a novel Binaural Audio Spectrogram Transformer (BAST) model to address the limitations of Convolutional Neural Networks (CNNs) in modeling the binaural human auditory pathway for sound localization in reverberant environments.

The BAST model takes binaural audio spectrograms as input and predicts the sound azimuth (angular direction) using a transformer-based architecture. Two modes of BAST implementation are explored: BAST-SP with shared parameters and BAST-NSP with non-shared parameters.

The BAST model incorporates several key features:

Subtraction interaural integration: This operation enhances the model's ability to capture interaural cues, which are crucial for sound localization.
Hybrid loss: The model is trained using a combination of Mean Squared Error (MSE) and cosine similarity loss, which improves the overall sound localization performance.

The experimental results show that the BAST model significantly outperforms CNN-based models, achieving an angular distance of 1.29 degrees and a Mean Squared Error of 1e-3 at all azimuths, even in reverberant environments.

The paper also provides an exploratory analysis of the BAST model's performance in different acoustic environments, including left-right hemifields and anechoic vs. reverberant conditions. This analysis demonstrates the model's strong generalization ability and the feasibility of using binaural transformers for sound localization tasks.

Furthermore, the paper includes an analysis of the BAST model's attention maps, which provides insights into the localization process and the model's internal workings in a natural reverberant environment.

Critical Analysis

The paper presents a compelling approach to address the limitations of CNN-based models in sound localization by introducing the Binaural Audio Spectrogram Transformer (BAST) model. The key strengths of the research include:

Innovative architecture: The transformer-based BAST model shows significant performance improvements over CNN-based approaches, highlighting the potential of transformer models for auditory perception tasks.
Comprehensive evaluation: The paper thoroughly examines the BAST model's performance in both anechoic and reverberant environments, as well as its ability to generalize across different acoustic conditions.
Interpretability: The analysis of the BAST model's attention maps provides valuable insights into the localization process, which can aid in understanding and further improving the model.

However, the paper could be strengthened by addressing the following potential limitations:

Generalization to other datasets: The evaluation is primarily focused on a single dataset, and it would be beneficial to assess the BAST model's performance on a wider range of datasets to ensure its generalizability.
Computational complexity: While the transformer-based architecture offers improved performance, the inherent complexity of transformers may raise questions about their computational efficiency, especially for real-time applications.
Comparison to state-of-the-art methods: The paper could benefit from a more detailed comparison of the BAST model's performance against the latest state-of-the-art approaches in sound localization, to better contextualize the significance of the proposed method.

Overall, the paper presents a compelling and well-executed piece of research, demonstrating the potential of transformer-based models for binaural audio processing and sound localization. The insights gained from this work can contribute to the ongoing efforts to improve auditory perception capabilities in various applications.

Conclusion

The paper introduces a novel Binaural Audio Spectrogram Transformer (BAST) model that significantly outperforms traditional Convolutional Neural Network (CNN) approaches in sound localization, even in reverberant environments. The BAST model's ability to capture global acoustic features, combined with its subtraction interaural integration and hybrid loss, enables it to achieve state-of-the-art performance.

The paper's comprehensive analysis of the BAST model's performance across different acoustic conditions, as well as the insights gained from the attention map interpretation, provide valuable contributions to the field of binaural audio processing and sound localization. These advancements have the potential to impact a wide range of applications, such as audio telepresence, target speaker extraction, and sound event detection.

The paper's findings highlight the effectiveness of transformer-based architectures in modeling the complex binaural auditory system and pave the way for further research and development in this area.

This summary was produced with help from an AI and may contain inaccuracies - check out the links to read the original source documents!

Follow @aimodelsfyi on 𝕏 →

Related Papers

🤔

BAST: Binaural Audio Spectrogram Transformer for Binaural Sound Localization

Sheng Kuang, Jie Shi, Kiki van der Heijden, Siamak Mehrkanoon

Accurate sound localization in a reverberation environment is essential for human auditory perception. Recently, Convolutional Neural Networks (CNNs) have been utilized to model the binaural human auditory pathway. However, CNN shows barriers in capturing the global acoustic features. To address this issue, we propose a novel end-to-end Binaural Audio Spectrogram Transformer (BAST) model to predict the sound azimuth in both anechoic and reverberation environments. Two modes of implementation, i.e. BAST-SP and BAST-NSP corresponding to BAST model with shared and non-shared parameters respectively, are explored. Our model with subtraction interaural integration and hybrid loss achieves an angular distance of 1.29 degrees and a Mean Square Error of 1e-3 at all azimuths, significantly surpassing CNN based model. The exploratory analysis of the BAST's performance on the left-right hemifields and anechoic and reverberation environments shows its generalization ability as well as the feasibility of binaural Transformers in sound localization. Furthermore, the analysis of the attention maps is provided to give additional insights on the interpretation of the localization process in a natural reverberant environment.

8/9/2024

Binaural Selective Attention Model for Target Speaker Extraction

Hanyu Meng, Qiquan Zhang, Xiangyu Zhang, Vidhyasaharan Sethu, Eliathamby Ambikairajah

The remarkable ability of humans to selectively focus on a target speaker in cocktail party scenarios is facilitated by binaural audio processing. In this paper, we present a binaural time-domain Target Speaker Extraction model based on the Filter-and-Sum Network (FaSNet). Inspired by human selective hearing, our proposed model introduces target speaker embedding into separators using a multi-head attention-based selective attention block. We also compared two binaural interaction approaches -- the cosine similarity of time-domain signals and inter-channel correlation in learned spectral representations. Our experimental results show that our proposed model outperforms monaural configurations and state-of-the-art multi-channel target speaker extraction models, achieving best-in-class performance with 18.52 dB SI-SDR, 19.12 dB SDR, and 3.05 PESQ scores under anechoic two-speaker test configurations.

6/19/2024

💬

BAT: Learning to Reason about Spatial Sounds with Large Language Models

Zhisheng Zheng, Puyuan Peng, Ziyang Ma, Xie Chen, Eunsol Choi, David Harwath

Spatial sound reasoning is a fundamental human skill, enabling us to navigate and interpret our surroundings based on sound. In this paper we present BAT, which combines the spatial sound perception ability of a binaural acoustic scene analysis model with the natural language reasoning capabilities of a large language model (LLM) to replicate this innate ability. To address the lack of existing datasets of in-the-wild spatial sounds, we synthesized a binaural audio dataset using AudioSet and SoundSpaces 2.0. Next, we developed SpatialSoundQA, a spatial sound-based question-answering dataset, offering a range of QA tasks that train BAT in various aspects of spatial sound perception and reasoning. The acoustic front end encoder of BAT is a novel spatial audio encoder named Spatial Audio Spectrogram Transformer, or Spatial-AST, which by itself achieves strong performance across sound event detection, spatial localization, and distance estimation. By integrating Spatial-AST with LLaMA-2 7B model, BAT transcends standard Sound Event Localization and Detection (SELD) tasks, enabling the model to reason about the relationships between the sounds in its environment. Our experiments demonstrate BAT's superior performance on both spatial sound perception and reasoning, showcasing the immense potential of LLMs in navigating and interpreting complex spatial audio environments.

5/28/2024

🔄

A tunable binaural audio telepresence system capable of balancing immersive and enhanced modes

Yicheng Hsu, Mingsian R. Bai

Binaural Audio Telepresence (BAT) aims to encode the acoustic scene at the far end into binaural signals for the user at the near end. BAT encompasses an immense range of applications that can vary between two extreme modes of Immersive BAT (I-BAT) and Enhanced BAT (E-BAT). With I-BAT, our goal is to preserve the full ambience as if we were at the far end, while with E-BAT, our goal is to enhance the far-end conversation with significantly improved speech quality and intelligibility. To this end, this paper presents a tunable BAT system to vary between these two AT modes with a desired application-specific balance. Microphone signals are converted into binaural signals with prescribed ambience factor. A novel Spatial COherence REpresentation (SCORE) is proposed as an input feature for model training so that the network remains robust to different array setups. Experimental results demonstrated the superior performance of the proposed BAT, even when the array configurations were not included in the training phase.

5/15/2024