SPEC2SPATIAL: A TIME FREQUENCY SPATIAL ATTENTION NETWORK FOR BINAURAL AUDIO SYNTHESIS

Submitted to Interspeech 2026

The following video demonstrates the performance of our model in binaural audio synthesis for unseen speakers. Imagine yourself positioned as the blue listener at the center of the scene. A red speaker continuously speaks while moving along a trajectory around you. As the position and distance of the red figure change, the direction and distance of the sound you perceive will also change accordingly. Please enjoy it.
Note: The silence video is from the dataset in BinauralSpeechSynthesis. [1]

⚠️ Please turn on the sound and wear headphones while watching this video.

Some samples from Binaural Speech Dataset

Mono Groundtruth DSP WaveNet WarpNet NFS BinauralGrad Spec2Spatial (Ours)
Sample1
Sample2
Sample3
Sample4
Sample5

⚠️ Please use headphones to listen to these audios.

Note: To provide readers with a more intuitive comparison, we also present the results of traditional DSP methods. The HRTF dataset we used was collected by Bill Gardner and Keith Martin from MIT [2]. This dataset was recorded using KEMAR at a distance of 1.4 meters. However, since we do not know detailed information about the specific room environment in which the Binaural Speech Dataset was collected, the DSP method cannot accurately simulate room reverberation.

Our model can also perform binauralization on out-of-distribution audio such as music.

Music1
Music2

⚠️ Please use headphones to listen to these audios.