TTBA: Spatial Prompted Text to Binaural Audio Generation Using Transformer

Submitted to Interspeech 2026
Proposed End-to-End Binaural Audio Generation Model Architecture

Figure 1. Diagrams of the overall structure of our proposed text to binaural audio model (TTBA).

Model Configuration

  • Transformer-ALM: Autoregressive Transformer with num_layers = 24, dim = 1536, num_heads = 24.
  • Text Encoder: T5-large with max_length = 128, cond_dim = 1536.
  • Audio Encoder/Decoder: EnCodec 16 kHz with 4 codebooks × 2048 bins.
  • Optimizer: Adam with learning rate = 5e-5, betas = (0.9, 0.999), weight decay = 1e-3;
    Learning rate scheduler = InverseLR (inv_gamma = 1e6, power = 0.5, warmup = 0.99).
  • Teacher Model: Autoregressive Transformer with num_layers = 48, dim = 1536, num_heads = 24;
    Audio Encoder/Decoder: EnCodec 16 kHz with 4 codebooks × 2048 bins.

This demo page compares text-to-binaural audio generation by Stable_audio_open, SpatialSonic, Ours on BEWO-1M subsets (SS-Set, DS-Set, SD-Set, M-Set). Using identical text prompts, it provides a fair, side-by-side evaluation of spatialization quality and consistency. Please wear headphones while listening to these audios for the best experience.

Text to Binaural Audio Generation Comparison

SS-Set

"Vehicles at the front left of the scene hum and vibrate as they rev their engines."

Stable_audio_open

SpatialSonic

Ours

"Wind blows hard from the right of the scene."

Stable_audio_open

SpatialSonic

Ours

"The toilet flushing is located directly in front of the scene."

Stable_audio_open

SpatialSonic

Ours

DS-Set

"Water is turning on and running continuously at the front right of the scene, while a person is speaking amidst various laughter and clapping on the right."

Stable_audio_open

SpatialSonic

Ours

"At the right, a large dog barks while an engine revs directly in front of the scene."

Stable_audio_open

SpatialSonic

Ours

"A car engine revving on the left with a vehicle horn honking directly in front of the scene."

Stable_audio_open

SpatialSonic

Ours

SD-Set

"Thunderstorm sounds start on the right and slowly move to the front right."

Stable_audio_open

SpatialSonic

Ours

"A toilet flushes from the right to the left at a moderate speed."

Stable_audio_open

SpatialSonic

Ours

"Ocean waves move from front right to left at a fast speed."

Stable_audio_open

SpatialSonic

Ours

M-Set

"Whimpers of dogs can be heard from the front right, abruptly followed by a glass link, a thump, and clattering from the right side."

Stable_audio_open

SpatialSonic

Ours

"A man is laughing, slowly shifting from directly in front towards the left, while a horn honks several times on the left."

Stable_audio_open

SpatialSonic

Ours

"From the right side, the sound of a toilet flushing sweeps across to the left at a moderate pace, while gurgling water resounds from the front right."

Stable_audio_open

SpatialSonic

Ours

⚠️ Please use headphones for the best experience.