Foreground-Background

Ambient Sound Scene Separation

Common ambient sound scenes usually consist of a great variety of sounds, where multiple short duration audio events occur on top of background sounds. For many applications, it would be of great interest to separate such events from the mixture, a task that we refer to as foreground-background ambient sound scene separation. To investigate such task, we have built four different deep learning-based models which depending on the configuration, can rely on the use of an auxiliary network in the separation process. They were trained with the acoustic-frontends presented in the following table:

Model Log-Mel spectrograms PCEN spectrograms Auxiliary network
M1
M1+
M2
M2+

Each model was tested under evaluation setups comprising mixtures of sound classes seen and not seen during training according to the following table:

Setup Foreground class Background class
seen unseen seen unseen
C1
C2
C3
C4

...
Mixture
...
Foreground ground-truth
...
Background ground-truth
...
M1

Foreground:

Background:

...
M1+

Foreground:

Background:

...
M2

Foreground:

Background:

...
M2+

Foreground:

Background: