Sound Scene Generation (SSG) focuses on the automatic synthesis of artificial acoustic scenes. We introduce SsgCaps, a publicly available dataset of human‑engineered sound scenes, each paired with a precisely structured prompt that guides the sampling process. The prompts are drawn from a predefined action‑based typology, enabling extensive sampling while preserving plausibility.
SsgCaps originates from the unpublished reference dataset of Task 7 in the 2024 DCASE Challenge, which contained both private and public‑domain audio samples. In contrast, SsgCaps includes only public‑domain audio, allowing us to release the dataset openly to the community.
To make the dataset useful, we first explain the rationale behind the prompt and dataset structure. We then conduct a comparative quantitative analysis of the two dataset versions by comparing them to audio synthesized by SSG algorithms submitted to the challenge, using Fréchet Audio Distance (FAD), Kernel Audio Distance (KAD), and perceptual ratings. The analysis reveals only minor differences, supporting our recommendation to adopt the open version for future benchmarking of SSG algorithms.
Review