DSing

DAMP Sing

DSing (Roa-Dabike & Barker, 2019) is an unaccompanied singing dataset created to fill a gap in automatic lyrics transcription datasets. It is derived from the Smule DAMP-MVP 300x30x2 dataset, a collection of thousands of solo-singing karaoke recordings.

Construction Process

I built this dataset following these steps:

  1. Source data — I obtained the Sing! 300x30x2 dataset from the DAMP datasets collection.
  2. Language filtering — I filtered the recordings and kept only songs with English lyrics.
  3. Lyrics download — Using each song’s unique ID, I downloaded the corresponding lyrics from the Smule website.
  4. Lyrics alignment — I aligned the audio recordings with the downloaded lyrics.
  5. Audio segmentation — I split the audio into complete lyric phrases.

Download the Data

The DSing repository contains the following directories:

  • DSing-Kaldi-Recipe/: Kaldi recipe for the DSing ASR task (the baseline system).
  • DSing-preconstructed/: the DSing dataset segmentation.

Getting started

  • Get the audio dataset: Request permission and download the audio from Smule DAMP-MVP 300x30x2.
  • Get the dataset segmentation: The train/dev/test splits used in our paper are in DSing-preconstructed/.
  • Run the baseline system: See DSing-Kaldi-Recipe/ for instructions on training and evaluating the baseline Kaldi ASR system.

References

2019

  1. Conference
    interspeech2019.png
    Automatic Lyric Transcription from Karaoke Vocal Tracks: Resources and a Baseline System
    Gerardo Roa-Dabike and Jon Barker
    In Interspeech 2019, 2019