earshot

Put a song in. Take it apart, sing on it, or watch it with the words.

Fetched through the Internet Archive, because this machine cannot reach YouTube directly. It only works for videos the archive kept a copy of, and it checks before downloading so you find out in two seconds rather than after the wait.

or

Everything the model can find is separated in one pass, so there is nothing to choose up front. What it found comes back below.

This runs a source-separation model on a small machine with no graphics card, at roughly two to four times realtime — a three minute song takes about ten minutes. One job runs at a time and the rest queue, because two at once does not fit in the memory.

Anything the song does not contain comes back nearly silent rather than missing, so the level of each stem is shown and the empty ones are marked. A piano stem from a song with no piano is a real output and a useless one.

Files are deleted after 6 hours. There is also a dialogue meter for the other half of this project.