Audio recognition
Case study: how I turned a model that recognizes sounds into a product feature that has labeled over a million tracks. The code is private, so this page covers the decisions.
- 1M+
- tracks and clips recognized
- 92%
- labeled with no manual correction
- 100 ms
- per track or clip
- < 50 MB
- on disk
- Local
- runs on the user's Mac
The problem
Before a mix can start, someone prepares the session. A project can hold dozens of tracks, and each one gets named, colored and routed (sent to the right group of channels) by hand.
What it does
In fMusic, it recognizes the instrument on every track, for mix engineers and music producers. In fPost, it sorts every clip into dialogue, music or effects, for audio post production studios. Then it names everything, so the session opens already organized.
My role
I designed the recognition system around the model and built it side by side with Forte's chief technology officer. I designed the naming algorithm on my own. I also wrote the requirements an engineer followed to fine tune the model.
Decision 1
Show when the model is unsure. A wrong label that looks confident costs the engineer time, because they have to find it before they can fix it. So any prediction below 90% confidence appears in grey, and the user makes the call.
Decision 2
Combine several predictions. The system takes several of the model's predictions for each track and keeps the answer they agree on. A single wrong reading carries less weight.
Decision 3
Run everything on the user's Mac. Privacy matters in this industry, so the model runs entirely on the user's Mac. That set the budget: under 50 MB and about 100 ms per track.
Result
More than a million tracks and clips recognized since launch. Fewer than 8% needed a manual correction.