In this miniseries, we will discuss in more depth the sound classification technology behind our acoustic monitoring solutions. Which technologies are necessary to make sound classification work? And what are some of the major challenges we face? But, before we dive into the details, we should first ask ourselves: what is sound classification anyway?
Basically, a sound classification system labels sounds with one or more labels. For example, “this is the sound of a bird”. Or, more specifically, “this is the sound of a blackbird”. Sound classification is not limited to recognizing individual sound sources, though. For example, systems that classify sound scenes (“airport”, “park”), emotions (“sad”, “aggressive”) or musical genres (“rock”, “jazz”) exists as well.
As we know, humans are usually very good at classifying sounds using our ears and brains, although even humans can be fooled easily sometimes, as proven by this entertaining video. Nowadays, thanks to the rapid advances in machine learning, machines become better and better at this task too. Thanks to complex neural networks, sounds can often be classified with a very good accuracy nowadays (but more on that in later blogs).
Within the field of sound classification, two major sub-fields exist:
1. Audio tagging: assigns one or more tags on an audio clip, e.g., “this clip contains speech and a piano”
2. Sound event detection: assigns labels to an audio clip or audio stream including time limits, e.g., “this clip contains speech from 1.0 to 3.5 seconds, and a piano sound from 2.8 to 7.4 seconds”. Or, “we detected an alarm sound on Friday June 7th at 8:30 AM”.
The goals differ slightly between the two, and so do the technologies behind them. At Sound Intelligence, we usually aim at classifying sounds in live audio streams, so we always work with *sound event detection*.
Automated sound Event Detection, or sound classification in general, can be really helpful in a lot of cases compared to human evaluation. For example, when the amount of sounds is just too much to be reasonably evaluated by humans. Or, when the audio contains sensitive information. In the case of Sound Intelligence’s customers, a sound event detection often acts as an *assistant*: it can act as an ‘early warning’ sensor and increase situational awareness to e.g. security or nursing staff.
For example, with staff in a hospital or nursing home carrying out hourly rounds at night, an emergency that happens between rounds can be easily missed without sound event detection. Also, there is no need to listen in every time ‘a’ sound is detected in a room; only when a sound occurs that may indicate an emergency, an alert is sent to the staff. This way, the patients are being disturbed less, and privacy is increased.
Over the years, sound event detection using machine learning has gained more and more interest from the industry *and* the scientific community. So, in the next blog, we will go a bit more into detail on how sound event detection with machine learning actually works.



