How AI helps people with hearing loss detect crucial sounds at home

How AI helps people with hearing loss detect crucial sounds at home

For someone with hearing loss, sounds like a smoke alarm or doorbell can be life-critical information. That’s why Estonian company Soundfree, in collaboration with STACC, is developing an AI-powered solution that can recognize important household sounds.

STACC’s data scientist Johanna Kaare developed an initial algorithm that automatically extracts the relevant signal from a user-recorded audio clip, enabling the system to learn and recognise the specific sounds that occur in the user’s own home. Since then, the solution has evolved further: the latest development phase added features for assessing the quality of recorded clips and the confidence of the model’s sound classification. In other words, the system not only says what sound it is, but also how sure it is about that classification.

A solution that adapts to every home

No two homes sound alike. One’s oven might beep softly, while another’s doorbell may have a much lower tone. That’s why Soundfree is building a system that allows users to train the AI to recognise the specific sounds in their environment.

To do this, the user records a clip of a desired sound – like a microwave beep or doorbell chime. The algorithm developed by STACC then automatically isolates the relevant signal from the clip, filtering out background noise. This makes the system accessible even for people with hearing loss, who may not be able to listen to the clip or edit it themselves.

The system uses several complementary analysis methods to detect relevant audio signals and distinguish target sounds even in the presence of background noise. Loud and distinct signals (such as a smoke alarm) are generally easier to separate from the background, while softer sounds or high levels of background noise make the task more challenging. Using multiple analysis methods helps the system handle different types of sounds and recording conditions.

How to assess whether a recording is high quality?

One of the practical challenges is that recordings made at home aren’t always high-quality. The sound might be too quiet or mixed with unrelated noises – for example, a doorbell recording might also capture a knock or an ambulance siren.

To prevent such confusion, a system was developed that automatically assesses the quality of sound files. The system analyses the similarity between recordings of the same sound class. If one recording differs significantly from the others, it may indicate that the training dataset contains the wrong sound or a recording of unsuitable quality.

A model that knows how confident it is

Another important development is confidence estimation. When the system identifies a sound (such as a doorbell), it also estimates how confident it is in that classification. Confidence is expressed as a numerical score indicating how certain the model is about its prediction. The higher the score, the more confident the model is that its classification is reliable.

Why does that matter? If confidence is low, support teams can recommend adding more recordings or retraining the model. While these confidence scores aren’t currently shown in the user interface, they are logged in the system to help identify problematic cases.

Several different evaluation methods are used to assess the model’s reliability. These make it possible to check how consistently the model recognises sounds and identify situations where its prediction may require further attention. This helps ensure that the system’s decisions are as reliable as possible in real-world situations.

Technical challenge: automation that doesn’t stall

The process may sound complex – audio cleaning, trimming, evaluation, and model training – but it all needs to happen fast, especially if we want to ensure a smooth user experience.

That was the biggest technical challenge of this development phase. Today, the system processes a new audio sample in under one second, and training a personalised model takes less than a minute. At the same time, the required computational resources turned out to be significantly lower than expected.

To make everything work in real-time and with minimal delay, the pipeline was designed so that audio cleaning, trimming, and evaluation happen sequentially in an optimised order.

AI that improves everyday life

Ultimately, the goal is simple: to provide people with hearing loss with a safer and smoother everyday life. If the model can reliably recognise a smoke alarm or notify someone when the oven beeps, the risk of dangerous situations is reduced and quality of life improves.

STACC’s data scientist Johanna Kaare says that work on the solution continues, but the results so far are promising. “We’ve managed to bring the training time for one personalised sound model down to under a minute – and done so at a fraction of the expected computational cost,” she says. “It shows that even a technically complex solution can be fast, efficient, and user-friendly.”

If you’re looking for a partner who can turn a complex idea into a working, user-friendly data solution, get in touch with us! STACC’s data scientists can help bring your vision to life in a way that’s technically sound and practically usable.

KALEV KOPPEL

Contact us

KALEV KOPPEL

CEO
+372 515 9966
kalev.koppel@stacc.ee