Analysis and optimization of speech intelligibility

The technical transmission of speech is often overlaid with background noise and reverberation, for example in train station announcements, in mobile phone communications, or in-car infotainment systems. In media production and broadcasting, decisions about whether speech is sufficiently intelligible to listeners are often subjective.

We develop signal processing methods that ensure better speech intelligibility and reduced listening effort in a variety of applications.

The software solutions from Fraunhofer IDMT in Oldenburg ...

Fraunhofer IDMT’s developments are aligned with market needs and have already been integrated into various industry solutions through licensing agreements.

Through strong collaborations, concrete research results from Oldenburg’s hearing research community—such as those from Carl von Ossietzky University and the "Hearing4All" Cluster of Excellence—are incorporated into these solutions. As a result, our solutions deliver a personalized, optimized sound experience for people with and without hearing impairments.

Detection of speech in audio signals

Fraunhofer IDMT’s Speech Activity Detection (SAD) reliably identifies spoken segments in audio signals. For example, the technology streamlines the workflow for audio professionals, as it makes it easy to locate speech during editing without having to listen to every segment. The technology was integrated into Steinberg’s post-production software Nuendo as the "Dialog Detection" feature. It is also used in other solutions developed by the Oldenburg Branch, such as its in-house speech and speaker recognition, algorithms for reducing background noise, and as a privacy filter. SAD can be used to filter out non-speech segments during audio processing. Conversely, it can ensure that speech is not recorded in the first place, thereby protecting users’ privacy—for example, in public spaces.

»Dialog Detection« in Steinbergs Nuendo 12 mit Algorithmen des Fraunhofer IDMT in Oldenburg
© Steinberg Media Technologies GmbH
"Dialog Detection" in Steinberg's Nuendo 12: Algorithms developed by Fraunhofer IDMT in Oldenburg reliably detect speech activity in the audio signal, regardless of background noise.

Analysis and evaluation of speech intelligibility

using LEAP technology (Listening Effort Prediction for Acoustic Parameters)

By combining machine learning and psychoacoustic modeling, researchers at Fraunhofer IDMT in Oldenburg have developed new technologies to objectively measure and visualize listening effort for individual audio segments in real time. This makes it possible, for example, to verify consistently good speech intelligibility in media productions and streams. The technology is already supporting audio professionals in post-production across various production solutions, such as the "LE-Meter" plug-in from NUGEN Audio, the "Speech Intelligibility Meter" in RTW’s hardware-based audio measurement platform, and the Nuendo production solution from Steinberg.

The algorithms are based on the latest findings of the Oldenburg hearing research. They have been evaluated multiple times in recent years and validated through formal listening tests with participants of various age groups.

Das vom Fraunhofer IDMT entwickelte »LE-Meter«, das in das Plug-in »DialogCheck« von NUGEN integriert ist.
© NUGEN Audio Limited
The "LE-Meter", developed by Fraunhofer IDMT and integrated into NUGEN’s "DialogCheck" plug-in, enables sound engineers to visualize and optimize listening effort and speech intelligibility during post-production.

Optimizing Speech Intelligibility

using PSIO (Personal Speech Intelligibility Optimization) technologies

Various signal processing methods developed by Fraunhofer IDMT can be combined to improve speech intelligibility. The real-time prediction models described earlier detect speech activity and identify the objective listening effort required for an audio signal (e.g., a broadcast video). Speech separation technologies isolate speech content from background noise and split it into two sound objects that can be processed independently of one another. Background music and sound effects, for example, could be reduced according to personal listening preferences or needs to achieve better intelligibility of dialogue.

These technologies are ideally suited for use in media production, both for live broadcasts and in post-production. Media libraries, streaming services, and communication service providers can also offer their customers added value, ranging from alternative soundtracks with improved intelligibility to customizable solutions on devices. 

In addition, adaptive signal processing can enhance the speech signal when the technical transmission of speech is marred by background noise and reverberation. Especially when it matters most—for example, in emergency dispatch centers, public address systems, or conference and telephony solutions—not a single word should be lost.

Die Quellentrennungstechnologien trennen Sprachinhalte von Hintergrundgeräuschen und teilen sie in zwei Klangobjekte auf.
© Fraunhofer IDMT/Anika Bödecker
The source separation technologies developed by Fraunhofer IDMT separate speech from background noise and divide them into two sound objects.

Adjusting the sound to your personal preferences

Another technology developed by Fraunhofer IDMT allows users to customize speech and background sounds based on their own audio preferences to create a personalized listening experience. Using a particularly simple and elegant process, sound and dynamics can be personalized without requiring users to navigate complex submenus or adjust parameters.

The so-called "YourSound" technology is already being used for the individual configuration of media playback in cars and headphones. Once set via an intuitive user interface, the personalized parameters have a positive effect on the overall sound. This results in a better listening experience, regardless of playback volume or driving conditions. Thanks to Fraunhofer algorithms, the sound of music and movies is optimally tailored to individual sound preferences.

Sprachanteile und Hintergrundgeräusche können unabhängig voneinander nach persönlichen Klangvorlieben optimiert werden.
© Fraunhofer IDMT/Leona Hofmann
Speech segments and background noise can be optimized independently of one another according to personal sound preferences.

Application examples

 

Press Release / 22.10.2025

For optimal speech intelligibility in film and television

"Speech Intelligibility Meter" integrated into international broadcast production solution

 

Press Release / 27.5.2025

Making dialogs accessible for everyone

"Listening Effort Meter" helps in post-productions to avoid poor speech clarity.

 

Press Release / 1.6.2023

First-rate sound for every listener

Fraunhofer IDMT's method for personalized sound experience has now been successfully integrated in headphones as well.

 

Press Information / 7.7.2022

Intuitive sound personalisation in vehicles

Everyone hears differently – this applies inside a vehicle too. That is why Fraunhofer IDMT in Oldenburg has developed a technological concept for fast and individual sound adaptation, which has been integrated into the multimedia system of vehicles from the Mercedes-Benz Group AG.

 

Press Release / 23.5.2022

Seeing Speech

New algorithms from Fraunhofer IDMT form the basis for the »Dialogue Detection« in Steinberg Media Technologies’ latest version of its audio post-production software Nuendo. 

 

Press Release / 10.12.2020

Red when mumbling!

Intelligibility Meter enables objective measurement and display of speech intelligibility in media productions.

Simply explained: Evaluating and optimizing speech intelligibility

 

Video

Analyzing and optimizing speech intelligibility

What exactly do we mean when we talk about "better speech intelligibility"? Dr. Jan Rennies-Hochmuth explains how the analysis, evaluation, and improvement of speech intelligibility work.

 

Video

Source Separation at Fraunhofer IDMT

To separate the dialogue from the background noise and improve speech intelligibility in a new audio track, we use source separation technology. Dr. Jan Rennies-Hochmuth explains how it works.

Further information

 

Your sound, in every situation

Everyone hears differently. »YourSound« of Fraunhofer IDMT in Oldenburg enables users of audio devices to adjust audio to their own acoustic preferences – as easy as never before!

 

Voice Filtering

Reliable speaker differentiation in a few seconds

 

SI-Live – Real-time monitoring of speech intelligibility

Optimal speech transmission and smooth transitions between interlocutors in telephone calls and video conferences are important for high user acceptance.

 

In-ear AI

We develop the hearable for the smart industrial workplace.

 

Press Release / 4.11.2021

Better understanding

Tonmeistertagung 2021: Fraunhofer IDMT presents solutions for analysing, evaluating and improving speech intelligibility.

 

AdaptDRC

Software solution for real-time optimization of speech intelligibility in noisy environments

 

SITA – Better sound, less noise!

SITA addresses the main factors for poor speech intelligibility along the whole distribution chain and aims to eliminate existing barriers for the widest possible variety of target groups, applications and hearing scenarios with the help of innovative software technologies.

R&D-Services and Licencing

Please contact us in case you are interested in our expertise and services.

All solutions at a glance

Here you can find further information about our solutions of the Oldenburg Branch for Hearing, Speech and Audio Technology HSA.

Customized sound quality and clear speech intelligibility

Learn more about our ongoing research and development.