AI-powered methods now make it possible to replicate voices, speech patterns, and emotional expressions with ever-greater realism and to manipulate spoken content in a targeted manner. Public figures are particularly vulnerable, as their voices are readily available and can be misused for disinformation, fraud, or targeted influence. Existing detection methods primarily look for general characteristics of synthetic or manipulated speech content. However, with the rapid advances in generative AI, these approaches are increasingly reaching their limits.
Objectives and Approach
The PADSE project is therefore developing a system for the detection, analysis, and classification of synthetic and manipulated speech content that employs a person-centric approach. The goal is to verify deceptively realistic speech manipulations not only based on general deepfake characteristics but also by comparing them to a comprehensive reference profile of the affected person.
Using various analytical technologies, the system creates reference profiles for persons to be protected, which include voice, speaking style, emotional expression, and linguistic patterns such as individual speaking style, authorship, and topic selection. The system then checks new content not only for general deepfake characteristics but also for consistency with the respective reference profile, while also taking into account the origin as well as the contextual and situational context of the content.
Innovations and Perspectives
With its person-centered approach, PADSE expands the traditional detection of audio deepfakes to include a new perspective. Instead of searching exclusively for general characteristics of synthetically generated speech, the system also evaluates how well a recording matches a person’s individual profile. This is intended to enable more reliable detection and better classification of audio manipulations and synthetically generated speech, even in the face of the challenges posed by new generative AI methods.
The modular platform developed in the project, featuring open interfaces, enables the integration and further development of individual components as well as adaptation to different use cases. Testing in a journalistic environment ensures a high degree of practical relevance and lays the foundation for evaluating the developed methods under real-world conditions and transferring them to other fields of application.
PADSE thus makes an important contribution to the reliable detection of manipulated and synthetically generated speech content and to the fight against disinformation.