PADSE – Person-centric, audio- and speech-based detection of deepfakes

AI-powered methods now make it possible to replicate voices, speech patterns, and emotional expressions with ever-greater realism and to manipulate spoken content in a targeted manner. Public figures are particularly vulnerable, as their voices are readily available and can be misused for disinformation, fraud, or targeted influence. Existing detection methods primarily look for general characteristics of synthetic or manipulated speech content. However, with the rapid advances in generative AI, these approaches are increasingly reaching their limits.
 

Objectives and Approach

The PADSE project is therefore developing a system for the detection, analysis, and classification of synthetic and manipulated speech content that employs a person-centric approach. The goal is to verify deceptively realistic speech manipulations not only based on general deepfake characteristics but also by comparing them to a comprehensive reference profile of the affected person.

Using various analytical technologies, the system creates reference profiles for persons to be protected, which include voice, speaking style, emotional expression, and linguistic patterns such as individual speaking style, authorship, and topic selection. The system then checks new content not only for general deepfake characteristics but also for consistency with the respective reference profile, while also taking into account the origin as well as the contextual and situational context of the content.
 

Innovations and Perspectives

With its person-centered approach, PADSE expands the traditional detection of audio deepfakes to include a new perspective. Instead of searching exclusively for general characteristics of synthetically generated speech, the system also evaluates how well a recording matches a person’s individual profile. This is intended to enable more reliable detection and better classification of audio manipulations and synthetically generated speech, even in the face of the challenges posed by new generative AI methods.

The modular platform developed in the project, featuring open interfaces, enables the integration and further development of individual components as well as adaptation to different use cases. Testing in a journalistic environment ensures a high degree of practical relevance and lays the foundation for evaluating the developed methods under real-world conditions and transferring them to other fields of application.

PADSE thus makes an important contribution to the reliable detection of manipulated and synthetically generated speech content and to the fight against disinformation.

Responsibilities of the Project Partners

Fraunhofer IDMT

 

  • Project coordination
  • Development of a modular and scalable system architecture
  • Development of core technologies for detecting speech synthesis and audio manipulation, speaker characteristics, faces, and emotion profiles
  • Source and decontextualization analysis
  • Selection and adaptation of privacy-enhancing technologies and C2PA-based content authentication

Deutsche Welle

 

  • Iterative requirements analysis in real-world newsroom workflows
  • Coordination of the creation and annotation of training and test data
  • Integration of automatic speech recognition and context enrichment, including person relationships
  • Development of the demonstrator’s GUI/front end and integration into existing verification platforms such as Truly Media
  • Conducting user evaluations and disseminating the results

Bauhaus-Universität Weimar (Webis Group)

 
  • Text classification and detection of style and topic patterns
  • Person-specific text synthesis detection and the bimodal coupling of speech and text analysis
  • Context and decontextualization analysis



Funding

Funded by the Federal Ministry of Research, Technology and Space BMFTR under the funding code 16KIS2468K