Monday, August 3, 2026


TECH


When the fake boss calls: a real-time warning system for fake videoconferences

Videoconferencing facilitates communication between different locations while still allowing people to look each other in the eye. However, more and more often we have to question whether the other person's eyes are real. This is because deepfake technologies are now capable of falsifying voices and images in real time more realistically than ever. In 2025, the CFO of a company in Singapore was ensnared by a deepfake Zoom call in which all the participants were AI-generated—including his boss. In the end, he got off lightly: The authorities were later able to recover the roughly 500,000 U.S. dollars he had transferred.

The Cyber Security Agency (CSA) in Singapore alone has recorded a massive increase in video fraud cases like this in the first quarter of 2026. The country saw over 1,200 cases in January alone. That number was 43 in the same month of the previous year. However, the problem is global. After analyzing over one billion cases worldwide, the 2026 Entrust Identity Fraud Report concludes: “Identity fraud is no longer a crime of opportunity. It has now become industrialized, globally organized and commercially optimized.” This is cause enough for Fraunhofer researchers such as Martin Steinebach to work intensively on this issue. Steinebach heads the Media Security and IT Forensics department at the Fraunhofer Institute for Secure Information Technology SIT in Darmstadt.

“Videoconferences represent a special challenge for deepfake detection,” says the expert. “Image and sound quality fluctuate with variations in the network. Because the data streams are compressed, those specific image artifacts often used by earlier detection methods to identify deepfakes are lost.” The challenge is further exacerbated by automatic blur filters, variations in lighting as well as noise and movement in the background. A detector cannot be allowed to produce false positives for these normal videoconferencing effects.

A local solution for real-time detection...In an ATHENE research project, Steinebach and his team of experts at Fraunhofer SIT and at the Fraunhofer Heilbronn Research and Innovation Center (HNFIZ) for Cybersecurity have developed a solution combining video and audio deepfake detection. The AI-based software is designed to continuously inform participants in security-critical video calls of the likelihood of a deepfake. The warning is visual: “If a meeting is currently raising many flags, the system indicates a probability that it is a deepfake,” the researcher says. “But then the person in the meeting still has to verify the other party's authenticity by asking specific questions or by calling back on a different channel.”

The analysis can run locally on a high-performance laptop without requiring the transfer of image or audio data to an external server. It is thus compliant with data protection requirements and is suitable for companies conducting confidential meetings, such as with executive boards, finance departments or external partners. “A computer with a modern graphics card and twelve gigabytes of graphics memory is sufficient for the local real-time analysis in our solution,” says Steinebach.

This solution should also prove interesting to videoconferencing systems providers who could integrate it as a plugin or as an embedded security feature in systems such as Teams, Zoom or comparable enterprise solutions. Another possibility would be a central corporate infrastructure with a secure server area that analyzes especially important meetings. The exact form depends on the use case, the available resources and the data protection requirements.

Researchers at Fraunhofer SIT first create deepfakes themselves to train the system. Shown here is an example comparison between the detection of the authentic researcher on the left and a version in which the face has been replaced with that of Tom Cruise-image above (© Fraunhofer SIT)

AI beats signal processing...“One observation we made right at the start was that traditional signal processing methods quickly reach their limits in real-time detection today. Machine learning was significantly faster and more efficient,” Steinebach reports. He feeds the self-learning system with real and manipulated audio and video data to enable it to independently recognize differences. In the audio domain, the team drew on resources including publicly available datasets containing roughly 19,000 real recordings and about 160,000 spoofed, i.e., manipulated recordings.

What is critical here is that the training data should replicate a realistic videoconferencing environment as closely as possible. Typical effects that occur in videoconferences, such as compression, fluctuating quality and blur filters, must therefore also be simulated. The AI term for this is augmentation, explains Steinebach: “We apply interference to the data such as occurs in the real world and train the system to deal with it.” Because this is a specialized model designed for a clearly defined task, the training is significantly less expansive than in large language models. “It was especially interesting that it was almost more time-consuming and difficult to create deepfakes than to detect them in a real-time system like this.”

Outlook: practical testing of technology and processes...The Fraunhofer SIT demonstrator is currently still in the proof-of-concept phase. In the next step, the researchers plan to collaborate with companies and videoconferencing system providers to reliably integrate this technology in real-world infrastructures with appropriate data protection and user-friendliness. Legal questions also have to be clarified, such as whether meeting participants should be required to consent to the analysis or how the terms of use for a videoconferencing system would reflect this function.

Martin Steinebach has mixed feelings about the future development of deepfakes. On the one hand, there is a risk that even security-critical processes such as video identification by banks could soon be falsified using deepfakes. On the other hand, the risk can be mitigated by cryptographic security, additional authentication and technical detection systems. However, it remains important to combine technology with clear processes. For example, if a bank transfer is suddenly requested during a videoconference, a second, independent communication channel should be used for verification.


© Fraunhofer SIT (sit.fraunhofer.de)

No comments:

Post a Comment

TECH When the fake boss calls: a real-time warning system for fake videoconferences Videoconferencing facilitates communication between diff...