
Freely available software that can mimic a specific individualâs voice produces results that can fool people and voice-activated tools such as smart home assistants.
Security researchers are increasingly concerned by deepfake software, which uses artificial intelligence to alter videos or photographs, for example by mapping one personâs face onto another.
at the University of Chicago and her colleagues wanted to investigate audio versions of these tools, which generate realistic English speech based on a sample of a personâs voice, after reading about such technology being used to in 2019.
Advertisement
Voice commands are now used to control digital home assistants like Amazonâs Alexa, as well as some automated phone systems run by businesses such as banks. âWe wanted to look at how practical can these attacks be, given that weâve seen some evidence of them in the real world,â says Wenger.
She and her colleagues used two deepfake voice synthesis systems, downloaded from the popular GitHub code repository, to mimic voices. One system, AutoVC, requires up to 5 minutes of speech to generate a passable imitation of the target voice, but the other, SV2TTS, only requires 5 seconds. âWe wanted to target the low-bar attacker mindset,â says Wenger.
They used the software to try and unlock speaker recognition security systems used by Microsoft Azure, WeChat and Amazonâs Alexa system. Microsoft Azureâs voice recognition system is certified by several formal industry bodies, WeChat allows users to log in with their voice and Alexa enables people to use their voice to make payments in third-party apps like Uber.
AutoVC was able to fool Microsoft Azure around 15 per cent of the time, while SV2TTS managed 30 per cent. However, Azure requires users to speak trigger phrases to authenticate themselves, and the team found that SV2TTS could successfully spoof at least one of 10 of these common phrases for 62.5 per cent of the people the researchers tried, suggesting a persistent attacker would have a higher chance of breaking through.
Given its lower performance, the team didnât try AutoVC against WeChat and Amazon Alexa, but SV2TTS was able to successfully fool both systems around 63 per cent of the time.
Results varied, but deepfakes were more successful at spoofing womenâs voices and those of non-native English speakers. âWhy that happened, we need to investigate further,â says Wenger.
Microsoft declined to comment, while WeChat didnât respond to 51¶ŻÂțâs request to comment. An Amazon spokesperson said: âAlexa is built with multiple layers of privacy and security designed to keep customer information safe.â
The deepfake voices werenât only successful against computer systems. In a separate experiment, the team asked 200 people to identify whether voices were fake or real, with the fakes fooling them around half the time.
âWeâve already seen synthetic voice deployed âin the wildâ to compromise both humans and biometric targets, with this research reinforcing the technologyâs viability even at this early stage in its development,â says at Metaphysic, a company developing tools for deepfakes.
âThe realism and accessibility of voice synthesis is only going to improve, bringing with it profound implications for the cybersecurity landscape as our voices increasingly become biometric keys to our digital lives,â says Ajder.
Reference: