Spokes.wiki Search About
Defined Term concept source ↗ source url updated Tue Jun 09 2026 00:00:00 GMT+0000 (Coordinated Universal Time)

Audio deepfake

An audio deepfake is AI-synthesized speech that “convincingly mimics specific individuals … phrases or sentences they have never spoken.” This is the safety/consent axis the synthesis flagged as a coverage gap — the dark side of the voice-cloning and text-to-speech capability the wiki otherwise celebrates. Source: Wikipedia.

How it’s made

Two routes: synthetic TTS (text-analysis → acoustic model → vocoder, à la wavenet) and imitation / voice conversion (alter a real signal’s style/prosody, often via GANs). I.e. the same text-to-speech/voice-cloning tech this wiki tracks, pointed at impersonation.

Harms — and why “rights outweigh quality” bites

Concrete harms: a 2019 CEO-voice scam (€220k); a McAfee survey where “one person in ten … targeted by an AI voice cloning scam, 77% … lost money”; the Feb 2024 fake-Biden robocalls; and consent violations (Jeff Geerling’s and Gayanne Potter’s voices cloned without permission). This is the lived form of the spoke’s thesis that rights/consent can outweigh raw quality.

Responses — provenance & detection

Detection classifies “Spoof” vs “Bonafide” (ASVspoof, DARPA SemaFor) — “limited non-English research” remains a gap. Provenance is the other lever: watermarking like SynthID (the same scheme gemini-live-3-5-translate stamps on output). Policy is moving: the FCC banned AI voices in robocalls (Feb 2024); China mandates deepfake labeling. The counterweight to unconstrained voice-cloning.

What the supply side says about it: nothing (added 2026-08-03)

The responses above — detection, watermarking, the FCC ban — all sit downstream of the models. Upstream, in the catalogues people actually use to choose a cloner, the subject is absent.

free-voice-clone-list is 58,997 characters covering 53 open audio models, 35 of its 36 TTS entries advertising voice-cloning from a three-to-ten-second sample. Searched for consent, ethics, watermark, misuse, deepfake, responsible, abuse and disclaimer, it returns zero matches. Not a thin treatment — none. The single licence in it carrying behavioural use restrictions (OpenRAIL-M, supertonic-2) belongs to the one model that cannot clone a voice.

This wiki had already flagged luxtts on 2026-07-29 for shipping three-second cloning with no consent language. One project is an oversight; a whole catalogue is the norm. The asymmetry is worth stating plainly: the harms on this page are documented in journalism, regulation and academic detection work, while the distribution layer that puts the capability in reach carries no corresponding line at all. Nothing here shows the omission causes harm — it is a gap in what gets said, recorded because this page exists to track exactly that gap.

One project did write it down (added 2026-08-06)

voicecraft breaks the pattern above. Its repository prohibits generating or editing anyone’s speech without consent and names political figures and celebrities as the case it cares about, and its licences (CC BY-NC-SA for code, Coqui CPML for weights) at least block commercial use. So the gap is a norm, not a law — a project can ship the sentence.

What makes it the more interesting example is which project it is. VoiceCraft’s headline capability is editing an existing recording, not synthesizing a new one: changing words inside real audio while the speaker and the room stay intact. That is the capability this page’s harms actually describe — a doctored recording of something that happened beats a fabricated clip of something that did not — and it is the one project here that paired it with a consent statement. The correlation this page half-expected, that the loudest capability comes with the quietest ethics, does not hold.

voice-cloning · text-to-speech · wavenet · gemini-live-3-5-translate · free-voice-clone-list · supertonic-2 · luxtts · voicecraft · speech-audio-ai