Vox Trust › Comparison
Vox Trust vs Google SynthID, Meta AudioSeal and deepfake detectors
Google SynthID, Meta AudioSeal, Microsoft and Resemble watermark audio their own AI generates, so it can be recognised later. Deepfake detectors estimate whether a voice is synthetic. Vox Trust does the opposite: a real speaker seals their own voice, and anyone verifies it with ordinary cryptography.
Features side by side
| Vox Trust | Google SynthID | Meta AudioSeal | Deepfake detectors | C2PA | |
|---|---|---|---|---|---|
| Vouches for a real human recording | Yes | Nomarks AI output | Nomarks AI output | Estimates | If the app signs |
| Proof bound to the speaker's own key | Yes | No | No | No | Yes |
| Points to the seconds that were altered | Yes | No | Partly | No | Nowhole file |
| Anyone can verify, offline | Yesin the browser | NoGoogle's detector | Yes | No | Yes |
| Open specification and code | Yes | Nofor audio | Code | No | Specification |
| Survives MP3, AAC, Opus | Experimental100 % measured | Claimed | Measured | n/a | Nometadata stripped |
| Survives phone calls | Not yet11 % | Not published | No0 % measured | Yes | No |
| Cost to verify | Millisecondsno AI model | Cloud | Neural network | Cloud | Milliseconds |
"Measured" means run on our public benchmark with the same corpus and codecs; "claimed" is the vendor's published figure. Every cell is sourced in the comparison notes.
Watermarks measured on the same test
13 recordings in 10 languages, the same real codecs and quality metrics. AudioSeal and WavMark carried one 16-bit message repeated over the whole recording, an easier task than a 102-bit seal every 9.6 s, so their numbers are an upper bound.
| Condition | Vox Trust stdm-2 | Meta AudioSeal | WavMark |
|---|---|---|---|
| MP3, AAC, Opus 24 kbit/s | 100 % | 100 % | 100 % (Opus) |
| G.722 (VoIP) | 100 % | 0 % | not run |
| Start of the audio trimmed | 100 % | 8 % | not run |
| 1 % speed change | 84 % | 62 % | 100 % |
| Noise at 30 dB | 51 % | 100 % | not run |
| Echo | 0 % | 100 % | 100 % |
| Noise reduction | 7 % | 0 % | 100 % |
| Phone call (AMR-WB 12.65 kbit/s) | 11 % | 0 % | 31 % |
| False alarms on unmarked audio | 0 | 0 | 0 |
| Cost to verify | milliseconds | neural network | minutes on CPU |
Full data and how to reproduce: benchmark results. Google SynthID could not be measured: its audio detector is not public.
Frequently asked questions
Is Vox Trust an alternative to Google SynthID?
It solves a different problem. SynthID marks audio made by Google's AI so it can be recognised as synthetic. Vox Trust lets a real person seal their own recordings, so a receiver can check they are genuine and unaltered. The two can coexist.
Does Vox Trust detect deepfakes?
No. It does not guess whether a voice is fake. It proves that a recording was sealed by a known key and shows which seconds changed. Unsealed audio is neutral, not "fake".
How is it different from AudioSeal or WavMark?
Those are watermarks: they hide a short message in audio. Vox Trust is a protocol on top: a cryptographic seal, a trust policy and exact tamper localisation. A watermark like AudioSeal or WavMark could carry the Vox Trust seal; we measured both as candidates.
Is Vox Trust open source?
Yes: the specification, the Rust implementation, the test vectors and the benchmark, including the conditions where it still fails.