Vox Trust
Don't detect fake voices. Prove real ones.
An open protocol and Rust reference implementation that seals a human voice at the source, so anyone can verify it later.
The problem
Cloning a voice now takes seconds of audio. Our ears can no longer tell a real voice from a cloned one, and detectors chase a moving target: as detection improves, so does generation.
Attest the real, don't guess the fake
Instead of asking "is this voice fake?", Vox Trust asks "was a real one sealed?"
- Seal. The speaker's device adds a signed seal to the audio at the source.
- Carry. The seal can ride in a pluggable audio watermark, so it can survive re-encoding. A first, experimental carrier is built and measured in public: it survives MP3, AAC and Opus, but not phone-call codecs or noise.
- Verify. Anyone with the protocol checks which key sealed the audio, when, and which seconds were altered.
- Decide. A local trust policy turns the result into one of four verdicts.
Four verdicts
Verified
A valid seal from a key you trust.
Unsealed
No seal, from someone who never used the protocol. Neutral, not "fake".
Warning
No seal from a contact who usually seals. Be careful, and ask them to send it again.
Alert
A broken seal, a seal from a different key than the one you have for that contact, or none from a contact who always seals in strict mode.
Illustration of the policy verdicts. The live demo runs the real Rust core in your browser; nothing is uploaded.
How it compares
Others watermark the fake. Vox Trust seals the real.
Google, Meta, Microsoft and Resemble mark audio their own AI generates. Deepfake detectors estimate whether a voice is synthetic. Vox Trust lets a real speaker seal their own voice, and anyone can check it with ordinary cryptography.
| Vox Trust | Google SynthID | Meta AudioSeal | Deepfake detectors | C2PA | |
|---|---|---|---|---|---|
| Vouches for a real human recording | Yes | Nomarks AI output | Nomarks AI output | Estimates | If the app signs |
| Proof bound to the speaker's own key | Yes | No | No | No | Yes |
| Points to the seconds that were altered | Yes | No | Partly | No | Nowhole file |
| Anyone can verify, offline | Yesin the browser | NoGoogle's detector | Yes | No | Yes |
| Open specification and code | Yes | Nofor audio | Code | No | Specification |
| Survives MP3, AAC, Opus | Experimental100 % measured | Claimed | Measured | n/a | Nometadata stripped |
| Survives phone calls | Not yet11 % | Not published | No0 % measured | Yes | No |
| Cost to verify | Millisecondsno AI model | Cloud | Neural network | Cloud | Milliseconds |
"Measured" means run on our public benchmark with the same corpus and codecs; "claimed" is the vendor's published figure. Deepfake detectors: for example Pindrop Pulse. Every cell is sourced in the comparison notes.
Build it into your app
Vox Trust is an open protocol meant to live inside the apps people already use: messengers, voicemail, podcast and newsroom tools, support platforms. Seal on send, verify on receive. A 93 KB WebAssembly core with no dependencies, in the browser or in Node, or a Rust crate.
npm install vox-trust
cargo add vox-trust-core
import { load } from "vox-trust";
const vt = await load();
const sealed = vt.seal(wav, {
mode: "public", key: seed, createdUnix, chunkFrames: 16000,
});
const report = vt.verify(sealed, { pinnedPublicKey });
vt.decide(report.check, contact); // "verified" | "unsealed" | "warning" | "alert"
The integration guide takes about 10 minutes, with examples for Node, the browser and Rust. Apache-2.0, no service to call, no account.
Also: pip install vox-trust (Python), Android and iOS libraries attached to each release, a browser extension for Chrome, Firefox and Safari, and an MCP server so AI assistants can verify recordings.
What it is not
- It is not deepfake detection and not voice biometrics.
- It does not prove a speaker is human or is who they claim. It proves that a key sealed the audio. A stolen key makes valid seals.
- "No seal" does not mean "fake": compression and noise suppression can erase a watermark.
- It does not protect against someone who controls the sender's device.
The full list of attackers and limits is in the threat model (in English).
Where we are
- A specification (version 0.2: file mode a release candidate, in-band parts experimental) and a threat model, open for criticism.
- A Rust reference implementation with file mode (a signed per-chunk manifest inside a WAV file), trust policy, a command-line tool and a WebAssembly build. Test vectors are published and cross-checked by a second, Python implementation written by the same author.
- A live demo that really verifies. Sealed files survive only bit-exact copies; verdicts come from file mode only.
- Not audited. No independent review yet. Do not use it to protect anyone.
- Measured: the experimental watermark (stdm-2) keeps the seal through MP3, AAC, Opus 24 and 32 kbit/s and G.722 (100 %), MP3 re-shared as Opus (98 %) and a 1 % speed change (84 %), with no false alarms, but still fails on phone calls (AMR-WB), strong noise and echo. A seal inside audio can be copied into other audio, so it gives no verdicts yet. Full results.
See the roadmap (in English).
Built on, and next to, existing work
Vox Trust is meant to complement these rather than replace them:
- C2PA content credentials, which already use watermarks as "soft bindings".
- Audio watermarks like AudioSeal and WavMark, which can serve as carriers.
- Research on public-key speech provenance, e.g. MerkleSpeech.
Help us break it
The most useful thing right now is scrutiny: read the threat model, find the attacker we missed, or reproduce a measurement. Open an issue on GitHub. Report vulnerabilities privately, as described in the security policy.