Vox Trust
v0.6.0 · file mode works · not audited

Vox Trust

Don't detect fake voices. Prove real ones.

An open protocol and Rust reference implementation that seals a human voice at the source, so anyone can verify it later.

The problem

Cloning a voice now takes seconds of audio. Our ears can no longer tell a real voice from a cloned one, and detectors chase a moving target: as detection improves, so does generation.

Attest the real, don't guess the fake

Instead of asking "is this voice fake?", Vox Trust asks "was a real one sealed?"

  1. Seal. The speaker's device adds a signed seal to the audio at the source.
  2. Carry. The seal can ride in a pluggable audio watermark, so it can survive re-encoding. A first, experimental carrier is built and measured in public: it survives MP3, AAC and Opus, but not phone-call codecs or noise.
  3. Verify. Anyone with the protocol checks which key sealed the audio, when, and which seconds were altered.
  4. Decide. A local trust policy turns the result into one of four verdicts.

Four verdicts

Verified

A valid seal from a key you trust.

Unsealed

No seal, from someone who never used the protocol. Neutral, not "fake".

Warning

No seal from a contact who usually seals. Be careful, and ask them to send it again.

Alert

A broken seal, a seal from a different key than the one you have for that contact, or none from a contact who always seals in strict mode.

Illustration of the policy verdicts. The live demo runs the real Rust core in your browser; nothing is uploaded.

How it compares

Others watermark the fake. Vox Trust seals the real.

Google, Meta, Microsoft and Resemble mark audio their own AI generates. Deepfake detectors estimate whether a voice is synthetic. Vox Trust lets a real speaker seal their own voice, and anyone can check it with ordinary cryptography.

Vox TrustGoogle SynthIDMeta AudioSealDeepfake detectorsC2PA
Vouches for a real human recordingYesNomarks AI outputNomarks AI outputEstimatesIf the app signs
Proof bound to the speaker's own keyYesNoNoNoYes
Points to the seconds that were alteredYesNoPartlyNoNowhole file
Anyone can verify, offlineYesin the browserNoGoogle's detectorYesNoYes
Open specification and codeYesNofor audioCodeNoSpecification
Survives MP3, AAC, OpusExperimental100 % measuredClaimedMeasuredn/aNometadata stripped
Survives phone callsNot yet11 %Not publishedNo0 % measuredYesNo
Cost to verifyMillisecondsno AI modelCloudNeural networkCloudMilliseconds

"Measured" means run on our public benchmark with the same corpus and codecs; "claimed" is the vendor's published figure. Deepfake detectors: for example Pindrop Pulse. Every cell is sourced in the comparison notes.

Full comparison, measured results and FAQ →

Build it into your app

Vox Trust is an open protocol meant to live inside the apps people already use: messengers, voicemail, podcast and newsroom tools, support platforms. Seal on send, verify on receive. A 93 KB WebAssembly core with no dependencies, in the browser or in Node, or a Rust crate.

npm install vox-trust
cargo add vox-trust-core

import { load } from "vox-trust";
const vt = await load();
const sealed = vt.seal(wav, {
  mode: "public", key: seed, createdUnix, chunkFrames: 16000,
});
const report = vt.verify(sealed, { pinnedPublicKey });
vt.decide(report.check, contact); // "verified" | "unsealed" | "warning" | "alert"

The integration guide takes about 10 minutes, with examples for Node, the browser and Rust. Apache-2.0, no service to call, no account.

Also: pip install vox-trust (Python), Android and iOS libraries attached to each release, a browser extension for Chrome, Firefox and Safari, and an MCP server so AI assistants can verify recordings.

What it is not

  • It is not deepfake detection and not voice biometrics.
  • It does not prove a speaker is human or is who they claim. It proves that a key sealed the audio. A stolen key makes valid seals.
  • "No seal" does not mean "fake": compression and noise suppression can erase a watermark.
  • It does not protect against someone who controls the sender's device.

The full list of attackers and limits is in the threat model (in English).

Where we are

  • A specification (version 0.2: file mode a release candidate, in-band parts experimental) and a threat model, open for criticism.
  • A Rust reference implementation with file mode (a signed per-chunk manifest inside a WAV file), trust policy, a command-line tool and a WebAssembly build. Test vectors are published and cross-checked by a second, Python implementation written by the same author.
  • A live demo that really verifies. Sealed files survive only bit-exact copies; verdicts come from file mode only.
  • Not audited. No independent review yet. Do not use it to protect anyone.
  • Measured: the experimental watermark (stdm-2) keeps the seal through MP3, AAC, Opus 24 and 32 kbit/s and G.722 (100 %), MP3 re-shared as Opus (98 %) and a 1 % speed change (84 %), with no false alarms, but still fails on phone calls (AMR-WB), strong noise and echo. A seal inside audio can be copied into other audio, so it gives no verdicts yet. Full results.

See the roadmap (in English).

Built on, and next to, existing work

Vox Trust is meant to complement these rather than replace them:

  • C2PA content credentials, which already use watermarks as "soft bindings".
  • Audio watermarks like AudioSeal and WavMark, which can serve as carriers.
  • Research on public-key speech provenance, e.g. MerkleSpeech.

Help us break it

The most useful thing right now is scrutiny: read the threat model, find the attacker we missed, or reproduce a measurement. Open an issue on GitHub. Report vulnerabilities privately, as described in the security policy.