Skip to main content
Capability

Catch a cloned voice while the call is still live

Detect cloned and synthetic voices on conference and VoIP calls as the conversation happens, whether the voice belongs to an executive, a vendor, or a candidate.

30 minutes with a founder. We will sign your NDA first if you want one.
442%
rise in voice phishing attacks
Source · CrowdStrike
$35M
moved on a single cloned-voice call
Source · Reported, 2024
+220%
year-over-year rise in hiring fraud
Source · Industry reporting
The verdict

From signals to one action your team can take.

What drove this verdict
  • Audio
    Synthesis artifacts consistent with a voice-cloning or text-to-speech model
    Cloning detected
  • Coverage
    Enough clean speech to evaluate end to end, not a partial pass
    Full call scored
  • Channel
    No video on this call, so nothing corroborates the audio
    Voice only
  • Conversation
    Pressure and an escalating ask alongside the audio flag
    Urgency rising

Hold the line. A synthetic-voice verdict reaches your team while the call is live, with a recommended next step.

How it works

How voice scoring runs

The same pipeline runs on a live call and on an audio file uploaded to the detector. It analyses the audio signal, not the speaker.

  1. 1

    Segment the audio

    Speech is split into short consecutive segments so the call can be scored as it happens rather than judged once at the end.

    Continuous · live or uploaded file

  2. 2

    Score each segment

    Every segment is scored independently for the artifacts left by voice-cloning and text-to-speech models. No reference recording of the real person is needed.

    No enrollment · no voice profile stored

  3. 3

    Keep scoring to the end

    Scoring runs for the length of the call, not just an opening sample, so a handoff to a different speaker or a voice that only appears at the ask is still scored.

    Real time · full call duration

  4. 4

    Resolve with coverage

    Segments combine into one band, and how much usable speech was actually evaluated is reported alongside it rather than assumed.

    Bands: AI · mixed · clean · inconclusive

The risk

Where voice cloning shows up

  • 01

    Cloned executives by phone

    A few seconds of public audio is enough to clone a voice convincing enough to move a wire or force an exception.

  • 02

    Spoofed vendor and candidate calls

    A cloned vendor confirming new banking details, or an altered candidate voice on a screening call.

  • 03

    Callbacks that confirm nothing

    The callback is supposed to be the control that catches a fraudulent change. It only works if the voice that answers is real, and a clone answers just as convincingly as the person it copies.

The attack playbook

How a voice-cloning attack unfolds

These attacks move through a recognizable sequence. Diopter scores that sequence while the call is still in progress.

01
Authority

A familiar voice calls

An executive, a vendor contact, or a candidate, recognizable enough that the request feels routine.

02
Urgency

Urgency arrives early

A closing window, an overdue invoice, or a competing offer compresses the time to verify.

03
Isolation

The call moves off-channel

The conversation shifts to a private line or a follow-up that keeps others out of it.

04
Escalation

The asks escalate

A small confirmation becomes a larger request as the call builds on each prior yes.

05
The ask

The action is taken

A wire, a banking change, or an offer is acted on while the voice is still trusted.

See this run against your own approval flow.
30 minutes with a founder. We will replay a real incident end to end.
Book a walkthrough
Where it shows up

The calls where there is no camera to check.

Voice-only channels are where this capability matters most, because every other verification signal a person would use is absent.

Inbound phone calls

Voice-only calls with no video to fall back on, which is where cloning is cheapest for an attacker and hardest for a person to catch.

Treasury authorization

The verbal approval on a transfer, scored while it is being given rather than reconstructed afterwards.

Supplier confirmations

The callback that is supposed to confirm a banking change, which only helps if the voice answering it is real.

Remote screening calls

Early-round candidate calls where a cloned or altered voice is doing the talking.

Why Diopter

This is not voice recognition, and that is the point.

Almost every tool in this category works by matching a voice against an enrolled sample of the real person, which means it needs a recording of everyone it protects, a database of voiceprints to secure, and an enrollment process nobody completes. Diopter does none of that. It scores the audio itself for the artifacts that cloning and text-to-speech models leave behind, so it works on a caller you have never heard before, on the first call, with nothing stored about anyone's voice.

A cloned voice passes a single listen. The script it runs, urgency, isolation, and the ask, gives it away across the call.

Side by side

Where single-layer tools stop.

Each category below covers one part of the attack and is blind to the rest. The last column is the only one that correlates them into a single verdict.

Detects synthetic voice on a live call

Awareness training
Single-frame deepfake
Identity / reputation
Live-call detection
Diopter Arc

Detects deepfake video frames

Awareness training
Single-frame deepfake
Identity / reputation
Live-call detection
Diopter Arc

Verifies caller identity (reputation/biometric)

Awareness training
Single-frame deepfake
Identity / reputation
Live-call detection
Diopter Arc

Models the conversation arc (pressure → ask)

Awareness training
Single-frame deepfake
Identity / reputation
Live-call detection
Diopter Arc

Correlates identity, media, and conversation signals on live calls

Awareness training
Single-frame deepfake
Identity / reputation
Live-call detection
Diopter Arc

Forensic evidence chain for incident review

Awareness training
Single-frame deepfake
Identity / reputation
Live-call detection
Diopter Arc
Supported Partial Not supported
Honest limits

What voice detection does not claim

The second item is the one most often assumed about voice tools, and it is worth being explicit that we do not do it.

It is not speaker identification

Diopter answers whether the audio is synthetic, not who is speaking. It does not build voiceprints, does not store voice profiles, and needs no prior recording of the real person to work.

It does not judge accent or origin

The score comes from generation artifacts in the signal. How someone sounds, where they are from, and what language they speak are not inputs, and must never be treated as risk signals.

Short or degraded audio is inconclusive

A few seconds of speech, heavy background noise, or a badly compressed line may not carry enough signal to judge. That is reported as inconclusive rather than passed off as clean.

Deployment & trust

Light to deploy, clear about what runs where.

Pilot in days, roll wider through MDM, and keep sensitive call media inside your perimeter.

Deployment & trust
  • On-prem and hybrid deployments supported
  • No caller-side install
  • Bot or bot-free capture
  • Configurable retention, including ZDR
  • MDM rollout (Intune, Jamf)
  • SOC 2 Type II in progress
Walkthrough · 30 min

Walk an attack arc with Diopter.

We will replay a real incident, show the signals Diopter scored, and map the verdict your team would act on. We will sign your NDA first if you want one.

Common questions

What security and fraud teams ask first.