Skip to main content
Capability

Voice deepfake detection: Catch a cloned voice while the call is still live

Diopter detects cloned and synthetic voices on conference and VoIP calls as the conversation happens, whether the voice belongs to an executive, vendor, or candidate.

A familiar voice can make a new payment request, a change in banking details, or hiring decisions feel routine. Voice deepfake detection checks the audio when the call is still live, so your team can decide before acting on what they hear.

30 minutes with a founder. We will sign your NDA first if you want one.
442%
rise in voice phishing attacks
Source · CrowdStrike
$35M
moved on a single cloned-voice call
Source · Reported, 2021
+220%
year-over-year rise in hiring fraud
Source · Industry reporting
The verdict

From signals to one action your team can take.

What drove this verdict
  • Audio
    Synthesis artifacts consistent with a voice-cloning or text-to-speech model
    Cloning detected
  • Coverage
    Enough clean speech to evaluate end to end, not a partial pass
    Full call scored
  • Channel
    No video on this call, so nothing corroborates the audio
    Voice only
  • Conversation
    Pressure and an escalating ask alongside the audio flag
    Urgency rising

Hold the line. A synthetic-voice verdict reaches your team while the call is live, with a recommended next step.

How Diopter decides a voice is synthetic

Diopter scores every participant on a call continuously for traces that voice-cloning and text-to-speech models have left behind. The audio is split into short consecutive segments, and each segment is scored for any cloned or manipulated media. The scoring continues throughout the call rather than stopping just after an opening sample.

That matters because synthetic speech may not appear from the beginning of the call.

Models such as ElevenLabs, OpenAI TTS, AWS Polly, PlayHT, and Respeecher can generate or manipulate synthetic speech and Diopter looks for traces these models may have left in the audio signal.

Diopter's enterprise deepfake voice detection does not need a recording of the real person to make the assessment. It does not create a voiceprint, store a voice profile, or require speaker enrollment. Rather, Diopter runs on Zoom, Microsoft Teams, Google Meet and Webex, as well as VoIP and conference phone calls.

The result is one of four bands: AI, mixed, clean or inconclusive, each with a confidence level. If the call could not be fully checked because the audio was too short, noisy, or degraded, the result is inconclusive rather than clean.

This synthetic voice detection is based on the audio and not who the speaker is.

How it works

How voice scoring runs

The same pipeline runs on a live call and on an audio file uploaded to the detector. It analyses the audio signal, not the speaker.

  1. 1

    Segment the audio

    Speech is split into short consecutive segments so the call can be scored as it happens rather than judged once at the end.

    Continuous · live or uploaded file

  2. 2

    Score each segment

    Every segment is scored independently for the artifacts left by voice-cloning and text-to-speech models. No reference recording of the real person is needed.

    No enrollment · no voice profile stored

  3. 3

    Keep scoring to the end

    Scoring runs for the length of the call, not just an opening sample, so a handoff to a different speaker or a voice that only appears at the ask is still scored.

    Real time · full call duration

  4. 4

    Resolve with coverage

    Segments combine into one band, and how much usable speech was actually evaluated is reported alongside it rather than assumed.

    Bands: AI · mixed · clean · inconclusive

The risk

Three places voice cloning shows up

For businesses, voice cloning can show up in three common situations: executive calls, vendor and candidate calls, callbacks that confirm nothing.

  • 01

    Cloned executives by phone

    Take this situation. Your treasury team receives a call with the familiar voice of the CFO asking for a wire to be transferred immediately. Eleven seconds of public audio can be enough to clone a voice convincingly. The attacker does not need to sound strange. The point is that the request and the audio should be familiar enough for nobody to question it. Diopter's deepfake voice detector can add another layer of checks here.

  • 02

    Spoofed vendor and candidate calls

    A supplier calls to confirm new banking details. The voice sounds like the person your accounts payable team normally deals with. Or, a candidate joins a remote screening call, and the voice on the other end has been generated or altered. These are two different situations. The voice is used to convince the receiver that the person or the request is genuine. Voice cloning detection can step in here as an additional layer of protection before a high-risk situation is encountered.

  • 03

    Callbacks that confirm nothing

    The callback can be a way to catch a fraudulent change. But it only works if the voice that answers is real. For example, for a finance team, that could mean a callback confirming a new beneficiary still confirms the wrong beneficiary.

The attack playbook

How a voice-cloning attack unfolds

These attacks move through a recognizable sequence. Diopter scores that sequence while the call is still in progress.

01
Authority

A familiar voice calls

An executive, a vendor contact, or a candidate, recognizable enough that the request feels routine.

02
Urgency

Urgency arrives early

A closing window, an overdue invoice, or a competing offer compresses the time to verify.

03
Isolation

The call moves off-channel

The conversation shifts to a private line or a follow-up that keeps others out of it.

04
Escalation

The asks escalate

A small confirmation becomes a larger request as the call builds on each prior yes.

05
The ask

The action is taken

A wire, a banking change, or an offer is acted on while the voice is still trusted.

See this run against your own approval flow.
30 minutes with a founder. We will replay a real incident end to end.
Book a walkthrough
Where it shows up

Which calls need voice deepfake detection

Voice deepfake detection is relevant to four types of calls where a trusted voice can influence a financial, access, hiring, or support decision. The teams acting on the verdict are usually treasury and accounts payable, recruiting, the IT help desk, and fraud operations or the CISO's team reviewing what was flagged.

Inbound phone calls

An executive, supplier, customer, or external contact calls, and the voice becomes the main signal your team has to judge who is on the other end. This includes automated application-to-person (A2P) calling, where a synthetic voice reaches many people rather than one target at a time. The scoring is the same either way, because it reads the audio rather than the campaign behind it.

Treasury authorization

A caller asks for a wire transfer, confirms a beneficiary change, or gives verbal approval for a payment. A routine authorization can also lead to wrong transactions.

Candidate calls

A candidate may join a remote screening call, and your HR personnel may not recognize that a cloned or altered voice is doing the talking.

Call center and support team calls

A caller contacts a support team about an account, access request, or other sensitive action. Diopter can detect a synthetic voice in support calls while the conversation is still live, so the agent has a verdict before granting access or resetting credentials.

Why Diopter

How voice deepfake detection differs from voice recognition

An audio deepfake detector determines whether audio is synthetic, and voice recognition determines who is speaking. Voice recognition typically relies on a sample, voiceprint, or stored media to match a speaker with a known person. On the other hand, a voice clone detection system examines the audio to check whether the speech is manipulated or AI-generated. Diopter does not identify the speaker; rather, it scores the audio signal for artifacts associated with voice-cloning and text-to-speech models. It also does not build voiceprints, does not store voice profiles, and does not require a prior recording of the real person. The score comes from generation artifacts in the signal. Accents, where the speaker is from, and what language they speak are not inputs and are not treated as risk signals.

A cloned voice passes a single listen. The script it runs, urgency, isolation, and the ask, gives it away across the call.

Side by side

Where single-layer tools stop.

Here is how Diopter compares to traditional training, single-layer detection, and identity tools. It brings together all these layers to give your security team a unified, actionable verdict.

Detects synthetic voice on a live call

Awareness training
Not supported
Single-frame deepfake
Not supported
Identity / reputation
Partial
Live-call detection
Supported
Diopter Arc
Supported

Detects deepfake video frames

Awareness training
Not supported
Single-frame deepfake
Supported
Identity / reputation
Not supported
Live-call detection
Partial
Diopter Arc
Supported

Verifies caller identity (reputation/biometric)

Awareness training
Not supported
Single-frame deepfake
Not supported
Identity / reputation
Supported
Live-call detection
Partial
Diopter Arc
Supported

Models the conversation arc (pressure → ask)

Awareness training
Partial
Single-frame deepfake
Not supported
Identity / reputation
Not supported
Live-call detection
Not supported
Diopter Arc
Supported

Correlates identity, media, and conversation signals on live calls

Awareness training
Not supported
Single-frame deepfake
Not supported
Identity / reputation
Partial
Live-call detection
Not supported
Diopter Arc
Supported

Forensic evidence chain for incident review

Awareness training
Not supported
Single-frame deepfake
Partial
Identity / reputation
Partial
Live-call detection
Partial
Diopter Arc
Supported
Supported Partial Not supported
Honest limits

What voice detection does not claim

The second item is the one most often assumed about voice tools, and it is worth being explicit that we do not do it.

It is not speaker identification

Diopter answers whether the audio is synthetic, not who is speaking. It does not build voiceprints, does not store voice profiles, and needs no prior recording of the real person to work.

It does not judge accent or origin

The score comes from generation artifacts in the signal. How someone sounds, where they are from, and what language they speak are not inputs, and must never be treated as risk signals.

Short or degraded audio is inconclusive

A few seconds of speech, heavy background noise, or a badly compressed line may not carry enough signal to judge. That is reported as inconclusive rather than passed off as clean.

Deployment & trust

Light to deploy, clear about what runs where.

Pilot in days, roll wider through MDM, and keep sensitive call media inside your perimeter.

Deployment & trust
  • On-prem and hybrid deployments supported
  • No caller-side install
  • Bot or bot-free capture
  • Configurable retention, including ZDR
  • MDM rollout (Intune, Jamf)
  • SOC 2 Type II in progress
Walkthrough · 30 min

Walk an attack arc with Diopter.

We will replay a real incident, show the signals Diopter scored, and map the verdict your team would act on. We will sign your NDA first if you want one.

Common questions

What security and fraud teams ask first.

SOC 2 Type II in progress.

There is no single accuracy number that holds across products, because most published scores are measured on clean uploaded audio files rather than on a compressed live call. When comparing tools, the questions that separate them are: does it score the whole call or only an opening sample; does it need speaker enrollment or a prior recording of the real person; does it report coverage when audio was too short or too noisy to check; and does it return a confidence level rather than a binary verdict. Diopter scores continuously for the full call duration, needs no enrollment, and returns one of four bands (AI, mixed, clean, inconclusive) with a confidence level, so an unscoreable call reads as inconclusive rather than clean.

Diopter is deployed as a detection layer alongside the call rather than as a self-serve API. It runs as an endpoint app or inline at the trunk, scores each participant continuously, and routes verdicts into your console, your SIEM, or a ticket by webhook. If your team needs programmatic integration, that is scoped during the walkthrough.

The United States regulates deepfakes through a patchwork of state-level laws and evolving federal laws. The TAKE IT DOWN Act criminalizes the publication of non-consensual intimate digital forgeries and requires platforms to remove them within 48 hours. There are also a few other law proposals, like the DEEPFAKES Accountability Act, the DEFIANCE Act, and the AI Labeling Act, that aim to target different vulnerabilities in the deepfake ecosystem. For organizations using audio deepfake detectors, some useful resources include the CISA, NSA, and FBI Guidance, the FinCEN Fraud Alerts, and the FTC Voice Cloning Guidelines.

No. Diopter scores the manipulation pattern, not isolated artifacts, so a normal call with real urgency does not trip it. Only the combination, an authority claim plus pressure plus an escalating ask, crosses the threshold. Your team sees fewer alerts with higher signal.

Diopter supports on-prem and hybrid deployments, with configurable retention including a zero-data-retention option. It runs with a meeting bot or bot-free, and needs no caller-side install. SOC 2 Type II is in progress.

Diopter works alongside the video and voice tools your team already uses, and rolls out through your existing MDM such as Intune or Jamf. There is no caller-side install and no change to how your team takes calls.

No. Every verdict carries a confidence level, not a flat flag, and if a call or file could not be fully checked, Diopter reports that rather than defaulting to a clean result. You can see this directly: every report from the Deepfake Detector shows the same confidence scoring behind Diopter's verdicts.

Research behind voice deepfake detection

All research →