Diopter
Sign in Try the Detector

While Resemble AI classifies the media, Diopter AI scores the live conversation.

Our detectors use a two-axis matrix for executive impersonation, MFA reset attacks, candidate fraud on live calls, and wire fraud. An upcoming update to our verdict matrix is likely to strengthen our detection.

Key Takeaways

  • The two-axis framework: Media detection + Conversation arc risk score
  • Bot-free capture on live video platforms like Zoom, Teams, Google Meet, Webex, and VoIP
  • Identity and payment verification tied to the call, not just the media
  • Policy alignment layer that maps live calls to existing controls
  • On-prem and hybrid deployment with configurable retention
  • MDM rollout via MS Intune and Jamf
  • Detection-only posture with no voice-generation product line

Where Resemble AI and Diopter AI Diverge

During procurement, enterprises usually evaluate deepfake detection vendors based on the accuracy of their classifiers, their generator coverage, benchmark rankings, and latency. Resemble AI scores strongly on all four aspects. However, their numbers describe the media verdict, and that verdict fires only when synthetic media is present. It cannot speak to what happens when an MFA is reset, or just before a wire clears, on a call where the media is clean.

Diopter AI puts a second axis to the media verdict graph: a conversation arc score that reads the urgency, authority, isolation, escalation, and the shape of the entire conversation on a live call. The two axes come together to create four verdict classes.

This divergence is what this article compares.

The Two-Axis Verdict Matrix

In February 2024, Arup’s finance worker joined a video call, saw the CFO and several upper-level executives on screen, and approved fifteen transfers worth roughly $25.6 million. Unknown to him, every other participant he saw onscreen was synthetic. While the delivery of the attack was a convincingly deepfaked video call, the main attack was carried by a conversational arc, that is, the sequence of authority signals, urgency framing, and channel narrowing that manipulated the trust of a competent finance worker into signing off transfers he would not have approved in any other context. This sequence of signals is what our detectors score, and Resemble AI does not.

Our core construct is a two-by-two verdict matrix that runs across every protected call. One axis is the media dimension most detection vendors cover: AI Detection. The other axis is the Conversation Arc that scores the behavioral dimension that Resemble AI does not.

Our detection is likely to be further enhanced and strengthened with our upcoming next matrix update.

This matrix produces four verdict classes:

AI Detection Arc Risk What to Look For On a Call Verdict Action
Clean Low Real human, normal request flow. No synthetic media detected, no pressure pattern. Verified Allow
Clean High Real human voice, but the conversation is being shaped towards a certain ask. Urgency, authority framing, and off-policy timing. A social engineering attempt by a real person or insider. High-Risk Threat Hold and flag for review
Synthetic Low Synthetic voice / deepfake video detected. Conversation is not pushing towards high-risk action. Chance of it being a benign AI agent, a voice filter, or a probe. Suspected Threat Hold until Verified
Synthetic High Synthetic media + conversation arc closing on a wire/MFA reset/credential handoff/new hire. Both axes confirm. High-Risk Threat Block and flag for security review

A single-axis detector is unable to flag Clean/High because no synthetic signals fire, but Diopter can. We included this category because an insider case is the hardest to prosecute since the identity is always legitimate, but it is a fraud that your security team already spends the most time manually reviewing.

Approved AI notetakers, translation bots, and voice filters that run inside every modern enterprise also produce synthetic signals that may trigger a false alarm. A single-axis detector flags all of them, and thereby drowns your SOC in low-context alerts, before eventually quietly turning them off. The Synthetic-Low class exists so that the notetaker does not become a blocker.

A block on a live call proves expensive, generates a ticket, creates a reason for escalation, and disrupts trust. Our detector, therefore, reserves the block for a Synthetic-High category, where the arc has been observed, and the media confirms the manipulation. This verdict will be defensible in a post-incident review.

Most buyers have been evaluating deepfake detection vendors on the assumption that better classifiers stop complex attacks. While metrics like accuracy percentages, benchmark rankings, latency numbers, and model coverage lists do matter, they are not what determines whether the wire clears or not in the end.

Features Comparison at a Glance

Capabilities Resemble AI Diopter AI
Conversation Arc Scoring Not offered Yes, two-axis verdict with four verdict classes
Bot-free Live Capture Not offered Yes. Across Webex, Teams, Google Meet, VoIP, Zoom
Bot-enabled Live Capture Yes Yes
Payment and Identity Verification Tied to a Call Yes. Identity and claim verification. Yes. Payment instruction on a live call.
Policy Alignment Not publicly documented1 Yes. Dedicated capability layer.
Published Benchmark Ranking Yes Not published
Detection Architecture DETECT-World Multi-layered: media forensics, liveness, generator fingerprinting
Voice generation product line Yes. Full voice AI product line. Not offered
On-prem and Hybrid Yes Yes
Retention Zero Retention Mode (ZDR) available. Configurable, including ZDR
MDM Rollout via Intune and Jamf Not publicly listed1 Yes
SOC 2 Type II Yes. Enterprise only. In progress
Conference Platform Coverage Zoom, Teams, Meet, Webex Zoom, Teams, Meet, Webex, VoIP
Content Provenance Watermarking PerTh multimodal, C2PA support Not offered
Explainability of Frame-level Verdict Yes. Forensic explanations, exportable audit trails, speaker profiling. Yes. Multi-layer votes + arc reasoning

1 Rows marked “Not offered”, “Not publicly documented” or “Not publicly listed” reflect the information published on Resemble AI’s website and public product documentation as of 7 September 2026. They reflect what each vendor publishes on that date and are not a statement of confirmed product limits.

Where Resemble AI is the Better Fit

Resemble AI’s product surface is a closer match if these three points resonate with your obligations:

  • Content Provenance and EU AI Act Ready: The Article 50 of the EU AI Act came into force on August 2, 2026. It requires machine-readable provenance on AI-generated content. Resemble AI’s PerTh Multimodal watermarking is a direct instrument meant for this specific obligation. If your company is positioned to be transparent on content your organisation produces as per the EU AI Act, then Resemble AI is the asset that will suit your requirements.
  • Published Third-Party Benchmarks and Forensic Explainability: Enterprises looking to weigh published benchmark rankings, real-time forensic explanations, reverse image search, fraud classification with attack vector, speaker profiling, and exportable audit trails, will find Resemble Intelligence closer to their criteria.
  • Broad Multimodal Media Forensics: If your organization’s workload requires asynchronous reviews of uploaded documents that may include images, audio, and video assets, Resemble AI’s DETECT-World architecture and its 250+ generator coverage claim will be your best fit.

Where Diopter AI is a Better Fit

  • Stops Live-call Payment Fraud: Companies that conduct high-value calls that include wire transfers, merger and acquisition negotiations, and vendor payment discussions, the fraud verdict must be able to change the control before the wire is executed. Our conversation arc scoring detects coached humans even on calls where no synthetic media have been explicitly detected. If a verdict returns where both the media evidence and the arc detect fraud, a recommendation to block the wire is routed to your workflow.
  • Bot-Free Fraud Detection on Sensitive Calls: Bots may not always be allowed on call during extremely sensitive board meetings, M&A negotiations, and vendor discussions. Resemble Meetings is integrated with the calendar, joins calls as a named participant, and can be removed by the host at the time of highest exposure risk. Instead, our detectors run in the background on the device being used for the call, much like antivirus software, and are able to capture audio and video directly from behind the camera and microphone, with no visible participant or calendar invites.
  • Help-desk and Identity Workflow Defense: In circumstances such as new-hire onboarding calls, MFA reset requests, or credential handoffs, a genuine human caller may be secretly coached by an attacker who tries to manipulate the call outcome. Our detector scores a Clean/High verdict and recommends a video callback before the requested action is approved.
  • Detection-only Posture: We offer a detection-only product. If a vendor also produces a media-generation product, while it may raise legitimate governance, procurement, and conflict-of-interest review concerns, they do not necessarily disqualify their detectors.

Four Scenarios Where the Arc Matters Most

These scenarios show how our detectors work in different use case settings and how our two-axis conversation arc goes further than a single-axis classifier:

  1. Financial Wire Fraud: An attacker cold-calls a treasury executive posing as a known vendor. The voice is human but coached. As the call escalates from just a routine banking-detail update to a same-day wire, it switches to a demand for secrecy. A frame classifier stays quiet through each frame because nothing on the call is identified as synthetic.

    Our detector’s arc closes on Clean/High, flags for review, and recommends holding the wire until an out-of-band callback confirms it.

  2. Executive Impersonation: A synthetic version of a top-level executive joins a scheduled cross-team call to authorize an unbudgeted acquisition transfer. An attacker can generate the video using HeyGen, while developing a realistic voice model using ElevenLabs.

    Our video and audio detectors fire on Synthetic. The arc registers authority framing as well as an off-policy urgency. The verdict lands on Synthetic/High, and the recommended action is to block, routed to the approver before the transfer is executed. This helps avoid a repeat of incidents such as the Arup deepfake attack.

  3. Candidate Fraud: A remote software engineer candidate joins a final-round interview using an identity that runs across three previous companies, generated with a face-swap and a live voice filter.

    Our video detector fires on Synthetic. The arc registers isolation signals, including the refusal to enable a second camera as well as vague answers on prior employment specifics. The verdict lands on Synthetic/High before the offer letter is shared.

  4. Help-desk Defense: Incidents like a live caller impersonating an executive, requesting an MFA reset, where the voice is real but the caller is not.

    Our detector notices the arc and registers the urgency, authority, and channel-narrowing. The verdict lands on Clean/High, flags the ticket for review, and recommends a video callback before the credential is reset. This is a case that a media classifier alone cannot catch owing to the lack of a synthetic signal.

Four Questions to Ask Deepfake Detection Vendors

Before signing a contract, it’s best for enterprise evaluation committees to run through these four questions with your choice of vendors. These questions will help you understand the limits of your vendor’s products.

  1. Can you show me the verdict output on a real human running a coached social engineering attack?

    If vendors say their detectors only fire on synthetic media, they offer partial control. The most expensive fraud in your organization will come from a real voice on a real call.

  2. What is your deployment model on an executive board call where a visible bot is not acceptable?

    If the answer requires renaming a bot to a benign label, you are being handed a workaround that will not survive a legal review. Resemble AI’s own product documentation states that its detection bot can be “named and positioned as a note taker”.

  3. Which independent NVLAP-accredited lab has tested your PAD to ISO/IEC 30107-3, and at what level?

    If the vendor is positioning against biometric identity flows, the answer needs to be a confirmation letter, not a benchmark leaderboard position.

  4. Does your company sell a voice-generation product alongside your detection product?

    While the answer shouldn’t be a disqualifier, it must be a valid data point that is included in your governance review policy.

Choosing What Suits Your Stack Best

At the end of the day, choosing the correct vendor isn’t about a benchmark or latency figure, coverage list, or scoreboard, but a product that successfully stops the wire.

The most expensive fraud your company may suffer will not be caught by a deepfake detector alone. It will be caught by a tool that is watching the shape of the conversation. The wire stops when someone in the loop sees the ask being shaped before it lands. When a real or synthetic voice requests something while emphasizing urgency on a particular channel, at a specific time, and in a premeditated sequence, the detector should surface that verdict while a control can still act on it, and prove the manipulation in a post-incident review.

Our multi-layer detector is built to withstand the most sophisticated deepfake attacks. The video, audio, payment, identity, and policy layers of our detector keep the verdict inside the few crucial seconds where a control can still act. This ability to score the conversation arc, therefore, becomes the determining factor that decides what happens between the signal firing and the wire clearing.

Book a walkthrough of an attack arc with Diopter AI.It takes 30 minutes and is NDA-safe. We replay a real deepfake incident, show the signals we would have scored, and map the verdict your team could have acted on.

Book a walkthrough →

Disclaimer: The comparison shared in this article is based on publicly available information as of the article’s publication date. Resemble AI® is a registered trademark of Resemble AI, used here solely for the purpose of identification and lawful comparative reference. Diopter AI is not affiliated with, endorsed by, or sponsored by Resemble AI.


FAQs

How is Diopter AI different from Resemble AI?

Resemble AI classifies media on and off a call, while Diopter AI classifies the media and scores the shape of the live conversation, so that even a coached human trying to run a social-engineering script offers a verdict that can prevent fraud before the wire is approved.

Do Diopter AI and Resemble AI detect the same media types?

Yes. Both Resemble AI and Diopter AI detect audio, images, and video deepfakes. The only difference lies in the live-call layer and the identity verification action that the call closes on.

How does Diopter AI stop deepfake wire fraud compared to Resemble AI?

Diopter AI runs a two-axis verdict on every protected call: a media detection axis and a conversational arc score that reads authority, urgency, isolation, the shape of the ask, and the presence of any coached human requests that may otherwise pass being flagged as synthetic. Capture is bot-free, so detection stays alive during sensitive calls. Resemble AI covers the media axis through its DETECT-World and Resemble Intelligence, while Resemble Meetings joins live calls with a named participant via calendar sync.

What is an injection attack, and does liveness detection catch it?

An injection attack bypasses the camera, and instead of showing a photo or mask to the lens, allows an attacker to pipe synthetic video directly into the application through a virtual camera or manipulated SDK call. Liveness detection, on the other hand, is designed to catch attacks at the lens but cannot see injections further downstream. Diopter runs both across each frame and scores the conversational arc in real time.

Is a single detection model enough, or does the stack require multiple layers?

Both approaches are in the market. Resemble AI’s DETECT-World is a world-architecture model aimed at generalizing across unseen generators through physics-based consistency checks. Diopter AI runs a layered stack that works independently of each other to keep detection active even when one layer lags. The practical buyer questions are retraining velocity and failure behaviour: how fast the detector adapts to a new generator, and what the verdict does when one signal disagrees with the others. A head-to-head analysis on the buyer’s own traffic against the specific generators worrying the security team, will separate them better than parameter counts or architecture labels.

When should I pick Diopter AI over Resemble AI?

Diopter AI’s detector proves to be the best option when your exposure window can exist within a live conversation that is closing a payment, MFA reset, or identity handoff. Diopter also enables bot-free high-exposure calls and is able to score coached-human social-engineering attacks that may carry no synthetic signals thanks to its conversation arc model.

Does Diopter AI offer on-prem deployment?

Diopter offers on-prem and hybrid deployment with Zero Data Retention configurable where regulatory environments demand it. Resemble AI also offers on-prem deployment.

Get the Diopter threat brief

Monthly analysis of AI social engineering, voice fraud and deepfake attacks on enterprises.

One email a month. No spam, and we never share your address.

Ask AI about this articleClaudeChatGPTPerplexity
Cite this articleAPA · MLA · BibTeX
APA 7
Gupta, S. (2026, September 7). Diopter AI vs Resemble AI: Scoring the Arc, Not Just the Frame. Diopter AI. https://diopter.ai/blog/diopter-vs-resemble-ai/
MLA 9
Gupta, Surojoy. "Diopter AI vs Resemble AI: Scoring the Arc, Not Just the Frame." Diopter AI, 7 September 2026, https://diopter.ai/blog/diopter-vs-resemble-ai/.
BibTeX
@misc{diopter20267cad68, author = {Surojoy Gupta}, title = {Diopter AI vs Resemble AI: Scoring the Arc, Not Just the Frame}, year = {2026}, month = {sep}, howpublished = {Diopter AI}, url = {https://diopter.ai/blog/diopter-vs-resemble-ai/} }
SG
Security Researcher & Writer

Surojoy Gupta is a security researcher and writer with 8 years embedded in the cybersecurity industry, specializing in deepfake fraud, social engineering, and AI-driven threats. His work covers APT threat analysis, ransomware, and the evolving tactics attackers use to exploit enterprise trust at the human layer.