Trend

Deepfake vishing: detect and stop AI voice attacks on Swiss organisations

Deepfake vishing combines AI-generated voices with voice phishing. Attackers clone the voice of a CEO, CFO or IT admin from a few seconds of public audio. They then call employees to trigger payments, password resets or MFA approvals. For Swiss organisations the vector is dangerous because classical email filters, MFA and EDR do not see it. Effective defence requires process controls (call-back, second channel) and awareness with realistic voice samples. It also requires a SOC that correlates accompanying signals in M365 and Entra ID.

All

Scope: deepfake vishing vs BEC vs AI attacks

This page covers deepfake vishing as a concrete attack vector: cloned voices over phone or voice-over-IP, often combined with a preparatory email. The text-based path via M365 mailboxes sits on Business Email Compromise in M365. The broader trend perspective on how AI accelerates attacks overall is on AI-driven cyber attacks. For the defence side with AI in the SOC see AI in the SOC. Human preparation through training belongs to Security Awareness.

Short version

Deepfake vishing is the vocal variant of social engineering. It bypasses email filters entirely and targets the person on the phone, usually under time pressure and with authority framing.

How a typical attack unfolds

  1. Reconnaissance: attacker collects voice material of the target (interviews, podcasts, webinars, LinkedIn videos, voicemail greetings). 10 to 30 seconds are already enough for modern voice-cloning models.
  2. Context building: research on active projects, travel times and absences via LinkedIn, commercial register and company news.
  3. Preparation: optionally a preparatory email from a compromised or spoofed mailbox announcing the call.
  4. Call: cloned voice of the CEO or CFO calls finance, HR or IT. Typical requests: urgent payment to a new supplier, password reset, MFA approval, consent to an OAuth app.
  5. Follow-up: on success the trail is obscured across several accounts; for payments often via foreign accounts with a short recovery window.
Reality check

Voice cloning is achievable in minutes today with publicly available tools. Even close colleagues often cannot reliably distinguish the cloned voice on the phone. A poor line or emotional pressure makes this harder.

Why Swiss organisations are especially exposed

  • Executives in SMEs and enterprises are publicly present in media; voice material is freely available.
  • Payments to new suppliers are routine; a plausible CFO call under time pressure often goes unchallenged.
  • Cyber insurers scrutinise social-engineering cases. Without documented process controls (call-back, four-eyes), insurers may classify a payment as gross negligence. See SOC and cyber insurance.
  • revFADP (revised Federal Act on Data Protection) traceability also applies to processes where staff release personal data, such as HR or customer data, via a phone call. See revFADP and SOC.

Effective defence in three layers

LayerMeasureWhy it works
ProcessMandatory call-back: confirm payments and sensitive actions only through a call-back on a stored number or a second channel (Teams, Signal).Removes the attacker's only runway; works independently of voice quality.
ProcessFour-eyes principle for new payees and changes to account data; documented approval workflow in the ERP.Breaks the single-call path; creates an audit trail for insurance and revFADP (revised Federal Act on Data Protection).
HumanAwareness with realistic voice samples and live drills, not only email phishing. See Security Awareness.Only those who have heard the attack once in a calm setting recognise it under pressure.
Technology / SOCCorrelation of accompanying signals in M365 and Entra ID: unusual mail rules, consent grants, MFA bombing, atypical sign-ins shortly before or after the call.Deepfake vishing is rarely isolated; the preparation leaves traces in the identity layer, visible via ITDR.
Technology / SOCIncident playbook for vishing cases: immediate session invalidation, block new payment runs, contact bank, forensic capture of call metadata.Recovery window for wire transfers is hours, not days; playbook must be rehearsed in advance. See Incident response plan.
What deepfake-detection tools deliver today

Automatic detection of synthetic voices in real time is an active research area but not yet a reliable standard for enterprise telephony. Do not rely on a detector product alone; process controls and awareness remain the load-bearing pillars.

Frequently asked questions

How much audio do attackers need to clone a voice?

Current voice-cloning models produce usable results from 10 to 30 seconds of clean audio. For executives with public interviews, podcasts or conference appearances this bar effectively does not exist.

Is MFA enough against deepfake vishing?

MFA alone is insufficient against deepfake vishing. The attack often aims exactly at getting an employee to approve an MFA prompt or reset a password. Phishing-resistant MFA (FIDO2, passkeys) makes technical abuse harder. Process controls such as call-back and four-eyes remain necessary.

What is the difference between deepfake vishing and CEO fraud?

CEO fraud is the umbrella term for fraud in the name of executive management, traditionally by email. See [BEC in M365](/en/soc/business-email-compromise-m365). Deepfake vishing is the current variant via phone with an AI-cloned voice. Both vectors increasingly appear combined.

Can a SOC even detect a vishing call?

A SOC cannot directly detect the call itself unless telephony is integrated with the SIEM. It can detect and correlate accompanying traces: compromised mailboxes, MFA bombing, consent grants and unusual payment runs in ERP logs. A mature SOC links these signals into one case and stops the follow-up actions.

What to do when a deepfake call has come in?

Verify immediately through a second channel using a stored number. If actions were already triggered, stop the payment through your bank, lock affected accounts, engage the SOC and launch the incident-response playbook. Capture call metadata (CLI, time, duration, trunk) forensically. Assess revFADP (revised Federal Act on Data Protection) notification duties if personal data is involved.

Continue reading in this cluster
Detecting and stopping Business Email Compromise in Microsoft 365
Business Email Compromise (BEC) in Microsoft 365 rarely involves malware. The attack chain involves phishing, session or token theft, inbox rules and OAuth consent abuse. A SOC detects BEC by correlating signals from Entra ID, Exchange Online and Defender for Cloud Apps. The email body alone is insufficient for detection. Responders revoke sessions, remove inbox rules, withdraw OAuth consents and enforce MFA again. They document these actions in line with ISG and insurance requirements.
AI-driven cyber attacks: what changes, and what is marketing
Generative AI changes cyber attacks in scale, language quality and personalisation. The underlying techniques remain unchanged. Phishing in flawless Swiss German, deepfake vishing against the finance team, auto-generated malware code and LLM-driven reconnaissance are today's reality. The kill chain, detection logic and the importance of fast response remain unchanged. Attack volume, quality and the ability to bypass security controls based on linguistic or behavioural anomalies are changing.
Security awareness training: turning click risk into reporting behaviour
Security awareness training is a measurable behavioural process, beyond a mandatory e-learning module. The target is not a zero click rate but a reporting rate for suspicious mail in the 40 to 50 percent target band before the SOC escalates. ANOMAL combines short role-specific modules, realistic phishing simulations and a reporting interface in Microsoft 365 so awareness becomes a detection source. The 'Regular security awareness training with phishing simulations' requirement of most Swiss cyber insurers is documented in the process.
Creating an incident response plan: template, roles, notification chains
An incident response plan defines who decides, who receives notifications and the order for shutting down, isolating and restoring systems before a crisis. It references the five core controls cyber insurers require: MFA, EDR, segregated backups, a patch process and the documented IR plan itself. It also assigns binding notification deadlines under revFADP (revised Federal Act on Data Protection), ISG, FINMA and DORA to specific roles and response times. Without this plan, teams improvise the 72-hour response, and insurers can contest the payout.
ITDR: Identity Threat Detection and Response for Switzerland
ITDR (Identity Threat Detection and Response) detects attacks on the identity itself, not just endpoints or networks. Targets are accounts, tokens, sessions, permissions and identity providers such as Entra ID or Okta. ITDR extends EDR and SIEM with signals only visible in the identity layer. These include impossible travel, consent phishing, refresh-token abuse, role abuse and attacks on federation and directory objects. For Swiss organisations running M365, Entra ID and regulated processes, ITDR today is as important as EDR was five years ago.