AI Vishing Simulations: How to Run Voice Phishing Drills That Actually Change Behavior
Step-by-step guide to running AI vishing simulations: personas, targets, consent, metrics that matter, and the training research mistakes to avoid.
Ask a room of security leaders how many run phishing simulations and nearly every hand goes up. Ask how many run voice phishing simulations and the hands drop.
Attackers noticed. They clone a voice with as little as three seconds of audio scraped from a conference video or voicemail greeting. They spoof a caller ID. They call your finance team or help desk with an urgent request that sounds exactly like the person it should sound like. The FBI logged $893 million in AI-related scam losses in 2025, and voice-based attacks on businesses rose roughly 1,300 percent as cloned executive voices became a standard upgrade to business email compromise.
Your employees have trained reflexes for suspicious emails. They have none for suspicious phone calls. This guide covers how to build vishing simulations that change what employees actually do under pressure.
What a Vishing Simulation Is (and What It Is Not)
A vishing simulation is a controlled voice phishing drill. A security team places a call to employees using an AI-generated voice, a scripted or live-adaptive attack scenario, and a defined goal: get the target to share a credential, approve a payment, or disclose sensitive information. Then the team measures what happened.
Did the employee comply? Did they hang up and call back on a number from the internal directory? Did they report the call?
A simulation measures behavior under pressure, not quiz knowledge. Everyone scores 90 percent on a multiple-choice test about social engineering. The useful data is what a stressed person does when a voice that sounds like their CFO demands a wire before close of business.
Not all voice drills are equal. Some platforms drop a scripted voicemail recording and call it a simulation. That tests very little, because real attacks adapt. A modern vishing simulation uses a live AI conversation that responds to pushback, escalates pressure, and tries to talk around the employee’s objections, the same way a real attacker would.
A simulation also isn’t a gotcha. If your program exists to catch people failing so you can point at them, you’ll get exactly one honest data point before employees stop picking up unknown numbers and stop reporting anything.
The Scenarios Your Drills Need to Cover
You can’t design a drill until you understand the playbook it rehearses against. Security researchers at Group-IB and Mandiant have documented the workflow, and it’s simple enough to reproduce at scale:
- Harvest audio. Podcasts, keynote videos, webinars, voicemail greetings. Three seconds is sometimes enough; more samples push clone accuracy to 95 percent.
- Clone the voice. Commercial tools like ElevenLabs or open-source frameworks produce eerily accurate replicas, including emotion and conversational tics. Some cost under $50 a month.
- Spoof the caller ID so the call looks like it comes from the executive or an internal number.
- Call with an urgent pretext. A granddaughter in jail. An overdue invoice the CEO needs paid today. An IT outage that requires an immediate password reset.
Mandiant’s red team ran this exact play against a real organization. The cloned voice was so trusted that the victim bypassed browser and Windows security prompts to download a payload. MGM Resorts lost roughly $100 million from a single vishing call to its help desk. LastPass caught a WhatsApp audio deepfake of its own CEO only because a trained employee spotted the red flags.
The pretexts your drills should mirror:
- Payment authorization from the CFO, often referencing a real invoice or vendor
- Password or MFA resets from IT support, sometimes riding a real outage for credibility
- Vendor verification calls to procurement
- Executive assistants getting urgent requests that route around normal process
The multi-channel attack deserves its own drill. A spoofed email arrives first to prime the target. Hours later, a voice call references it. A text confirms the detail. Each corroborating signal erodes skepticism, and an employee who would dismiss a lone suspicious call complies when email, phone, and SMS all point the same direction. Testing one channel at a time doesn’t test the attack.
Consent, Ethics, and Legal Ground Rules
Vishing simulations put employees under real psychological pressure. That’s the point, and it’s also why you need ground rules.
Get written approval from an executive sponsor before the first call goes out. Define the scope: which roles, which scenarios, which channels, and what the exercise measures. This protects the program when someone asks why the CFO’s voice called them asking for gift cards.
Never simulate a personal crisis. Family emergencies, medical situations, and anything touching protected characteristics are off limits, and not just for optics. A drill built on a fake dying-relative scenario will poison the reporting culture you’re trying to build. Stick to legitimate business risks: urgent payment changes, credential resets, confidential data requests.
If you clone a real executive’s voice, which is the most realistic option, you need documented consent, access controls on the voice model, and a clear internal notice path so authorized administrators know the drill is running. Use fictional accounts, non-functional links, and payment instructions that cannot trigger real transfers.
Non-punitive design is also the more effective choice. The phishing research is blunt on this point, and we’ll get to why in a moment.
Building Your First Simulation, Step by Step
Step 1: Establish your risk baseline
You can’t simulate everything, so start where the money and credentials live. Finance teams, IT help desks, executive assistants, and HR carry the highest-value targets because of what they can authorize or access. 73 percent of AI impersonation attacks hit financial or payroll staff. Map who reports to whom, which executives have public audio, and which payment workflows a cloned voice could exploit.
Step 2: Select scenarios per role
The same drill can’t serve a finance analyst and a help desk technician. The pressure tactics differ, the verification protocols differ, and the attack goals differ. A help desk scenario might dangle a locked-out user and a VPN outage. A finance scenario references a vendor payment that’s “already late.” Role-specific scenarios are the only way the right reflexes land with the right people.
Step 3: Configure the caller
Decide the persona, the voice, and the pressure profile. A generic professional voice tests baseline response. A cloned executive voice, with consent, tests whether employees treat familiarity as proof. Set the urgency level and how the AI caller behaves when pushed: does it deflect a verification request, apply more pressure, or try a different angle? Live-adaptive platforms handle this in real time. If your tool can only play recordings, note that limitation in your results.
Step 4: Define success criteria before you dial
Decide in advance what counts as a fail (shared a credential, approved a transfer), what counts as a pass (independent callback to a directory number), and what counts as a report (flagged the call to security). Vague criteria after the fact make every result arguable, which defeats the purpose of measuring.
Step 5: Give every scenario a trusted verification route
Most programs skip this step. Every scenario needs a real path to safety: a directory number to call back, a code word to request, a second approver to consult. The drill should rehearse the correct action instead of only catching the wrong one. An employee who ends the call and says “I need to verify this through our standard process” should pass, every time, no matter how convincing the fake was.
On build versus buy: dedicated platforms (Callstrike, Jericho Security, Living Security, among others) handle live adaptive calls, cloned voices, and analytics. But practitioners on r/cybersecurity have also run effective DIY drills: clone an executive voice from a public interview, target IT and finance with WhatsApp voice messages, then present the whole thing at a town hall. One team reported that the demo itself, showing how easy the clone was to make, was the most persuasive training artifact they’d ever used.
Follow-Up Training That Doesn’t Backfire
Voice simulation programs can inherit email’s mistakes. The phishing research has been rough on conventional training, and the findings transfer directly.
A randomized controlled trial at UC San Diego Health spanning 19,500 employees found that annual mandated training showed no significant relationship with phishing outcomes, and embedded post-click training reduced failure by only about two percentage points. Most people spent under a minute with the embedded material. Worse, ETH Zurich researchers found that immediate embedded training can make employees overconfident, which left some more susceptible later. And mandatory retraining didn’t help the most susceptible people at all.
Three moves work better at the moment of failure:
- Brief, non-punitive acknowledgment. The employee learns what happened and which tactic caught them, without shame and without a 40-minute module nobody reads.
- Immediate practice of the correct action. If someone complied with a vishing call, have them rehearse ending the call, finding the directory number, and completing the callback. Convert the near miss into muscle memory.
- Delayed, department-level education. Deliver the broader lesson to the whole team on a delay, not just the person who failed. This sidesteps the overconfidence mechanism while spreading the knowledge.
Skip public rankings entirely. A failed drill identifies a process gap that security, finance, and leadership close together.
The goal is a culture where reporting a suspicious call is a rewarded action, and where an employee who isn’t sure whether a voice was real reports it anyway. That employee did their job.
Metrics That Actually Tell You If It’s Working
Completion rates and quiz scores measure delivery. Behavior under pressure is what you actually care about, and it shows up in five numbers.
Fail rate. The percentage of simulations where the employee complied with the attack goal. This is your primary indicator. A mature program targets under 5 percent, and you track the trend over time, not any single result. One bad month against a harder scenario tells you less than six months of trajectory.
Answer rate. The percentage of simulation calls that get answered. If answer rates fall while fail rates stay high, you have avoidance, not awareness. Employees who stop picking up unknown numbers are just harder to test, and the real attacker will try a different door.
Report rate. The percentage of employees who flag the call through your approved reporting channel. This is the behavioral outcome to optimize. An employee who reports a call they couldn’t identify as fake did exactly the right thing.
Time-to-report. How fast reports come in after the call. Shrinking time signals that pause-and-report is becoming automatic rather than deliberate.
Verification-completion rate. The subtlest and most important metric. Distinguish an attempted verification from a completed one. Recording whether the employee actually called a known number, confirmed through a trusted system, or obtained a second approver before acting is what tells you the control works. A verification attempt that never confirms anything is theater.
Two disciplines keep these numbers honest. First, score deepfake video, vishing, SMS, and email tests with the same rubric, because an employee who reports suspicious emails but accepts a cloned voice call has an unguarded voice channel, not a security-aware workforce. Second, review results monthly by department. Persistently high fail rates in a role group mean you increase simulation frequency and specificity before the next real attack arrives, not after.
One documented data point to calibrate expectations: organizations running structured vishing simulation programs have reported verification behavior improving by around 65 percent, and continuous simulation-based training cutting successful compromises nearly in half over twelve months. Treat vendor numbers with healthy skepticism, since independent verification is thin in this category, but the direction is consistent: frequent, realistic practice moves the numbers that count.
Making Verification the Default
Simulations create the reflex. The surrounding system determines whether the reflex has anything to grip.
Write the verification protocols down and train them. A zero-trust voice policy: a familiar voice never confirms identity, and any request touching financial authorization, credential access, or sensitive data requires verification through an independent channel. Callback verification as a documented four-step procedure. Code words for executives’ direct reports and finance teams, rotated on a schedule and shared out-of-band. Dual approval for any payment change. These are organizational guardrails, and they protect even an employee who fully complies with a convincing attack, because no single person can authorize the transaction alone.
Make leaders part of it. When executives undergo the same simulations and visibly endorse verification, employees stop treating “hang up and call back” as an accusation. A real executive survives the verification process. A scam usually can’t.
Then test the system itself. Documentation satisfies an auditor. It does not satisfy an attacker. Until your callback procedure, dual approval, and help desk verification have been stress-tested by a controlled drill where someone’s job is to defeat them, the policy has never actually been proven.
Run the drill, measure the behavior, coach without blame, and repeat until a cloned voice calling your finance team hits a wall of callbacks, code words, and second approvers instead of a wire transfer.
Running These Drills with Brightside
If building all of this in-house sounds like more craft than your team has time for, Brightside’s AI vishing simulator was built for exactly this workflow. It places live, adaptive calls that respond to what the employee actually says rather than playing a recording at them. You set the attack goal and caller persona, and the platform can generate the persona, draft the opening line, and recommend a tactic mix with a short explanation of why each works. Executive-impersonation drills support self-serve voice cloning from a one-to-two-minute recording, under the consent rules described above. Hybrid scenarios pair the live call with a trackable phishing email in one campaign, so you can rehearse the multi-channel attack instead of one channel at a time. An in-browser preview lets your team hear the call before it goes out.
The reporting side matches the metrics in this guide: a vishing-specific dashboard tracks answer rate, failure rate, call duration, and trends. A cooling period prevents retargeting the same employee too soon, and a failed simulation triggers follow-up training automatically, which lands the brief, non-punitive coaching the research supports.
Get a complete live walkthrough
Book a call with our team for a full overview of the platform, and bring any questions you want answered. No obligation exploration call.
Try our vishing simulator
Experience the most advanced voice phishing simulator built for security teams. Create scenarios, test voice cloning, and explore automation features.
Latest articles
Live Vishing Simulation vs Pre-Recorded Calls: What the Difference Actually Trains
The Help Desk Callback Verification Script: What to Say and When
Product Update: What's New in the Brightside Vishing Simulator