Live Vishing Simulation vs Pre-Recorded Calls: What the Difference Actually Trains
A live vishing simulation holds a two-way conversation; a pre-recorded call plays a message. Here is what separates the five mechanisms sold under one name, and how to verify which one a vendor ships.
Read three vishing simulation datasheets and the same three words appear on all of them: conversational, interactive, realistic. The products behind those words do different things on the phone. One dials an employee and plays a recorded voicemail. One dials and reads a script through a text-to-speech engine, advancing through steps. One puts an AI agent on the line that pursues an objective through whatever the employee says back. The vocabulary doesn’t separate them, so two security teams can read the same word on two datasheets and buy products that behave nothing alike.
The thing worth comparing is the mechanism: what happens on the call at the moment the employee opens their mouth. That single property decides which human behavior the exercise can put under pressure, which numbers on the dashboard carry information, and whether the drill resembles the attack it’s meant to rehearse.
Key Takeaways
- People can’t reliably hear that a caller is synthetic. In a 4,100-person study, participants identified AI callers with 70.3% accuracy in voice conditions and frequently flagged real humans as AI. Familiarity with AI didn’t help.
- The skill that survives a real call is conversational: noticing pressure, declining the request as stated, ending the call, and verifying through a route the caller didn’t supply.
- Real vishing keeps the target talking because the attacker needs live input, most often an MFA code that doesn’t exist until the attacker submits the stolen password. In controlled research, successful automated vishing calls averaged 92.4 seconds of conversation.
- At least five different call mechanisms are sold as vishing simulation, from a voicemail drop to a live adaptive agent. Each one caps what the program can measure.
- Two questions separate them on a demo: what happens when the employee asks something the scenario didn’t anticipate, and can an admin hear the call before an employee does.
What an Employee Has to Do While the Call Is Happening
Most voice-security training rests on an assumption worth testing: that the employee’s job is to notice the call is fake. If that were the skill, the format of the simulation would barely matter. You could play people a synthetic voice, ask them whether it sounded off, and score the answers.
The evidence says people can’t do it. A July 2026 study of AI-automated voice phishing surveyed 4,100 US adults and ran twelve qualitative interviews, exposing participants to scam scenarios generated by leading voice models alongside human baselines. Participants achieved “70.3% accuracy in voice conditions” when asked to identify AI-generated callers, and they “frequently misidentified human callers as AI.” Exposure to AI made no difference. People who reported never using AI detected AI voices 54.4% of the time; people who use it often managed 51.2%. Both are close to a coin flip.
Compliance held up even so. Across five scam categories, 16.5% of participants said they would or might comply with the phishing request, rising to 36% for a cloned voice claiming to be a relative in distress. Those figures are stated intent under survey conditions rather than measured behavior on a live call, and they are the only large-sample numbers of their kind available.
The study’s authors draw the obvious conclusion. They argue for training that focuses on “recognizing manipulative conversational strategies, rather than detecting AI-generated speech,” and for teaching people to resist urgency and emotional pressure and to verify caller identity through independent channels. The market reached the same place on synthetic media generally, where detection accuracy drops sharply outside the lab and detection tooling works as a signal layer rather than a control you can build a program on.
The skill is behavioral, and it breaks down into four things a person does under pressure:
- Register that the request is unusual for the channel it arrived on
- Decline the specific thing being asked, while the person asking is still on the line
- End the call without resolving the caller’s apparent problem
- Re-establish contact through a number or portal the caller didn’t provide
All four happen while somebody is talking. Three require saying no to a person who has an answer ready. A voice program exists to build that behavior, which is why the mechanism of the simulation determines the value of the exercise.
Why Vishing Only Works as a Conversation
In May 2026, Google Threat Intelligence Group published its analysis of UNC6671, the extortion cluster behind BlackFile, which had targeted dozens of organizations across North America, Australia and the UK since early that year. Its summary of the technique runs to one sentence: “The vishing call functions as a live adversary-in-the-middle (AitM) attack.”
A caller reaches an employee on their personal mobile, posing as internal IT or the help desk, citing a mandatory passkey migration or a required MFA update. The employee is directed to a lookalike subdomain mirroring their own SSO portal. As they type their username and password, the attacker captures both and immediately submits them to the real identity provider. The real provider issues an MFA challenge. The employee, believing they’re completing a setup step, hands the code or the approval to the caller, who uses it before it expires. The attacker then registers their own MFA device and starts moving through connected SaaS.
That code doesn’t exist until the attacker submits the password. Nobody can record a request for it in advance, because the thing being requested is generated seconds earlier by an action the attacker takes during the call. The attack needs a live human on the line, listening and responding, or it doesn’t function at all.
The pattern repeats in a different sector with a different objective. Mandiant’s June 2026 report on UNC3753, tracked elsewhere as Silent Ransom Group, documents a data-theft extortion campaign against dozens of US legal, financial and professional services organizations between January and May 2026. Callers pose as the internal IT help desk or security team and, per the report, “use a variety of verbal instructions to guide target behavior,” building trust and steering the employee into a screen-sharing session and then a remote monitoring and management install. An email using a data-migration or invoice pretext arrives first, so the story is already sitting in the inbox when the phone rings. The full chain can complete inside a single business day.
Both campaigns put the decisive moments mid-call: the employee is being talked through a sequence, and every step is an opportunity to stop. Both also start in email, which is why a coordinated phishing and vishing exercise tests something a call alone doesn’t.
ViKing, an AI-powered vishing system built at Instituto Superior Técnico from publicly available components and published at AsiaCCS 2025, extracted sensitive information from 52% of 240 participants in a controlled experiment. Across all calls, the average duration was 160.2 seconds. For the calls that succeeded, it was 92.4 seconds. One study with participants role-playing employees at a fictitious company sets no benchmark for anyone’s program. The scale is still instructive: the operative part of a successful automated vishing call is about ninety seconds of back-and-forth.
The same study measured how far a warning gets you. In the wave given minimal caution, 77% disclosed sensitive information. In the wave given the strongest warnings, 33% still did. Warnings move the number a long way and leave a third of the population disclosing, which is a rehearsal problem rather than an information problem.
Cost is the encouraging part. ViKing’s calls ran between $0.50 and $1.16 per successful attack, and the figure rose with the target’s awareness level, because the underlying speech and language services bill per minute and per character. A workforce that keeps callers talking and gives up less is measurably more expensive to attack. The August 2026 attempted wave against major hedge funds shows what happens when firms are ready for the call.
Five Things Sold as “Vishing Simulation”
The market isn’t split between live and pre-recorded. At least five distinct mechanisms are in circulation, and they form a ladder, ordered by how much the call can respond to the employee. The gap between adjacent rungs is often larger than the gap in price.
| Mechanism | What happens when the employee speaks | What the employee practices |
|---|---|---|
| In-browser or content-only scenario, no outbound call | Nothing. The employee clicks or types through a learning module at a moment they chose | Recognizing a described tactic. Useful as knowledge, not as rehearsal |
| Pre-recorded voicemail drop | Nothing. The recording plays to the end regardless | Deciding whether to act on a voicemail, and whether to report it |
| Inbound callback lure | The employee dials a number and reaches a recording, a menu, or an operator | Whether they’ll call a number that arrived unsolicited |
| Outbound text-to-speech call on a template or branch tree | The script advances, ignores the input, or ends the call. Unanticipated answers fall through | Listening to a plausible ask and hanging up or complying |
| Outbound live conversation, AI agent pursuing an objective | The agent answers the question, handles the objection, and keeps working toward its goal | Declining a request to someone who pushes back |
Voice cloning appears at several of these rungs, so a cloned voice tells you nothing about whether the call is a conversation. A platform can hold a two-minute recording of your CFO, generate a convincing synthetic version, and then play it as a fixed message. Realism of the voice and adaptivity of the call are separate purchases.
The word “conversation” gets applied across the bottom four rungs too. Vendor copy describing “realistic conversations” built on “advanced text-to-speech” and “multi-step customizable scenarios” is describing a script with steps in it. That may be genuinely useful, and it’s a different mechanism from an agent that improvises. The datasheet won’t tell you which one you’re looking at.
Live adaptive calls aren’t rare, and no vendor owns them. Six of the fifteen platforms we track place live adaptive calls, a group that mixes voice specialists with broader suites. The dividing line is real, and it runs through the middle of the market rather than around any one product.
What Each Format Can Measure
The mechanism sets a ceiling on the dashboard, which is where the wrong rung eventually costs something.
Answer rate works at every rung. Picking up the phone is the one behavior every format provokes, and it’s genuinely useful: it tells you the phone data is correct and how reachable a population is.
Failed rate needs a request the employee can refuse. Where nothing is asked, as in a voicemail drop, there’s no failure to record. Where the ask is a keypad entry, the metric captures whether someone typed digits into a phone, and nothing about whether they’d have kept talking to a person who had an answer for their hesitation. They never met one.
Call duration only carries information when the employee’s own choices determined the length. On a recorded drop, duration is the runtime of an audio file, plus however long it took to hang up. On a live call, the same number becomes behavioral: how long the employee stayed engaged, whether they asked questions, how far they got before disengaging. Median duration on a scripted platform and median duration on a live platform are different measurements wearing the same label.
Reporting rate works everywhere and is the number a recorded drop measures honestly. If the exercise is “did anyone tell us a strange call came in,” a recording answers it.
Watch for a program that looks healthy on numbers describing the audio file rather than the workforce. A 4% failed rate on keypad entries and a two-minute median duration set by a script are real figures that don’t answer the question a CISO is actually being asked, which is whether the finance team would wire the money.
Where Recorded and Scripted Calls Still Earn a Place
The lower rungs do specific jobs, and they do them at a cost live calls can’t match.
A voicemail drop is a defensible way to baseline answer rate and validate phone-number data before you spend anything on live calls. Bad numbers are common, and finding out during a live campaign wastes the campaign.
Scripted text-to-speech calls give a large population cheap first exposure to the fact that the phone is an attack channel. For a low-risk cohort with no access to reset credentials, move money, or authorize applications, that may be all the voice coverage the risk justifies.
An inbound callback lure tests one real and specific behavior that no outbound call tests: whether people will dial a number that arrived unsolicited. Callback-style attacks work that way, so the exercise matches the threat.
None of the three can put a decision under pressure in front of an employee, because none of them can push back. So they can’t stand in for coverage of the roles attackers actually call, which is help desk and service desk staff, finance and treasury, executive assistants, privileged users, and executives themselves. Both the UNC6671 and UNC3753 campaigns went after people who could grant access or release documents.
The sensible split is usually broad low-cost coverage at the lower rungs and live calls aimed at the roles that get called. Settle that before you shortlist, because it’s a program design decision more than a product one, and it changes which vishing simulation software will fit.
Questions That Reveal Which Mechanism You Are Buying
A datasheet won’t answer any of this. A demo will, if you ask directly.
What happens if the employee asks something the scenario didn’t anticipate? The useful answer describes behavior: the agent answers and continues toward its goal. A vague answer about handling a wide range of inputs usually means a branch tree with a fallback.
Can the caller pursue a goal, or does it follow steps? These are different architectures. Ask whether the campaign is configured with an objective the agent works toward, or with an ordered sequence. A platform built on steps will describe steps when pressed.
Can I hear the call before an employee does? This separates platforms quickly, because it’s rare. If an admin can’t preview the live call, nobody at your company knows what your employees are about to hear, and the first person to evaluate the simulation’s quality is the person being tested.
Is the call outbound, or does the employee dial in? Callback flows and outbound calls test different behaviors. Both are legitimate; you should know which you’re buying.
Does the platform report median call duration, and what determines the length? Ask the second half of that question. If the length is set by the script, the metric is describing the vendor’s content rather than your employees.
What tier or add-on does the voice capability need? Voice frequently sits above the tier that covers email simulation, so confirm it before it becomes a budget surprise.
One shortcut covers most of the list: ask to hear a live call and put an unscripted question to the agent yourself. A vendor who can’t or won’t demonstrate that has answered the first question.
How Brightside Shapes a Live Simulation Call
A live conversation is the entry price rather than the differentiator. Several platforms place live adaptive calls, so the question worth asking a vendor that clears that bar is what an admin actually controls once the call is running.
The campaign is built around an attack goal, the agent’s objective for the whole call rather than a step in a sequence. An admin describes the specific information or action the agent should try to obtain, using a preset such as a 2FA phishing link or a fraudulent invoice approval, or writing their own. That goal is the mechanism behind the adaptivity: an agent holding an objective rather than a route can absorb a question or an objection and keep working, which is what makes the resulting conversation resemble the ones in the GTIG and Mandiant reports.
Around the goal sits the attacker persona: caller name, job position, organization, plus a free-text context field for the details that make a call sound internal, such as a ticket number or recent account activity. Then the social engineering tactics the agent will use, selected by the admin: applying pressure through risk, appealing to authority, using professional jargon. Tone is set separately, and delivery notes can ask the agent to speak slowly or leave more pauses. The first spoken line can be left as the default greeting, written by hand, or generated from the setup already configured. An academic review of 86 vishing attacks found deference to authority exploited in 95.3% of them, which is a reasonable place to start.
Voices in the library are built for phone calls, so they sound like a person on a handset rather than a studio recording, which matters more than fidelity for a call that’s supposed to be routine. A custom voice is cloned from a one to two minute recording, self-service, and that covers executive impersonation. Voice languages currently include English, French, German, Italian, Polish, Spanish and Dutch, with more added on request.
Real campaigns arrive on two channels, and a simulation can too. Voice + BEC runs the call and the phishing email as one coordinated attack, with the email scheduled before the call, so it’s already in the inbox when the agent refers to it, or after the call ends. That configuration reproduces the UNC3753 shape: invoice pretext by email, help desk voice on the phone.
Test launch is the final step before anyone else is contacted. The admin receives the email, hears the live AI call, or both, which is how you find out that a persona doesn’t hold up or a tactic reads as absurd before the audience hears it.
Vishing runs on Pro or the Voice add-on, and there’s a fuller walkthrough of Brightside’s vishing simulator if you want the mechanics in more detail.
Results come back as four numbers: failed rate, answer rate, median duration, and total simulations, with a failed-rate trend over 7, 30 or 90 days. Median duration is the one that repays attention on a live platform, because the length of each call was set by the employee’s own choices rather than by the runtime of a file. When that median moves, something about how your people handle a caller has changed.
Get a complete live walkthrough
Book a call with our team for a full overview of the platform, and bring any questions you want answered. No obligation exploration call.
Try our vishing simulator
Experience the most advanced voice phishing simulator built for security teams. Create scenarios, test voice cloning, and explore automation features.
Latest articles
The Help Desk Callback Verification Script: What to Say and When
AI Vishing Simulations: How to Run Voice Phishing Drills That Actually Change Behavior
Product Update: What's New in the Brightside Vishing Simulator