Written by: Kamil Ponicki, Director of Talent Acquisition, Digital Colliers
AegisAI closed a $36M round this month to fight AI-driven spear phishing, and the timing tells you everything. We're now at the point where three seconds of clean audio is enough to clone a caller's voice well enough to fool a human agent and most legacy voiceprint systems. If your call centre still treats "the voice matches" as a verification factor, you're running a control that attackers have already priced in.
This isn't a future problem. It's a right-now problem, and the regulators know.
Voice was never a strong factor, it was a convenient one
Voice biometrics worked because cloning used to be expensive and slow. That economic gap is gone. Consumer tools produce convincing clones from short samples pulled off LinkedIn videos, podcast clips, or the customer's own voicemail greeting.
So the honest read is this. Voice is now something the caller has, not something the caller is. It behaves like a password that leaked years ago. You can still collect it, but you can't lean on it alone to move money, reset credentials, or authorise a beneficiary change.
CTOs I talk to in financial services get this intellectually. The problem is that ripping voice out of the flow feels worse than leaving it in, because nothing obvious replaces it at the moment of the call.
The layered signals that should meet at verification
A modern verification decision isn't one factor. It's a small bundle of signals that arrive together and get scored in the same second the agent is on the line. The interesting engineering work is the join, not any single signal.
The layers that actually matter:
- Device and channel telemetry. Is this call coming from a number and carrier path consistent with the customer's history, or is it a spoofed originating number over a VoIP hop that showed up yesterday.
- Behavioural signal from the app or web session running in parallel. Real customers usually have the app open. Attackers usually don't.
- Transaction context. A password reset request from a device the customer has used for two years is not the same event as one from a fresh handset in a new geography.
- Liveness and challenge on the voice channel itself. Not "does the voice match", but "can this caller respond to a dynamic prompt in a way a pre-recorded clone can't fake".
- Out-of-band confirmation for anything high-value. Push to a known device beats voice every time.
None of these are new. What's new is the requirement to fuse them at call time, with a decision engine the agent can see and the auditor can inspect later.
The gap in mid-market call centres
Tier-one banks have been building this stack for a decade. The gap is in the mid-market, where the call centre platform is often a hosted product with voice biometrics bolted on, and the fraud team lives in a different tool that only sees transactions after they post.
The pattern I keep seeing is familiar. AML transaction monitoring already runs at 85 to 95% false positive rates in these shops, so the fraud analysts are drowning before you add real-time call verification to their queue. Meanwhile, the call centre supervisor's dashboard has no signal from the mobile app team, and the mobile app team has no idea a suspicious call is in progress. Three systems, three owners, no join.
This is a data model problem before it's a vendor problem. If your customer identity, session telemetry, and transaction stream don't share a key that resolves in under a second, no amount of AI on top will save the verification decision.
The exam question your auditor is about to ask
DORA has been in force across the EU since 17 January 2025. EU AI Act transparency obligations under Article 50 apply from 2 August 2026, and the high-risk obligations that catch a lot of financial services use cases land on 2 December 2027. The exam question is already written, and it goes like this. You knew voice cloning was cheap. What compensating controls did you put in place, when, and how did you evidence them.
The cost of getting this wrong is not abstract. GDPR exposure alone runs to €20M or 4% of global turnover, and EU AI Act fines for high-risk violations reach €15M or 3%. Stack a customer-harm event on top of a regulator finding that your verification control was known-broken, and the number gets uncomfortable fast.
The operators moving now are the ones who'll have a clean answer in the exam room. The ones still running voice as a primary factor in 2027 will be writing very expensive post-mortems.

