Daybreak - India's call centres are using AI to make agents sound American
Episode Date: September 3, 2026When you call an Indian call centre these days, the person answering might sound American — even if they've never left India. A new kind of AI can change an agent's accent in real time, mid...-call, to sound more familiar to the customer. It's meant to build trust and cut hiring costs. But the same software that smooths out an accent often smooths away everything else too — laughter, tone, the small things that make a voice sound human. Some customers have started wondering if they're even talking to a real person. We dig into how this technology works, and what it costs the people whose voices are being rewritten.Daybreak is produced from the newsroom of The Ken, India’s first subscriber-only business news platform. Subscribe for more exclusive, deeply-reported, and analytical business stories.
Transcript
Discussion (0)
Crashing your car is bad enough.
Now, imagine if it happened on a Friday the 13th.
That's exactly what happened to an American customer
whose call was picked up at an Indian call center for his insurance claim.
The call didn't go the way these ones usually do,
where you ask questions, the agent answers,
and then you read some kind of resolution together.
Instead, something very strange happened.
Here's my colleague, De Banjali Biswas,
sharing the story she heard from some.
someone who works at one of these centers.
At some point, you know, in the middle of this really stressful conversation of recounting
everything, to lighten the mood, the client jokingly says, oh my God, this is, you know,
this happened on a Friday that was also a 13th of the month.
So surely this is, you know, so Friday the 13th coded.
Like this surely couldn't be a worse luck thing.
The client waited for a laugh, a chuckle.
anything. But there was nothing from the other end of the line.
The client could not just keep sitting in silence. So at some point he just went, hello,
am I, is someone there? What's happening? And the agent then replied back saying,
oh yeah, of course, I'm here. Of course, the client felt embarrassed. They were sure that they had
just made an unfunny joke to a stranger who was probably already having a bad day.
Except that wasn't the case at all. On the other end,
someone who was auditing the call flagged it to the agent saying that,
why were you silent?
Why did you not say anything that was so awkward?
And the agent was like, no, I did.
I did laugh.
So the agent had laughed.
It's just that the laugh itself had never made its way down the line.
And that's because the agent wasn't really speaking to the customer directly.
Their voice had been live translated through an AI accent translator,
a software developed by Silicon Valley-based voice AI startup,
like Sanas. And to put it simply, it takes, say, an Indian accent and makes it sound like an
American one. The idea is to make international callers just the people they are talking to more.
From insurance companies like United Health to payment companies like American Express and
BPO's like Teleperformance and Alorica, they all use it. They often even outsource their work
to Vipro, the IT services giant, which also ends up using these translators by a
extension. Now, on the surface, the tech itself presents some great advantages. It allows companies
to hire cheaper talent from Tier 2 and Tier 3 towns and even helps them improve their customer
satisfaction ratings. In fact, the BPO I mentioned earlier, Teleperformance, which uses these
AI translators for about 42,000 of its Indian agents, claims that using Sunnis helped increase its
score by 26%. But on the other hand, like in the story you heard earlier,
The tech has real gaps. For example, it doesn't pick up on emotions very well.
And a lot of times, it even makes humans sound almost robotic.
Which is exactly what the customer on the other end thinks as well.
Many have even asked agents outright if they are in fact AI.
And these gaps have started to raise concerns.
The Banjali, who spent days working on the story, is joining me today for a deeper look
into what it means for the BPO sector's future when humans start sounding like AI at the same time as when AI starts sounding more human.
The voice translator fits under the larger category of voice AI products.
Some products are entirely automated agents with no humans in the loop, others fix the overall clarity and quality of a call,
but the translation products can change accents and even languages live.
They're obviously ideal tools for BPA.
to adapt because they get calls from all over the world.
And some other products also help them analyze calls without interfering in them to create
insights and other takeaways.
Which is why De Banjali said that the scale at which AI translators themselves are used in
Indian BPO's is hard to quantify because of the range of applications that these products
come with.
Plus, companies are also a little gaugy about the details.
She was, however, able to get a number from one of the companies,
leading in this space called Sanas AI.
The figure that we get from Sanas's end, which I got when I spoke to some of their employees,
was that specific to accent translation, they have released out around 200,000 licenses,
which means that basically 200,000 people, agents, are using accent translation.
The overall number of licenses that they've given out is roughly 500,000.
200,000 to 600,000 licenses.
So at this stage, the tech is still picking up steam, despite the problems it comes with.
Though, especially for Sanas, the motivation is almost noble.
They say that they are out to solve racism.
Their logic is basically that accent translation can build familiarity.
Think about it this way.
A US customer calling about health care benefits will find it easier to talk to someone who sounds American,
because the system and the way people relate to it is so specific to the country.
Satyam Vyvedi, the former data analytics head at Sanas, told De Banjali that in fact several
BPO's have actually moved to the Philippines for this reason, because even though it's
costlier, the people there sound more American.
So, while Sanis's logic makes sense, it does seem like a tall order for the company to
actually deliver on.
So I asked Debanjali, how much?
well it was holding up.
This is not just racism, right?
Sure, there's a part of it is that.
But the other part is also
the competence level of
the agent speaking. These are not
people who are sort of industry
experts in
the field that they are chatting
about. Many of them,
they learn about what they have to talk
about. They get manuals and they get handbooks
and they learn it. But a lot
of them might not be able to say
answer these questions on the fly in a non-formulic way.
They can stick to the script, but the minute you go off script, it's difficult.
They aren't that good at solving the issue.
And then on top of that, hearing the accent of someone, say, Filipino or someone Indian,
just sort of frustrates you beyond on top of that.
So on that front, yes, it does help solve racism,
because it sort of removes that immediate bias, right?
maybe the call center person is actually good at their job maybe they just need a little more time to figure out what your specific problem is right they might not know exactly all the context just of the top of their head um but when you hear a voice uh an accent from a different part of the world maybe you have an inherent bias and it takes that bias away and it sort of gives the conversation of fighting chance for you to actually do well right without
coming in with preconceived notions.
Honestly, that was a more optimistic outcome than I had personally been expecting.
But the company's success, even if it is just to give agents a fighting chance, like DeBanjali said, made sense.
Because the bias fixing aspect is actually part of the company's origin story.
Apparently, the founders had a friend, one of, either this was one of the co-founders or it was a friend of the co-founders,
who was from a Latin American country.
And during the pandemic, they needed to take a job.
And they were from, you know, an Ivy League school because they needed money and they took a call center job.
And despite being really qualified, really skilled, they were really struggling at the job.
They weren't doing well because of the accent bias that existed because they were from a Latin American country.
So this was sort of the inspiration to start a company like Sanas.
Fun fact, an employee told me that this is told to them in the interview rounds,
in every meeting they start with this as a core principle.
So you never forget that this is the starting story because it's repeated so often to everyone in the org.
And what this tech is allowing companies to do is hired differently.
Before, they were forced to ignore capable talent just because they had heavy regional accents.
In fact, they also had some very surprising priorities that Di Banjali found out about in her conversations with team leaders in the space.
You are looking for people from convent schools, people who have an English medium educational background.
One of them even mentioned this.
He basically just said, oh, we hire Catholics because they tend to speak English more at home.
So, you know, there are all these like social bias.
that come in even at the hiring level, whether they are filtering you by your school,
by whether you went to a Tier 1 city's missionary school or a Tier 1 city's Catholic school or not.
And now things have changed.
Pooja Ruth, a voice coach from one of India's top IT companies,
told Ebanjali that hirers have started looking for people with positive attitudes and leadership qualities.
It's also allowing them to cut down on costs quite a bit,
because English speakers with good accents charge about 30,000 to 40,000 rupees.
Regional talent, on the other hand, charge anywhere between 5,000 to 22,000 rupees.
With a player like Sana's coming in, one of the recruiters, she was telling me that they are able to cut costs a lot by hiring people who maybe aren't that fluent in English.
they are more comfortable speaking faster, communicating thoughts better in a different language, say Hindi.
And they're able to hire them at a much cheaper cost, like around $20,000 a month, right?
Some people, they're not great at a conversation, but they speak English really well, so they get hired.
But now the criteria has shifted where with these sort of AI accent translators, it's not just about whether you can speak the language,
but it's about are you actually good at holding a conversation?
Are you good at going back and forth and understanding what the client wants?
So this does help companies extend their talent pool.
And it saves agents themselves from having to go through the process of removing
Indianisms from the way they speak.
But the unintentional result, as we already know,
is that the tech is removing the non-verbal cues from conversations that make them feel more human,
like laughter, inflections or even fumbles.
And these issues in the tech are complicating another problem.
The fact that people are getting fed up with AI being forced into every part of their lives.
More on this in the next segment.
When a woman from Bangalore was trying to book a demo swimming class,
the agent on the other end kept trying to convince her that there was no slot available.
But she had already seen a free slot on the website.
And she didn't even realize that the agent was AI until she noticed that it kept comfortable.
at the exact same time during a certain sentence.
Pretty weird, right?
The woman thought so too.
She told Debanjali that she finally hung up the phone and booked the demo in person.
So we know now that the tech doesn't pick up laughing and emotion.
But actually, a manager at Vibro told DeBanjali that silence is actually one of the best case scenarios.
Sometimes something worse happens.
If you're laughing, these things might not be something the AI catches.
or maybe they do catch it but they don't translate it well.
Another thing that the AI agent sometimes does,
instead of just going completely silent,
is they will read the laugh that, you know,
someone has laughed,
and they'll translate it in like a very gurgled sort of robotic,
sounding glitchy sound, right?
That essentially sounds like the machine is malfunctioning.
And all that malfunctioning only increases
when the human agents themselves have less energy
to make up for what the tech misses.
One of the Sanas ex-employees mentioned to me
is that they saw a case study basically
where they realized that this was in the Philippines,
I believe, where towards the end of the shift
at a certain point of time and the day,
the call quality really dropped
and the amount of issues that kept coming up
where the agent was, sorry, the AI voice filter
was not catching things properly
or they were mistranslating things.
This was increasing at a certain point of time in the day.
And they realized that it was correlating with the fact that the workday was coming to an end.
And people were really exhausted and they just wanted to get home.
So the tech gets worse when the agents are tired.
And the consequences are quite real.
If you've ever been in these customer service loops, you know the drill.
By the time a customer like you or me even reaches a human agent,
we've already been through quite a few steps.
The first step in most calls is already an automated system.
Press one for this, two for that, three for something else.
The Bansji told me that these loops can drag on for 20, 40, even 60 minutes before a real
person ever picks up.
So when that person finally does answer, the customers are already at the limit.
And if they sound a little robotic, after all that, well, it makes sense that sometimes the
limit breaks.
But here's where the consequence comes in.
Because even when the actual conversation itself is fine, the frustration doesn't go away.
Instead, it starts reflecting in the satisfaction score.
BPOs have the system called NPS, which is a net promoter score.
In simple words, the client has to rate how well the call went, right?
It's a rating system.
You just rate one out of 10 on different parameters.
And when that rating step comes at the end of the call, once your issue has been resolved or maybe hasn't been resolved,
they mark it and if they've had a very frustrating call, even if the call itself was fine,
they'll sometimes rate it like a five or a four.
And in the comments, they will say things like, oh, the agent helped me out, but the process
was so frustrating.
So the agents end up being the ones taking the hit for a system that they don't even really
have a say in.
And these ratings are very important because if the scores fall too much, clients can shift
entire business elsewhere. For example, places like the Philippines. That would mean less reason
for BPO's here to hire and ultimately would result in layoffs. That's one stress that
agents need to deal with. And on top of that, with this new tech, they also need to ease customers'
suspicions about their humanness. A lot of times the agents clarify that they are humans
because the clients are suspicious that they are not human. Some of them might very directly just say,
am I speaking to an AI agent?
Am I speaking to a bot?
Some of them don't ask it directly because they are very suspicious.
And they go in roundabout ways.
They'll say things like, oh, where exactly are you based?
Right. Like with city are you based in?
What time is it there?
You know, things that might trip up a robot that does not have a physical existence to answer these things.
That's a pretty confusing spot to be in, right?
having to prove how human you are constantly.
Turns out, agents are getting quite creative with their own workarounds.
So a laugh or a scoff has to be literally spelled out.
Like someone instead of just laughing, they would physically say he on call.
Or that's funny, you know, like, oh, that was a great joke.
And some people take it just a little further.
There was an example where someone realized that, uh,
the client would understand that they're a person if they were really informal.
So some of them would say, yeah, bro, which is not something you tell a client.
You don't say yeah, bro, to a client.
And, you know, nine out of ten times they got away with it.
But at some point, a customer brought it up.
Yeah.
Someone complained being like, oh, that was very strange.
Like, I've never heard a call agent speak like that.
And then this was flagged, right?
So there are those sort of fringe downsides to these things.
things because it's again uncharted territory you're defining the rules by yourself.
Basically, agents are kind of writing their own rule books at this point. One improvised
yeah, bro, at a time. But all of this is obviously going to cost the industry. Though some jobs
are for now safer than others. Stay tuned. You see, some of what the tech does isn't a problem
and in some cases it's even necessary. Noise cancellation is one of those features. BPO's usually
have rows in cubicles filled with agents all speaking to customers at the same time. Without some
way to isolate an agent's voice, every call would sound like, well, it was being taken in a crowded
office. Some companies use Sannas or its competitors just for that feature. But the trade-off
is still there. You see, even things like a mic being hit or a scrape of a chair or just an ambient
murmur make a call sound real. Still, the tech is only.
spreading because the economic advantages of it are actually very tough to argue with.
There's a specific Gartner report that talks about the cost difference, right?
So when you are employing a human to come and help customers in customer service,
the cost for a human is more than seven times that of if you just used an AI bot.
So for very routine things, right, why would you pay that?
seven times more money when you can just have an automated bot.
Obviously, seven times is a pretty big deal, which raises the question.
If AI is seven times cheaper and it's increasingly sounding more human, where does that leave
the humans?
Di Banjali explained to me that when it comes to routine stuff like renewal update calls or
balance update calls or just reminders, AI is already taking over.
So, Niti Ayyog had put out a report in 2025 that talked about this threat.
And, you know, for them it was, they were just more so calling to action that something has to be done to address this issue.
And they estimated that currently the customer service sector employees, 2 to 2.5 million people.
And if something is not actively done to intervene, this headcount can shrink.
to as low as 1.8 million.
So we're talking about a lot of layoffs
and these are things that are already sort of happening.
So if the shrink is already happening,
what's going to happen to the BPO industry?
Di Banshi told me that it's finally going to come down
to a few sectors and cases where human sensitivity is still a necessity.
It's not that the BPO sector itself is going to disappear.
It's just that who staffs them, who runs them,
and who sits and makes the calls within them, that changes, that ratio changes.
So they will hire less people and they will buy licenses to more AI bots that can
autonomously do these calls.
So BPOs will survive.
The people within it, that dynamic itself will change.
Finance, healthcare, mainly, because these are ones that require.
Like finance, for example, has a lot of sensitive data, right?
Like, you're talking about people's money.
You're talking about less statements.
it's linked to, you know, identification details.
So these things require people in the mix, obviously,
to maintain certain levels of security
and also mainly accountability, right?
It's not just that people can also leak,
but you need a person to hold accountable.
And then, of course, in sectors like healthcare, similar issues, right,
where healthcare has a lot of data, patient data,
health conditions, family history, there is the emotional connect with both these sectors.
It's a lot of high priority, high risk, people's lives impacted.
So it's a mix of you need the human emotion, the human touch, as well as you need
accountability and a physical person to be held accountable.
What De Banjali just said, a blob of technology cannot be held accountable, isn't a new
idea. It's pretty much exactly what a code from a
1979 IBM training manual said. A computer
can never be held accountable and therefore
a computer must never make a management decision.
About 40 years later, it's surreal to think that
that idea is one of the last things standing between the
automation of an entire industry. And that brings us back to
the agent on the Friday the 13th call. They had laughed.
It was the system that failed to translate it.
And for now at least, maybe the fact that someone was in fact laughing on the other end of the line counts for something.
Daybreak is produced from the newsroom of the Ken India's first subscriber-focused business news platform.
What you're listening to is just a small sample of our subscriber-only offerings.
A full subscription offers daily long-form feature stories, newsletters and a whole bunch of premium podcasts.
To subscribe, head to the Ken.com and click on the red subscribe button on the top of the Ken website.
Today's episode was hosted and produced by my colleague Rachel Vargis and edited by Rajiv Sien.
