I spent 20 minutes counting everything I hate about Chat GPT’s new voice — and accidentally discovered why it works | Tech Radar
Overview
News, deals, reviews, guides and more on the newest computing gadgets
Start exploring exclusive deals, expert advice and more
Details
Unlock and manage exclusive Techradar member rewards.
Unlock instant access to exclusive member features.
Get full access to premium articles, exclusive features and a growing list of member rewards.
I spent 20 minutes counting everything I hate about Chat GPT’s new voice — and accidentally discovered why it works
I hate Chat GPT’s new 'ums', 'rights' and awkward pauses
When you purchase through links on our site, we may earn an affiliate commission. Here’s how it works.
I don’t like Chat GPT’s latest voice model. If you don’t use voice to interact with Chat GPT, this might seem like a strange thing to get worked up about, so let me explain.
Until recently, its voices have all sounded fairly advanced. But there was still something unmistakably AI-y about them. At times, they sounded smooth and even expressive, but also stunted and, well, robotic.
The latest voice update attempts to smooth away the rest of those more robotic edges. Which means that Chat GPT now pauses, hesitates, changes intonation and adds lots of little noises humans make without even thinking about it.
There are lots of “umms” and “hmms”. There are plenty of drawn out “yeeeahs”, as well as enthusiastic “ooohs”. Sometimes it sounds like it’s taking a breath or sighing. Sometimes it pauses halfway through a thought. And often its voice rises at the end of a sentence. In other words, it’s been changed to sound more human and I find it incredibly annoying.
When I first tried it, I thought it was broken. Especially because of the way it said a filler word then paused. Once I’d heard three long “hmmms” in a row, I couldn’t pay attention to anything else it was saying in response. So I decided to turn my irritation into an experiment. I opened the different voices in Chat GPT, started talking and counted all of the filler words.
I asked Chat GPT to change my mind about something I strongly believed
Chat GPT has stopped taking your prompts so literally, and that’s a big deal
I started with Maple, one of Chat GPT’s more cheerful, female-sounding voices, and within a conversation that lasted about 5 minutes I counted 20 “hmmms”. There were also lots of elongated “yeeeahs”, a couple of “uhhhhs”, an “ooooh”, a “hmm nice” and several enthusiastic “yeah, totally” responses.
There was also a lot of rising intonation, which is when the pitch of the voice goes up at the end of a sentence. That made me wonder whether I was just getting hung up on some speech quirks that were more associated with US English than Chat GPT, so I switched voices.
Next, I tried Vale. It’s a bright and inquisitive British female voice. Interestingly, Vale had a different set of verbal habits. There were fewer “hmmms”, but plenty of “ohhh yeeeahs”, long “suuuures” and a few “riiiights”. And its intonation also seemed to rise constantly, particularly when it was asking me questions.
I thought I’d try a more obviously male-sounding voice next. I chose Arbor, another British voice, which I found calmer but much more stilted. It gave me way less of those exaggerated conversational noises that I disliked in Maple, but replaced them with strange pauses and a staccato rhythm. At times it seemed to stop talking in really odd places.
After around 20 minutes of conversations across the different voices, I'd counted more than 100 filler words and conversational noises.
I tried to use Chat GPT to create fake evidence — and I came away worried
I gave Chat GPT 7 extra words and the answers instantly got better
I’d gone into these conversations with Chat GPT looking for things that irritated me. I was listening closely for every “umm”, strange pause and exaggerated “yeeeah”. I was intentionally trying to count all the quirks Open AI had added to Chat GPT’s voice to make it sound more like a person.
But, weirdly, the conversation started to feel different. I reluctantly began talking about priorities and my week ahead to test the voices. And instead of giving Chat GPT prompts and waiting for its responses, I soon found myself slipping into a much more natural conversation. This surprised me because I still didn’t like the voices.
The filler words reamined distracting and the pauses were still in all the wrong places. At no point was I convinced that there was a person on the other side of the conversation. But I had to admit that although on a conscious level I am wary about AI, and the whole point of this was to scrutinize the system, I still seemed to be responding to its more natural-sounding voice by getting increasingly wrapped up in the conversation.
Whether you find that cool or worrying probably comes down to your broader opinions about AI and its influence on us.
There’s an obvious reason for making Chat GPT sound more natural, which is that humans find natural conversation easier, and are therefore more likely to continue engaging with it.
Real, everyday speech isn’t a clean stream of information. All of us pause, hesitate, change pitch, make noises to show someone we’re listening and use filler words while we’re thinking about what to say next. These small but significant parts of speech do a lot of work.
So what are they doing when AI uses them? Because Chat GPT doesn’t need to say “hmm” when it’s gathering its thoughts or “yeeeeah” to show you it’s listening. It could just retrieve the information and respond. So these sounds are for making interactions feel more conversational, which is what makes me wary.
A calculator doesn’t need to add in a filler word before it gives you an answer and a word processor doesn’t need to pause and say “hmmm, yeah, totally!” when you start typing. But that’s what Chat GPT is now doing.
Some might argue that a voice with a more natural-sounding rhythm and intonation is simply just a nicer experience. But those same choices make interactions feel more social, potentially blurring the line between the role of AI as a tool and something we relate to more like another person.
After my experience, I wasn’t surprised to find research suggesting there might be more to this. Research suggests that voice-based interactions with AI might be associated with increase emotional engagement and anthropomorphism.
There’s also some suggestion that voice could lead people to overestimate AI, especially when a human-like voice is mistaken for understanding or competence.
This doesn’t automatically mean that adding some “umms” will make us all trust AI. But it seems that adding voice into the mix does change how we perceive and respond to it.
One more practical learning from this experiment was that if Chat GPT's voice suddenly drives you mad after the update, try another one. The differences between them felt bigger than I expected. Initially, I couldn't stand Maple's mannerisms and rising intonation, but Arbor's calmer delivery was easier for me to tolerate, even if its strange pauses created a different pacing problem.
I went into this experiment expecting to write about an annoying voice update. I thought I'd count some “umms”, complain about how artificial all of this manufactured naturalness sounded and then just return to typing. Instead, I spent 20 minutes deliberately studying the mechanisms designed to make AI sound more conversational and still found myself becoming more conversational with it.
I knew the hesitation wasn’t genuine hesitation. I knew there wasn’t anyone on the other side of the exchange thinking, listening and searching for the right words. Yet I still responded to the signals by talking more naturally and, eventually, being more open.
I don't think that means I was somehow fooled into believing Chat GPT was human. But I think what it suggests is that we don’t have to believe AI is conscious, or there’s some form of life there, for our deeply ingrained social instincts to start responding to human-like cues.
So, while I still wince at every “hmmm”, I don’t think whether I like the filler words or not really matters. It’s more about what they make us do and how surprisingly instinctive my response to them was.
Follow Tech Radar on Google News and add us as a preferred source to get our expert news, reviews, and opinion in your feeds.
Becca is a contributor to Tech Radar, a freelance journalist and author. She’s been writing about consumer tech and popular science for more than ten years, covering all kinds of topics, including why robots have eyes and whether we’ll experience the overview effect one day. She’s particularly interested in VR/AR, wearables, digital health, space tech and chatting to experts and academics about the future. She’s contributed to Tech Radar, T3, Wired, New Scientist, The Guardian, Inverse and many more. Her first book, Screen Time, came out in January 2021 with Bonnier Books. She loves science-fiction, brutalist architecture, and spending too much time floating through space in virtual reality.
You must confirm your public display name before commenting
Think CD players died in the 2000s? These 9 wonderfully weird new models prove otherwise
The US Army is training AI agents to work alongside human forces in 'work roles'
The US Air Force wants its next generation of target drones to mimic stealth jets
Quote of the day by Palantir co-founder Peter Thiel: "Anyone that has a monopoly will pretend that they're in incredible competition"
Tech Radar is part of Future US Inc, an international media group and leading digital publisher. Visit our corporate site.
© Future US, Inc. Full 7th Floor, 130 West 42nd Street, New York, NY 10036.
Key Takeaways
- News, deals, reviews, guides and more on the newest computing gadgets
- Start exploring exclusive deals, expert advice and more
- Unlock and manage exclusive Techradar member rewards
- Unlock instant access to exclusive member features
- Get full access to premium articles, exclusive features and a growing list of member rewards



