I’ve spent a good chunk of my life in rooms with people who are way smarter than me, building things that seemed impossible. We sold my first company, MovieLaLa, to Gfycat. My next, RemoteTeam, got acquired by Gusto in a deal we closed entirely over Zoom—not a single handshake. I’ve been lucky enough to write checks for over 200 companies, including some you might have heard of like Anthropic, OpenAI, and Scale AI. You learn a few things when you’re in the trenches for that long. One of the biggest? The gap between hype and reality can be massive.
And right now, the hype around AI voice is deafening. Everyone’s talking about cloning voices, creating AI-powered podcasts, and generating perfect, human-like speech from a block of text. The promise is incredible. The reality? Well, it’s a little more complicated.
The 40-Word Lie
Here’s a dirty little secret most text-to-speech companies won’t tell you: that amazing, mind-blowing demo you heard was probably the result of hours of fine-tuning and a team of audio engineers. They’ll tell you, “Just give us 40 words of any audio, and we’ll create a perfect clone!” I’ve tried it. It’s a lie.
What you get is a robotic, soulless version of the original voice. It might have the right pitch, but it’s missing the… well, the humanity. The subtle shifts in intonation, the pauses for effect, the way your voice speeds up when you’re excited. That’s the stuff that makes us sound like us. And that’s the stuff that most AI voice models are still struggling with.
I remember when we were building MovieLaLa. We were obsessed with creating a social experience around movies. We wanted to capture the feeling of talking to a friend about a film you just saw. We couldn’t have done that with a robotic voice. It would have felt fake, disingenuous. The same is true for podcasting. Your listeners are there for you. Your personality, your unique perspective. If you replace that with a generic AI voice, you’re losing the very thing that makes your podcast special.
I’ve seen this firsthand. I’ve advised a few startups in the audio space, and the temptation to use AI to cut corners is always there. One team I worked with spent weeks trying to generate podcast intros using a popular text-to-speech service. They had a great host with a really engaging voice, but they wanted to be able to generate new intros on the fly without having to bring him into the studio every time. The results were… not great. The AI-generated voice sounded flat and lifeless. It was like listening to a GPS navigator trying to tell a joke. We ended up scrapping the idea and going back to recording the intros with the actual host. It was more work, but it was worth it to preserve the authenticity of the show.
Behind the Curtain: Why AI Voices Fail the Turing Test
So why is it so hard to create a truly human-sounding AI voice? It comes down to a concept called prosody. Prosody is the rhythm, stress, and intonation of speech. It’s the music of language. It’s what gives our words meaning and emotion. And it’s incredibly difficult to teach a machine.
Think about all the different ways you can say the word “sure.” You can say it with a rising intonation to ask a question (“Are you sure?”). You can say it with a falling intonation to express certainty (“I’m sure.”). You can say it with a flat intonation to show you’re unenthusiastic (“Sure, whatever.”). An AI model has to learn all of those nuances from the data it’s trained on. And that’s a huge challenge.
When you look at a spectrogram of human speech, it’s a beautiful, chaotic mess. It’s full of peaks and valleys, sudden shifts in frequency, and subtle variations in timing. That’s the sound of a human being, with all of our imperfections and quirks. A spectrogram of AI-generated speech, on the other hand, is often too clean, too perfect. It’s like a perfectly straight line in a world of beautiful, messy curves.
This is why even the best AI voices still have that tell-tale robotic quality. They can get the words right, but they can’t get the music right. They’re like a musician who can play all the right notes but has no sense of rhythm or soul.
The Problem with “Perfect”
The other issue is that we’re chasing the wrong goal. Everyone’s trying to create a “perfect” AI voice, one that’s indistinguishable from a human. But what does that even mean? Human speech is messy. We stumble over our words, we use filler words like “um” and “uh,” we go on tangents. That’s what makes it authentic.
When you listen to a podcast, you’re not just listening for information. You’re listening for a connection. You want to feel like you’re in the room with the host, having a conversation. A perfectly polished, robotic voice can’t give you that. It creates a barrier between you and the listener.
Think about it. Would you rather listen to a podcast that’s a little rough around the edges but feels real, or one that’s perfectly produced but sounds like it was recorded by a machine? I know which one I’d choose.
This reminds me of the early days of computer-generated graphics in movies. There was a period where studios were so focused on creating “realistic” CGI that they lost sight of the art. The characters looked technically perfect, but they had dead eyes. They were in the uncanny valley. That’s where we are with AI voice right now. We’re so focused on the technical details that we’re forgetting about the soul.
Where We Go From Here
So, am I saying that AI voice is a dead end? Absolutely not. The technology is getting better every day. But we need to be realistic about where we are right now. And we need to be more thoughtful about how we use it.
Instead of trying to replace human voices, we should be looking for ways to augment them. Can we use AI to clean up audio, remove background noise, or even help with editing? Of course. But we shouldn’t be so quick to hand over the microphone entirely.
Here are a few ways I see AI voice being used effectively in the near future:
- Personalized audio at scale. Imagine a future where you can listen to an article from your favorite news source, read in the voice of your favorite narrator. Or a company that can create personalized audio messages for its customers, in a voice that’s consistent with its brand. I’m an investor in a company called Pika, which is doing some amazing things with AI-powered video generation. I can see a similar future for audio, where we can create highly personalized, dynamic content on the fly.
- Tools for creators. AI can be a powerful tool for podcasters and other audio creators. It can help with things like transcription, show notes, and even generating rough cuts of episodes. This frees up creators to focus on what they do best: creating great content. At RemoteTeam, we were always looking for ways to automate the tedious parts of our work so we could focus on the creative stuff. AI has the potential to do that for audio creators on a massive scale.
- Accessibility. For people with speech impairments, AI voice can be a life-changing technology. It can give them a voice, and allow them to communicate with the world in a way that was never before possible. This is where the technology has the potential to do the most good. And it’s an area where I’m actively looking to invest.
I’ve seen this movie before. With every new technology, there’s a period of inflated expectations. Then comes the trough of disillusionment, followed by a more realistic understanding of what the technology can and can’t do. We’re in the hype phase with AI voice right now. But the disillusionment is coming.
The Human Element
Ultimately, the success of AI voice will come down to one thing: the human element. We can have the most sophisticated technology in the world, but if it doesn’t connect with us on an emotional level, it’s just noise. The future of AI is not about replacing humans, but about augmenting our abilities and freeing us up to be more creative, more expressive, and more… well, human.
My advice? Don’t get caught up in the hype. Focus on creating great content and building a real connection with your audience. That’s the stuff that matters. And that’s the stuff that no AI, no matter how advanced, will ever be able to replicate.
I’m still incredibly bullish on the future of AI. I’ve put my money where my mouth is, with investments in some of the most promising AI companies in the world. But I’m also a realist. And the reality is that we’re still a long way from a world where AI voices are indistinguishable from human ones. And maybe that’s not such a bad thing. Maybe the imperfections are the point. Maybe the uncanny valley is a reminder that there’s still something special, something irreplaceable, about the human voice.
Frequently Asked Questions
What experience informs this perspective?
This perspective comes from over a decade of building companies in Silicon Valley, two successful exits (RemoteTeam to Gusto, MovieLaLa to Gfycat), and investing in 200+ startups including Anthropic, OpenAI, and Scale AI. I write about what I've lived.
How has this view evolved over time?
My thinking on most topics has changed significantly over the years. Early in my career, I held many conventional views that experience proved wrong. I try to update my beliefs when the evidence changes.
How can I apply this thinking to my own situation?
Start by identifying the core principle behind the opinion, not the specific example. Then ask yourself: does this principle apply to my context? If yes, test it in a small, low-risk way before going all in.