I almost laughed my co-founder out of the room. It was 2018, and he had just pitched me on integrating a new text-to-speech feature into our platform, RemoteTeam. “You want our users to listen to that?” I asked, pulling up a demo of the leading text-to-speech (TTS) service at the time. The voice that emanated from my speakers was flat, robotic, and utterly devoid of personality. It sounded like a hostage reading a script with a gun to its head. The idea of using that to read important notifications to our users seemed absurd, almost insulting.
We were trying to build a product that fostered human connection for remote teams, and this felt like the exact opposite. It was cold, impersonal, and frankly, a little creepy. I vetoed the idea on the spot. It was a hard no.
Fast forward eight years, and I’ve personally angel-invested over $5 million across 14 different startups in the AI voice space. Some of them are now household names, pioneers in a field I once dismissed. So what changed? The technology didn't just get better; it crossed a threshold. It went from a robotic novelty to something almost indistinguishable from human speech, and in some cases, even better. It’s a technology that has found its soul.
This is the story of my journey from a staunch non-believer to one of the biggest advocates for AI voice, and why I’m convinced we’re at the very beginning of a seismic shift in how we interact with every piece of technology in our lives.
The Early Days: A Robotic Wasteland
My early skepticism wasn't unfounded. For years, text-to-speech was a technological backwater. The voices were synthesized using a method called concatenative synthesis, which essentially involved stitching together tiny, pre-recorded snippets of speech from a single voice actor. The result was a choppy, disjointed mess. You could always hear the seams. The cadence was unnatural, the emphasis was often wrong, and there was zero emotional range. Think of the default voice on your old GPS, the one that would pronounce “St.” as “Street” in one context and “Saint” in another, often with hilarious, and sometimes disastrous, results.
But even in that robotic wasteland, the promise was there. As an entrepreneur, you learn to look for the underlying human need a technology is trying to meet. And the need for more accessible information was, and still is, massive. For the 2.2 billion people globally with vision impairment, or the 700 million with dyslexia, TTS wasn’t a novelty; it was a lifeline. It was the key to unlocking a world of written content that was previously locked away. I saw the 10x potential, but the technology itself was maybe at a 1.5x level of quality. I passed on investing in three early-stage TTS companies between 2017 and 2019. In hindsight, that was a multi-million dollar mistake for each of them.
The Turning Point: When AI Found Its Voice
For me, the “holy shit” moment arrived in late 2022. I was judging a pitch competition at Stanford, and a team of PhD students presented a demo that I initially thought was a prank. They had a text-to-speech model that was generating audio in real-time, and it sounded… human. Not just human-like, but genuinely human. It had pauses, inflections, and even the subtle, non-verbal cues that we use to convey emotion. They fed it a line from a sad movie, and the AI’s voice cracked with sorrow. They gave it a joke, and it delivered the punchline with perfect comedic timing.
They explained that they had abandoned the old concatenative methods. Instead, they were using a generative adversarial network, or GAN. Two neural networks were essentially locked in a battle. One network, the “generator,” would create synthetic speech from text. The other, the “discriminator,” would try to determine if the speech was real or fake by comparing it to a massive dataset of human speech. The generator’s goal was to fool the discriminator. Over millions of cycles, the generator got so good at mimicking human speech that the discriminator could no longer tell the difference. The result was audio that was startlingly realistic.
I cornered the founders after their presentation. My first question wasn’t about their business model, it was “How?” We spent an hour geeking out on the architecture of their model. I invested $250,000 on the spot, my first check into the new wave of AI voice. That company was acquired by a major tech giant 18 months later, returning 30x my initial investment.
The Rise of AI Voice Cloning: A Double-Edged Sword
That breakthrough opened the floodgates. The next frontier was not just generating a human voice, but generating any human voice. This is the world of AI voice cloning, and it’s one of the most exciting and terrifying technologies I’ve ever encountered.
I recently worked with a startup that can create a perfect digital clone of your voice from just 30 seconds of audio. I tried it myself. I uploaded a clip from a podcast I was on, and within a minute, I was listening to my own voice saying things I had never said. The fidelity was flawless. It had my exact accent, my pacing, even the way I tend to slightly trail off at the end of a sentence. It was uncanny.
The positive applications are mind-boggling. Think about a Hollywood studio being able to dub a movie into any language using the original actor’s voice. Or a video game where every character has a unique, dynamically generated voice. I’ve invested in a company that’s using voice cloning to create personalized audiobooks for children, where the child is the main character and the story is read in their parent’s voice. The emotional connection is incredibly powerful.
But the dark side is just as potent. The potential for misuse gives me pause. We’re already seeing deepfake audio being used to scam people, spread misinformation, and create fake endorsements from celebrities. The FTC even launched a challenge to spur the development of technology to detect AI-generated voices. It’s a classic cat-and-mouse game. As the fakes get better, so must our ability to detect them.
My stance on this is one of cautious optimism. You can’t stop technological progress, but you can build guardrails. The companies I invest in are all committed to ethical development, including features like watermarking to identify synthetic audio. It’s not a perfect solution, but it’s a start.
AI in Podcasting: My Secret Weapon
As someone who hosts a podcast, the impact of AI voice is not theoretical; it’s a part of my weekly workflow. Editing a podcast used to be a tedious, time-consuming process. Now, I use an AI tool that can automatically remove filler words, awkward pauses, and even background noise. It saves my audio engineer at least 5 hours of work per episode.
But the real revolution is in content creation. I’ve started using AI to generate entire segments of my show. For example, I recently did an episode on the history of venture capital. I fed a script into an AI voice generator, and it created a 10-minute audio documentary, complete with a historical-sounding narrator and background music. The feedback from my listeners was overwhelmingly positive. Not a single person suspected it was an AI.
This technology democratizes content creation. You no longer need a professional recording studio or a “radio voice” to create a high-quality podcast. You just need good ideas and a compelling script. It’s a massive unlock for creators worldwide. I’m also experimenting with translating my podcast into Spanish and Mandarin using a cloned version of my own voice. The potential to reach a global audience is immense.
The Future is Vocal, and It’s Here Now
We are at the dawn of the voice-first era. The clumsy, robotic voices of the past are being replaced by a symphony of AI-generated speech that is rich, diverse, and deeply human. The way we interact with our devices, our cars, and our homes is about to fundamentally change. Keyboards and screens will become secondary, and our voice will become the primary interface.
This isn’t some far-off sci-fi future. It’s happening right now. The AI voice revolution is not coming; it’s here. And for me, the entrepreneur who once laughed at the idea of a talking computer, it’s a powerful reminder that the most disruptive ideas are often the ones that sound the most ridiculous at first. The key is to listen closely. You might just hear the future speaking to you.
I remember one particularly painful board meeting where we were reviewing a product demo that used a TTS voice to guide new users through the onboarding process. The voice mispronounced our company name—not just once, but every single time. It was excruciating. We were trying to project an image of a polished, professional company, and our own product was butchering our name. The investors were not impressed. That was the day I declared a moratorium on all TTS projects until the technology was at least 50% better.
My skepticism was a running joke in my investment circle. I was the “anti-voice guy.” While my friends were pouring money into early voice assistant startups, I was investing in things I could see and touch, like SaaS platforms and developer tools. I just couldn’t get past the sheer cringe factor of the technology. It felt like a solution in search of a problem, a gimmick that would never achieve mainstream adoption.
The Technical Deep Dive: From Concatenation to Generation
To truly appreciate the leap that AI voice has taken, you have to understand the technology under the hood. The old concatenative synthesis was, in essence, a digital puppet show. You had a library of sounds, and you were just stringing them together. It was a brute-force approach, and it was incredibly limited. You could never create a voice that sounded truly natural because you were always working with a finite set of pre-recorded sounds.
The new generative models are a completely different beast. They don’t just play back sounds; they create them. They learn the underlying patterns of human speech—the rhythm, the pitch, the intonation—and they use that knowledge to generate entirely new audio. It’s the difference between a music box and a concert pianist. One is just playing a pre-programmed tune; the other is creating a unique and expressive performance.
The GAN architecture I mentioned earlier is just one example. There are other approaches, like variational autoencoders (VAEs) and diffusion models, that are also producing incredible results. The pace of innovation is staggering. Every few months, a new paper is published that pushes the state of the art even further. We’re in a golden age of AI voice research, and the results are speaking for themselves.
Beyond Podcasting: The Enterprise Revolution
While podcasting is a fun and exciting application of AI voice, the real money is in the enterprise. I’m talking about using AI voice to transform customer service, sales, and internal training. I’ve invested in a company that has developed an AI-powered call center agent that can handle over 80% of inbound customer queries. It can understand complex questions, access customer data in real-time, and provide personalized responses in a natural, conversational tone.
Think about the implications. You can provide 24/7 customer support in any language, without having to hire a massive team of human agents. You can scale your sales team by using AI to qualify leads and schedule appointments. You can create interactive training modules that are more engaging and effective than a traditional PowerPoint presentation.
This is not about replacing humans; it’s about augmenting them. It’s about freeing up human agents to focus on the most complex and high-value interactions. It’s about creating a better experience for both customers and employees. The ROI is massive. The company I mentioned earlier is already saving its clients an average of $2 million a year in customer service costs.
The Final Frontier: Emotional Intelligence
The next big challenge for AI voice is emotional intelligence. It’s not enough for an AI to just understand the words you’re saying; it needs to understand the emotion behind them. Is the customer angry? Frustrated? Delighted? Being able to detect and respond to these emotional cues is the key to creating truly human-like interactions.
This is an incredibly difficult problem, but it’s one that I’m confident we will solve. There are already startups that are using sentiment analysis and other techniques to detect emotion in speech. The results are still early, but they’re promising. I believe that within the next five years, we will have AI voice assistants that can not only understand what you’re saying, but also how you’re feeling.
When that happens, it will be a game-changer. It will usher in an era of truly personalized and empathetic technology. It will be the final step in the journey from a robotic wasteland to a world where technology speaks our language, both literally and figuratively. And I, for one, am incredibly excited to be a part of it.
Frequently Asked Questions
How has this view evolved over time?
My thinking on most topics has changed significantly over the years. Early in my career, I held many conventional views that experience proved wrong. I try to update my beliefs when the evidence changes.
Do all experts agree with this view?
No, and that's fine. The best ideas in business are often contrarian. I share my perspective based on my experience and data, but I encourage you to seek out opposing viewpoints and form your own conclusions.
How can I apply this thinking to my own situation?
Start by identifying the core principle behind the opinion, not the specific example. Then ask yourself: does this principle apply to my context? If yes, test it in a small, low-risk way before going all in.
What experience informs this perspective?
This perspective comes from over a decade of building companies in Silicon Valley, two successful exits (RemoteTeam to Gusto, MovieLaLa to Gfycat), and investing in 200+ startups including Anthropic, OpenAI, and Scale AI. I write about what I've lived.