I still remember the sweat beading on my forehead. We were in a stuffy boardroom in Mountain View, pitching MovieLaLa to a VC firm that could make or break us. Our big idea? A “Siri for movies.” You’d talk to your phone, and it would find you the perfect movie to watch. The demo was supposed to be the knockout punch.
I held up the phone and said, with as much confidence as I could muster, “Show me action movies with Tom Cruise.”
The app whirred for a second and then proudly displayed… a list of romantic comedies starring Tom Hanks.
I tried again. “Action. Movies. Tom. Cruise.”
It gave me a documentary about cruise ships.
The room was silent. The VCs exchanged glances. We did not get the funding.
That little disaster cost us more than just the deal. We had burned through nearly $50,000 on the best speech recognition API available at the time, and it was utterly, comically useless. I honestly had no idea what I was doing, other than lighting money on fire. Back then, which feels like a century ago in tech years, voice recognition was a gimmick. A clunky, frustrating, expensive gimmick.
Fast forward to last week. I called my bank. A voice answered immediately—no menu, no “press one for English.” It sounded like a friendly, competent human. I asked a complicated question about a wire transfer, it understood the nuance, and gave me a perfect answer. It even had a slight, reassuring cadence. I was so impressed that I asked, “Are you a real person?” It replied, cheerfully, “I’m a virtual assistant powered by AI! Pretty cool, right?”
Cool doesn’t even begin to cover it. We’ve gone from Tom Hanks rom-coms to flawless, human-like conversational AI in just a few years. The technology that humiliated me in that boardroom is now one of the most exciting and disruptive fields I’m seeing in my 200+ angel investments. It’s why I’ve put money into companies at the forefront of this, like Anthropic and OpenAI. They aren’t just building better voice apps; they’re fundamentally changing how we interact with technology.
What Changed? It’s Not Just About Hearing Better
So, what happened? How did we get from there to here? The change wasn’t incremental. It was a complete reinvention of the underlying technology.
The old systems were based on phonetics. They broke words down into tiny sound components (phonemes) and tried to match them to a dictionary. It was a rigid, brittle process. Think of it like trying to understand a conversation by having someone spell out every single word for you. You miss the tone, the intent, the context. You just get the raw letters.
Today’s AI voice systems work on a completely different principle: deep learning. They don’t just listen for sounds; they learn the patterns, context, and structure of human language. They train on massive datasets of audio—thousands of hours of conversations, podcasts, and audiobooks. It’s the difference between learning a language from a phrasebook and growing up speaking it natively. The new way understands.
This allows for two incredible breakthroughs:
Speech Recognition That Actually Works: Modern AI can understand speech with near-human accuracy. It can handle different accents, background noise, and rapid speech. It’s not just transcribing words; it’s grasping intent. This is the foundation for everything that follows.
Generative Voice That Feels Real: This is the part that gives me goosebumps. AI can now generate speech that is indistinguishable from a human voice. It can replicate tone, emotion, and individual vocal tics. This isn’t the robotic “Stephen Hawking” voice of the past. This is a voice that can tell a story, express empathy, or sell you a product.
This Isn’t a Fad. It’s a Platform Shift.
I get it. In Silicon Valley, we love to call everything the “next big thing.” But this is different. When a core human function like speech becomes a programmable, scalable technology, it changes everything. It’s not just about building better customer service bots.
Look, I’ve seen a few platform shifts in my time. The move from desktop to web with RemoteTeam, and from web to mobile with MovieLaLa. This feels just as big. Voice is becoming a primary interface. We’re moving from a world of tapping on glass to a world of conversations.
Here’s where I see the most immediate impact:
Hyper-Personalized Media: Imagine a podcast that’s generated in real-time, just for you, in your favorite narrator’s voice. Or an audiobook where you can choose the voice actor. We’re already seeing the start of this, and it’s going to completely upend the content industry. For more on this, you can check out my thoughts on the future of AI in content creation.
Truly Smart Assistants: The dream of the conversational computer from Star Trek is finally within reach. An assistant that doesn’t just follow commands but understands context, anticipates your needs, and communicates like a partner. This will be built into our phones, our cars, our homes.
The End of Language Barriers: Real-time voice translation is already here, but it’s about to get a lot better. Imagine having a seamless conversation with someone in another language, with the AI translating in real-time, in a voice that sounds like your own. This has massive implications for global business and human connection.
Of course, this technology is not without its risks. The potential for misuse in creating deepfakes and spreading misinformation is terrifying. We need to be having a serious conversation about ethics and safeguards right now. I’m not an expert on policy, but as a builder and investor, I know we can’t afford to ignore the downside. We need robust verification systems and clear regulations. If you're interested in the security side of things, I wrote about some of my early-stage security investments in my post on building a digital fortress.
The Takeaway: Don’t Get Left Behind
I made a mistake all those years ago. I saw the promise of voice but underestimated how hard the problem was. I was too early, and I paid the price. The opposite mistake is being made today. People are seeing the technology, but they’re thinking too small. They see a better Siri, not a new computing paradigm.
My advice is simple. Start paying attention. Think about how conversational AI could change your business, your industry, your job. Don’t just think about how to improve your existing processes; think about what entirely new things become possible when you can talk to computers and they can talk back.
Because this time, it’s not a demo in a stuffy boardroom. It’s real. And it’s happening now.
Frequently Asked Questions
How can I apply this thinking to my own situation?
Start by identifying the core principle behind the opinion, not the specific example. Then ask yourself: does this principle apply to my context? If yes, test it in a small, low-risk way before going all in.
What's the most common pushback you get on this?
People often push back by citing exceptions or edge cases. And they're usually right that exceptions exist. But building a strategy around exceptions rather than patterns is a losing game for most founders.
How has this view evolved over time?
My thinking on most topics has changed significantly over the years. Early in my career, I held many conventional views that experience proved wrong. I try to update my beliefs when the evidence changes.
What experience informs this perspective?
This perspective comes from over a decade of building companies in Silicon Valley, two successful exits (RemoteTeam to Gusto, MovieLaLa to Gfycat), and investing in 200+ startups including Anthropic, OpenAI, and Scale AI. I write about what I've lived.