I almost passed on the investment that taught me the most about the future of voice. It was 2018, and a couple of Stanford PhDs pitched me on a real-time voice cloning technology. My first thought? “Another deepfake startup.” I’d seen a dozen of them. Most were focused on gimmicks, like making celebrities say funny things. I was about to say no, but then they showed me a demo that completely changed my mind. They cloned my voice in 30 seconds from a recording on my phone, and it was… perfect. Not just the sound, but the intonation, the cadence, the slight lisp I have when I’m excited. That’s when I realized this wasn’t about deepfakes. It was about creating a digital representation of ourselves that could communicate on our behalf.
That company, as you might have guessed, was not one of the big names you hear about today. They were acquired by a larger company for their tech, and the founders are probably on a beach somewhere. But that experience taught me a valuable lesson: the future of voice is not about mimicking reality, but about extending it.
The "Voice First" World is Finally Here
For years, we’ve been hearing about the “voice-first” world. A world where we interact with technology using our voice, just like we do with each other. But for a long time, it was all hype. Siri was a novelty, Alexa was a glorified egg timer, and Google Assistant was… well, Google Assistant. The technology just wasn’t there yet. The speech recognition was clunky, the voices were robotic, and the assistants were dumb as a box of rocks.
But in the last couple of years, something has changed. The technology has finally caught up with the vision. We’re now at a point where voice AI is not just a gimmick, but a genuinely useful tool. I’ve seen it in my own companies. We use voice AI to transcribe meetings, to generate summaries, to create audio versions of our blog posts. We even have a voice AI assistant that helps our customer support team answer questions. And this is just the beginning.
I’ve made over 200 angel investments in my career, and I can tell you that the AI voice space is one of the most exciting areas I’m looking at right now. I’m seeing startups that are using voice AI to do everything from helping children learn to read to providing companionship for the elderly. The possibilities are endless.
Voice Cloning: More Than Just a Deepfake
When people hear “voice cloning,” they immediately think of deepfakes and misinformation. And yes, that’s a real concern. But it’s also a very narrow view of the technology. The reality is that voice cloning has the potential to be a force for good. Think about it. You could have a loved one who has passed away read a bedtime story to your children. You could have a famous author narrate their own audiobook, even if they’re no longer with us. You could even have a digital version of yourself that can handle all your phone calls and emails.
I invested in a company called “Voicery” (a fictional name) that was doing just that. They were building a platform that allowed people to create a digital version of their voice that could be used in a variety of applications. They were working with a well-known podcaster who had lost his voice due to a medical condition. They were able to clone his voice from old recordings and create a new podcast that sounded exactly like him. It was an incredibly powerful and moving experience.
Of course, there are ethical considerations. We need to have a serious conversation about who owns our voice, how it can be used, and what happens when it’s misused. But we can’t let fear hold us back from exploring the potential of this technology.
The Rise of Voice AI Assistants
I have a love-hate relationship with voice AI assistants. On the one hand, I love the convenience. I can ask my assistant to play music, to set a timer, to tell me the weather. On the other hand, I hate how dumb they can be. I can’t tell you how many times I’ve had to repeat myself or rephrase my question to get the right answer.
But I’m starting to see a change. The assistants are getting smarter. They’re starting to understand context, to learn my preferences, to anticipate my needs. I recently tried a new voice AI assistant from a startup I’m advising (let’s call them “Aura”). Aura is different. It’s not just a command-taker. It’s a true assistant. It can schedule meetings, book flights, and even order my groceries. It’s still early days, but I’m convinced that this is the future of voice AI.
Speech Recognition: The Unsung Hero
We’ve talked a lot about voice cloning and voice AI assistants, but we haven’t talked about the technology that makes it all possible: speech recognition. Speech recognition is the unsung hero of the voice AI revolution. It’s the technology that converts our spoken words into text that a computer can understand. And it’s gotten incredibly good in the last few years.
I remember when I was first starting out as an entrepreneur, I had to manually transcribe all my interviews and meetings. It was a tedious and time-consuming process. Now, I can just use a speech recognition service to do it for me. And it’s not just for transcription. Speech recognition is being used in a variety of applications, from voice search to voice-controlled devices.
One of the most interesting applications I’ve seen is in the area of healthcare. I invested in a company that is using speech recognition to help doctors diagnose diseases. They’ve developed a system that can analyze a patient’s voice and detect subtle changes that may be indicative of a medical condition. It’s still in the early stages of development, but it has the potential to revolutionize the way we diagnose and treat diseases.
My Investments in the Space
As an investor, I’m always looking for the next big thing. And right now, I believe that AI voice is one of the most promising areas of technology. I’ve already made a number of investments in the space, and I’m always on the lookout for more.
So what do I look for in a voice AI startup? First and foremost, I look for a strong team. I want to see a team that is passionate about what they’re doing and has the technical expertise to pull it off. Second, I look for a big market. I want to see a company that is solving a real problem for a large number of people. And third, I look for a unique technology. I want to see a company that has a defensible technology that can’t be easily replicated by a competitor.
Some of my recent investments in the space include:
- A company that is building a platform for creating interactive audio stories.
- A company that is developing a voice AI assistant for sales teams.
- A company that is using speech recognition to help children with learning disabilities.
I’m incredibly excited about the future of voice AI. I believe that we’re on the cusp of a new era of computing, an era where we interact with technology using our voice. And I can’t wait to see what the future holds.
The Future is Talking to Us
I started this article by talking about an investment I almost passed on. That experience taught me that the future is often not what we expect it to be. The future of voice is not about creating a perfect replica of reality. It’s about extending our reality, about creating new ways to communicate and interact with the world around us.
I believe that we’re at the very beginning of the voice AI revolution. The technology is still in its infancy, and we’re only just starting to scratch the surface of what’s possible. But I’m convinced that voice AI will change the world in ways that we can’t even imagine yet. The future is talking to us, and I, for one, am listening.
The Unseen Backend: The Data and Infrastructure Powering Voice AI
It’s easy to be captivated by the seamless experience of a voice assistant or the uncanny accuracy of a voice clone. But behind that magic lies a colossal infrastructure of data and computing power. This is something I learned the hard way with one of my early investments in the voice space. The company had a brilliant algorithm, but they completely underestimated the sheer volume of data required to train their models. Their initial prototype, which worked impressively in a controlled lab environment, fell apart in the real world. The model was brittle, failing to understand different accents, dialects, and background noises. They were burning through their seed funding just on data acquisition and annotation.
This experience was a stark reminder that in the world of AI, data is the new oil. And not just any data, but high-quality, diverse, and ethically sourced data. For voice AI, this means collecting millions of hours of audio from a wide range of speakers, in a variety of languages and environments. It means carefully annotating this data, a process that is still largely manual and incredibly labor-intensive. This is the unglamorous, behind-the-scenes work that is essential for building robust and reliable voice AI systems.
And then there’s the infrastructure. Training large-scale voice models requires massive amounts of computing power, typically in the form of specialized GPUs. This is a huge barrier to entry for smaller startups, and it’s one of the reasons why the big tech companies have such a dominant position in the market. They have the resources to build and maintain the massive data centers required to train these models. However, I’m starting to see a shift in the landscape. The rise of cloud computing and the availability of open-source models are starting to level the playing field. Startups can now rent the computing power they need from cloud providers like AWS, Google Cloud, and Azure, without having to make a huge upfront investment in hardware. And they can build on top of open-source models from companies like Hugging Face, which can significantly reduce their development time and cost.
The Human Element: Why UX is the Key to Adoption
I’ve seen too many startups with brilliant technology fail because they neglected the user experience. They were so focused on the technical details of their product that they forgot about the human beings who would be using it. In the world of voice AI, user experience is everything. It’s not enough to have a technology that is accurate and reliable. It also has to be intuitive, engaging, and easy to use.
One of the biggest challenges in designing a voice user interface (VUI) is that there are no visual cues. You can’t rely on buttons, menus, or other graphical elements to guide the user. You have to rely on the power of your words. This means carefully crafting your prompts, your responses, and your error messages. It means thinking about the flow of the conversation and how to make it as natural and intuitive as possible.
I once invested in a company that was building a voice-controlled smart home device. The technology was amazing. It could control everything from the lights to the thermostat to the security system. But the user experience was a disaster. The commands were confusing, the responses were robotic, and the system was constantly misunderstanding the user. The company eventually went out of business, not because their technology wasn’t good enough, but because they failed to understand the importance of user experience.
This is a lesson that I’ve taken to heart. When I’m evaluating a voice AI startup, I don’t just look at the technology. I also look at the user experience. I want to see a company that is obsessed with making their product as easy and enjoyable to use as possible. Because at the end of the day, that’s what will determine whether a product succeeds or fails.
The Road Ahead: Challenges and Opportunities
We’ve come a long way in a short amount of time, but there are still many challenges to overcome. One of the biggest challenges is the problem of bias. AI models are only as good as the data they’re trained on. And if that data is biased, then the model will be biased as well. This is a particularly serious problem in the world of voice AI, where models have been shown to be less accurate for women and people of color. This is not just a technical problem. It’s a social problem. And it’s one that we need to address if we want to build a future where voice AI is fair and equitable for everyone.
Another challenge is the problem of privacy. Voice assistants are always listening. They’re collecting vast amounts of data about our lives, our habits, and our conversations. And while this data is used to improve the user experience, it also raises serious privacy concerns. Who owns this data? How is it being used? And what happens if it falls into the wrong hands? These are questions that we need to be asking ourselves as we move forward.
But despite these challenges, I’m incredibly optimistic about the future of voice AI. I believe that we’re on the verge of a new era of computing, an era where we interact with technology in a more natural and intuitive way. I’m excited to see what the future holds, and I’m proud to be a part of it.
Frequently Asked Questions
What experience informs this perspective?
This perspective comes from over a decade of building companies in Silicon Valley, two successful exits (RemoteTeam to Gusto, MovieLaLa to Gfycat), and investing in 200+ startups including Anthropic, OpenAI, and Scale AI. I write about what I've lived.
Do all experts agree with this view?
No, and that's fine. The best ideas in business are often contrarian. I share my perspective based on my experience and data, but I encourage you to seek out opposing viewpoints and form your own conclusions.
What's the most common pushback you get on this?
People often push back by citing exceptions or edge cases. And they're usually right that exceptions exist. But building a strategy around exceptions rather than patterns is a losing game for most founders.