It was 2018, and the AI world was buzzing. We were all obsessed with GANs, reinforcement learning, and the endless possibilities of deep learning. At RemoteTeam, my second startup, we were building tools to help remote teams collaborate better. We were a small, scrappy team, and we were always looking for an edge.
One afternoon, I stumbled upon a research paper on voice cloning. The idea was simple: feed a neural network a few seconds of someone’s voice, and it could generate new audio that sounded just like them. My mind immediately started racing. I thought about all the ways we could use this. Personalized onboarding for new hires? Automated sales calls that sounded like me? The possibilities were endless.
I had to try it. I found an open-source implementation on GitHub and spent the next few hours tinkering with it. The setup was a bit clunky, but I finally got it working. I recorded a few sentences of my own voice and fed them into the model. The first few attempts were… not great. The voice was robotic and distorted. But I kept tweaking the parameters, and after a few more hours, I had something that sounded surprisingly like me.
Now for the real test. My co-founder, Alex, was working late that night. I used the cloned voice to generate a short audio clip: “Hey Alex, can you push the latest build to production? I just found a critical bug.” I sent it to him on Slack and waited.
A few seconds later, he replied: “On it.”
I couldn’t believe it. It worked. I had fooled my own co-founder with a cloned version of my voice. I immediately called him and confessed. He was shocked, and a little freaked out, but he also saw the potential. We spent the rest of the night brainstorming all the ways we could use this technology.
That was the moment I realized that AI voice was going to be a huge deal. It wasn’t just a cool party trick; it was a technology that could fundamentally change the way we interact with computers and with each other.
The Cambrian Explosion of AI Voice
Fast forward to today, and the AI voice landscape is unrecognizable. The technology has advanced at an exponential rate. What took me hours to set up in 2018 can now be done in a few seconds with a web app. The quality of the cloned voices is also much better. It’s not just about mimicking a voice; it’s about capturing the nuances of human speech – the intonation, the emotion, the rhythm.
We’re seeing a Cambrian explosion of AI voice applications. Here are just a few of the areas where this technology is making a huge impact:
- Content Creation: AI voice generators are making it easier than ever to create high-quality audio content. Podcasters can use it to create professional-sounding intros and outros. YouTubers can use it to narrate their videos. And authors can use it to create audiobooks.
- Customer Service: AI-powered voice assistants are transforming the customer service industry. They can answer customer questions, resolve issues, and even provide emotional support. This is freeing up human agents to focus on more complex and creative tasks.
- Healthcare: AI voice is being used to help people with speech impediments communicate more easily. It’s also being used to create personalized voice assistants for the elderly and people with disabilities.
- Entertainment: The entertainment industry is using AI voice to create more realistic and immersive experiences. Video game characters can now have unique voices that adapt to the player’s actions. And movies can be dubbed into different languages with a level of quality that was previously impossible.
The Dark Side of AI Voice
Of course, with any powerful technology, there’s a dark side. The same technology that can be used to create amazing new experiences can also be used for malicious purposes. We’ve already seen examples of AI voice being used to create deepfake audio of politicians and celebrities. And it’s only a matter of time before we see it being used for more sophisticated scams and phishing attacks.
This is something that I’ve thought a lot about. As an entrepreneur and an investor, I’m always looking for new technologies that can change the world. But I also have a responsibility to think about the potential negative consequences of those technologies.
That’s why I’m a big believer in responsible innovation. We need to build safeguards into these systems to prevent them from being misused. We need to educate the public about the dangers of deepfake audio. And we need to work with policymakers to create a regulatory framework that encourages innovation while protecting people from harm.
The Future is Vocal
Despite the risks, I’m incredibly optimistic about the future of AI voice. I believe that we’re on the cusp of a new era of human-computer interaction. For decades, we’ve been forced to communicate with computers on their terms, using keyboards and mice. But with AI voice, we can finally communicate with computers in the most natural way possible: by talking to them.
This is going to have a profound impact on every aspect of our lives. It’s going to change the way we work, the way we learn, and the way we play. It’s going to make technology more accessible to everyone, regardless of their age or ability.
I’ve been fortunate enough to have a front-row seat to the AI revolution. I’ve seen how this technology can be used to solve some of the world’s most pressing problems. And I’m more convinced than ever that AI voice is one of the most important technologies of our time.
The future is vocal. And I can’t wait to see what we build with it.
The Nitty-Gritty: How Does This Stuff Actually Work?
I get asked this a lot. People are fascinated by the idea of cloning a voice, but they often assume it’s some kind of black magic. The reality is a bit more down-to-earth, but no less impressive. It all comes down to a type of AI called a neural network, which is loosely modeled on the human brain.
Think of it like this: you show a child a thousand pictures of a cat, and eventually, they learn to recognize a cat. A neural network does something similar, but with audio. You feed it hours and hours of audio data, and it starts to learn the patterns and characteristics of human speech. It learns the difference between a hard “t” and a soft “t,” the way a question rises in pitch at the end, and the subtle pauses and hesitations that make speech sound natural.
To clone a specific voice, you use a technique called “transfer learning.” You take a pre-trained neural network that already knows a lot about human speech, and then you fine-tune it on a small sample of the target voice. This is why you only need a few seconds of audio to clone a voice. The network already has a solid foundation, and it just needs a little bit of extra data to adapt to a new voice.
Of course, there’s a lot more to it than that. There are different types of neural networks, like recurrent neural networks (RNNs) and generative adversarial networks (GANs), each with its own strengths and weaknesses. And there’s a whole field of research dedicated to making these models more efficient and more accurate. But at its core, that’s how it works. It’s not magic; it’s just a lot of data and a lot of math.
Where I'm Placing My Bets: Investing in the Future of Voice
As an investor, I’m always looking for companies that are not just building cool technology, but are also solving real-world problems. And in the world of AI voice, there’s no shortage of them. I’ve had the privilege of investing in some of the most innovative companies in this space, and I’m incredibly excited about what they’re building.
One of my early investments was in a company called Descript. They started out as a simple tool for transcribing audio, but they’ve since evolved into a full-fledged audio and video editing platform. One of their most impressive features is called “Overdub,” which allows you to create a text-to-speech model of your own voice. This means you can edit audio by simply editing the text. If you make a mistake in a recording, you don’t have to re-record the whole thing. You can just type in the correction, and Descript will generate the audio in your own voice. It’s a game-changer for anyone who works with audio.
Another company I’m excited about is ElevenLabs. They’re building a platform for creating high-quality, natural-sounding AI voices. They have a library of pre-made voices, but you can also use their technology to clone your own voice. The quality of their voices is truly remarkable. They’re able to capture the subtle nuances of human speech in a way that I’ve never heard before. I believe they have the potential to become the “Stripe for voice,” providing the underlying infrastructure for a whole new generation of voice-powered applications.
These are just a couple of examples, but they illustrate a broader point. The AI voice industry is not just about creating deepfakes and party tricks. It’s about building tools that can help us communicate more effectively, create more engaging content, and make technology more accessible to everyone. And that’s something I’m proud to invest in.
The Next Frontier: Where Do We Go From Here?
So, what’s next for AI voice? If the last few years are any indication, the pace of innovation is only going to accelerate. Here are a few of the things I’m most excited about:
- Real-time voice conversion: Imagine being able to speak in any voice you want, in real-time. You could be on a video call and sound like a celebrity, or you could be giving a presentation and have your voice translated into another language on the fly. This technology is still in its early stages, but it has the potential to be truly transformative.
- Emotional AI: One of the biggest challenges in AI voice is capturing the emotional content of speech. But as these models become more sophisticated, they’re starting to get better at understanding and generating emotion. This will open up a whole new range of applications, from more empathetic customer service bots to more engaging and immersive storytelling experiences.
- Personalized voice assistants: In the future, I believe we’ll all have our own personalized voice assistants. These assistants will know our preferences, our habits, and our goals. They’ll be able to help us with everything from scheduling appointments to managing our finances. And they’ll communicate with us in a voice that is tailored to our own personality and preferences.
Of course, there are still many challenges to overcome. We need to make these models more efficient, more robust, and more secure. We need to develop better ways to detect and prevent malicious use of this technology. And we need to have a broader public conversation about the ethical implications of AI voice.
But I’m confident that we can overcome these challenges. The potential benefits of this technology are simply too great to ignore. AI voice has the power to change the world, and I’m thrilled to be a part of it. The journey that started with a late-night experiment in 2018 is far from over. In fact, it’s just beginning.
Frequently Asked Questions
Do all experts agree with this view?
No, and that's fine. The best ideas in business are often contrarian. I share my perspective based on my experience and data, but I encourage you to seek out opposing viewpoints and form your own conclusions.
What's the most common pushback you get on this?
People often push back by citing exceptions or edge cases. And they're usually right that exceptions exist. But building a strategy around exceptions rather than patterns is a losing game for most founders.
How can I apply this thinking to my own situation?
Start by identifying the core principle behind the opinion, not the specific example. Then ask yourself: does this principle apply to my context? If yes, test it in a small, low-risk way before going all in.
How has this view evolved over time?
My thinking on most topics has changed significantly over the years. Early in my career, I held many conventional views that experience proved wrong. I try to update my beliefs when the evidence changes.