We paid a team of elite hackers to break our AI. They succeeded in ways we never imagined.
That’s not hyperbole. For a few weeks, we invited a red team—a group of ethical hackers—to attack our latest AI model. What they found was both terrifying and enlightening. It was a raw, unfiltered look at the vulnerabilities that can exist even when you think you’re doing everything right. This isn’t a story about a perfect process. It’s the story of our biggest security failure and the lessons we learned from it.
Why We Did It
Let's be honest, the term "AI safety" gets thrown around a lot. It’s easy to talk about building responsible AI, but it’s much harder to actually do it. We’ve always been committed to security, but as our AI systems became more complex, we knew we had to go beyond standard penetration testing. We needed to bring in people who thought differently, who could see the flaws that we, the creators, were blind to.
We were about to launch a new feature that used a sophisticated new AI model. The pressure was on. But a nagging thought kept me up at night: what if we were missing something? What if there was a backdoor, a vulnerability so subtle that we hadn’t even considered it? The potential for misuse, for our own creation to be turned against us, was a risk we couldn’t ignore. So, we decided to do something that felt counterintuitive: we hired a team to destroy our work.
Finding the Right People
Finding the right red team wasn't easy. We weren't looking for standard pentesters. We needed a team with a very specific skillset: a deep understanding of machine learning, a creative and adversarial mindset, and a proven track record of breaking complex systems. We interviewed a dozen firms, from big-name security consultancies to small, boutique outfits.
We eventually settled on a small, under-the-radar team of three. They weren't the most famous, but their approach was different. They didn't just talk about running automated scans. They talked about understanding the psychology of the system, of finding the blind spots in our own thinking. They were part hackers, part artists. They were perfect.
The rules of engagement were simple: no holds barred. They had full access to our code, our models, and our internal documentation. We wanted them to simulate a real-world attack from a determined, well-funded adversary. The only thing off-limits was our production data, for obvious reasons.
The First Cracks Appear
The first few days were quiet. Too quiet. The red team was heads-down, analyzing our architecture, our data pipelines, our model weights. We were on edge, waiting for the first shoe to drop. And then it did.
It started with a simple prompt injection attack. They found a way to bypass our input filters and feed the model a malicious prompt that made it reveal its own system prompt. It was a classic, almost textbook attack, but it was a stark reminder that even the most basic vulnerabilities can be overlooked. We had focused so much on the complex threats that we had neglected the simple ones.
But that was just the beginning. The red team then used that initial foothold to launch a more sophisticated attack. They found a way to poison our training data, to subtly manipulate the model's behavior over time. They could make it say things it wasn't supposed to say, to generate outputs that were biased or offensive. It was a slow, insidious attack, the kind that you wouldn't notice until it was too late.
The Deepfake Nightmare
The real “oh shit” moment came a week later. The red team had been quiet for a few days, and we were starting to think we had weathered the worst of the storm. We were wrong.
One of our security engineers received a video call from me. It was my face, my voice, asking for emergency access to a sensitive database. The request was unusual, but the video looked and sounded completely real. It was a deepfake, and it was flawless. The only reason we caught it was because the real me was in a meeting with the security team at that exact moment.
The red team had used a combination of publically available photos and audio of me to create a hyper-realistic deepfake. They had then used that deepfake to try and social engineer their way into our systems. It was a chilling demonstration of how easily our own identities could be weaponized. We had been so focused on protecting our AI from external threats that we hadn't considered the possibility that the threat could be wearing our own face.
The Aftermath and the Fixes
The red teaming exercise was a humbling experience. It exposed a number of critical vulnerabilities in our systems and forced us to confront some uncomfortable truths about our own security practices. But it was also one of the most valuable things we’ve ever done.
In the weeks that followed, we worked tirelessly to address the issues the red team had found. We implemented stricter input validation and sanitization. We developed new techniques for detecting and preventing data poisoning attacks. We completely overhauled our authentication and access control systems to be more resilient to social engineering and deepfake attacks.
But the changes weren't just technical. The exercise also led to a fundamental shift in our company culture. We now have a dedicated AI red team that is constantly trying to break our own systems. We've made AI safety a core part of our development process, not just an afterthought. And we're committed to being more transparent about our security practices, to sharing our learnings with the broader community.
Lessons for the Future
Building responsible AI is not a one-time task. It's an ongoing process of vigilance, of humility, and of constantly questioning your own assumptions. Here are a few of the key lessons we learned from our first AI red-teaming exercise:
- Think like an attacker. You can't defend against threats you don't understand. You need to bring in people who can think like your adversaries, who can see the vulnerabilities that you can't.
- Don't neglect the basics. It's easy to get caught up in the hype of complex, sophisticated attacks. But often, the most damaging vulnerabilities are the simplest ones. Don't forget to patch your systems, to validate your inputs, and to follow basic security best practices.
- The human element is your weakest link. The most sophisticated security systems in the world can be bypassed with a simple social engineering attack. You need to train your employees to be skeptical, to question unusual requests, and to be on the lookout for phishing and other social engineering tactics.
- Be prepared for the unknown. The threat landscape is constantly evolving. You need to be prepared for new and unexpected attacks. That means investing in research, staying up-to-date on the latest security trends, and building a culture of continuous learning and improvement.
Our first AI red-teaming exercise was a wake-up call. It was a stark reminder that building safe and responsible AI is not just a technical challenge, but a cultural one. It requires a commitment to transparency, a willingness to confront uncomfortable truths, and a healthy dose of paranoia. The path to truly responsible AI is a long and difficult one, but it's a path we must all walk together.
Frequently Asked Questions
What's the most common pushback you get on this?
People often push back by citing exceptions or edge cases. And they're usually right that exceptions exist. But building a strategy around exceptions rather than patterns is a losing game for most founders.
How has this view evolved over time?
My thinking on most topics has changed significantly over the years. Early in my career, I held many conventional views that experience proved wrong. I try to update my beliefs when the evidence changes.
What experience informs this perspective?
This perspective comes from over a decade of building companies in Silicon Valley, two successful exits (RemoteTeam to Gusto, MovieLaLa to Gfycat), and investing in 200+ startups including Anthropic, OpenAI, and Scale AI. I write about what I've lived.
Do all experts agree with this view?
No, and that's fine. The best ideas in business are often contrarian. I share my perspective based on my experience and data, but I encourage you to seek out opposing viewpoints and form your own conclusions.