We paid a team of hackers to destroy our AI. They succeeded.
It was one of the most humbling and valuable experiences of my career. For a few days, we invited some of the smartest people I know—people who think in ways that are completely alien to most engineers—to find the cracks in our systems. And they found them. Not just small cracks, but gaping holes that made my stomach turn.
This wasn’t a theoretical exercise. This was a full-contact, no-holds-barred assault on a system we had spent years building. I’m sharing the raw, unfiltered story because the single biggest threat to the future of AI isn’t the technology itself, but our own arrogance in believing we can control it without this kind of rigorous, adversarial testing.
Why We Did It
Let’s be honest. Most companies approach AI security as a compliance checkbox. They run some automated scans, patch the obvious stuff, and call it a day. That’s not nearly enough. We’re building models that are increasingly autonomous, and the attack surfaces are things we can’t even predict. The EU AI Act is coming, and it has teeth. But even if it didn’t, this is just the right thing to do.
I’ve been fortunate to invest in companies like Anthropic and OpenAI, and I’ve seen firsthand how seriously they take alignment. But for the rest of us, the thousands of companies building on top of these foundational models, we have a responsibility to be just as vigilant. We decided to put our money where our mouth is. We hired a red team.
The Setup: Inviting the Wolves
The team we brought in wasn’t your typical pentesting firm. These were specialists in adversarial machine learning. One was a former NSA operator, another a PhD who specializes in data poisoning attacks, and the third was a 22-year-old who had won one of the biggest bug bounties in history. We gave them a simple mandate: “Break it. Do whatever it takes.”
We set up a sandboxed environment with our latest models and gave them full access to our code, our documentation, and our engineers. The only thing we didn’t give them was a lot of time. They had one week.
The Attack: A Masterclass in Deception
For the first two days, there was silence. I was starting to get a little cocky. Maybe our defenses were better than I thought.
Then, on Wednesday morning, it all went to hell.
The first attack was a classic prompt injection, but with a twist. They figured out a way to embed commands in a seemingly innocuous user query that tricked our model into revealing its own system prompts and internal instructions. It was like a magician revealing their trick. Suddenly, the black box was wide open.
But that was just the warm-up. The really scary part came next. The team used a sophisticated data poisoning technique. They created a series of inputs that looked like normal data but were subtly engineered to corrupt the model’s training. Over a few hours, they were able to create a backdoor. They could feed the model a specific, secret phrase, and it would bypass all of our safety filters, generating harmful, biased, and completely unrestricted content.
They owned it. We had lost control.
The Aftermath: Humility and Hard Work
Seeing that backdoor in action was a punch to the gut. It felt like a personal failure. We had spent so much time on this, and a small team of experts unraveled it in a few days.
But after the initial shock wore off, I felt a sense of clarity. We now had a roadmap. We knew where the real vulnerabilities were, not just the theoretical ones.
The red team didn’t just break things; they showed us how to fix them. We’re now implementing a multi-layered defense system. We’re building new classifiers to detect adversarial inputs, we’re completely overhauling our data sanitization process, and we’re making red-teaming a permanent, ongoing part of our development cycle, not a one-time event.
This Is the Way
Building responsible AI is not about writing a pretty ethics statement. It’s about embracing a culture of paranoia. It’s about having the humility to admit that you don’t know what you don’t know. It’s about inviting the smartest people you can find to challenge your assumptions and break your creations.
If you’re building AI, you need to do this. Don’t wait for a regulator to force your hand. Don’t wait for a real-world incident to expose your weaknesses. Hire a red team. It will be terrifying. It will be expensive. And it will be the best investment you ever make.
Frequently Asked Questions
What's the most common pushback you get on this?
People often push back by citing exceptions or edge cases. And they're usually right that exceptions exist. But building a strategy around exceptions rather than patterns is a losing game for most founders.
What experience informs this perspective?
This perspective comes from over a decade of building companies in Silicon Valley, two successful exits (RemoteTeam to Gusto, MovieLaLa to Gfycat), and investing in 200+ startups including Anthropic, OpenAI, and Scale AI. I write about what I've lived.
Do all experts agree with this view?
No, and that's fine. The best ideas in business are often contrarian. I share my perspective based on my experience and data, but I encourage you to seek out opposing viewpoints and form your own conclusions.