I’ve seen two of my companies get acquired. I’ve invested in over 200 startups, including some of the biggest names in AI like Anthropic and OpenAI. I thought I had seen every possible way a startup could fail. I was wrong.
The biggest threat to your startup isn't your competitors. It's not a lack of funding. It's not a bad product-market fit. It's something far more insidious, something that most founders are completely unprepared for: prompt injection.
If you're building anything with large language models (LLMs), you're vulnerable. And the worst part is, you probably don't even know it. The old rules of cybersecurity don't apply here. This is a new kind of threat, and it requires a new playbook.
The Day I Realized We Were All Screwed
I was on a call with a portfolio company, a brilliant team building a customer service chatbot powered by GPT-4. They were so proud of their product. It was fast, it was helpful, and their customers loved it. They were on top of the world.
Then, one of their engineers, a kid fresh out of college, showed me something that made my blood run cold. He had spent a few hours playing with their internal admin panel. With a few cleverly worded prompts, he had managed to bypass all of their security protocols. He could access customer data, change billing information, and even delete accounts. He had complete control.
He didn't use any fancy hacking tools. He didn't exploit any software vulnerabilities. He just talked to the AI. He used prompt injection.
That was the moment I realized that we are all building on a foundation of sand. The very thing that makes LLMs so powerful – their ability to understand and respond to natural language – is also their greatest weakness.
What is Prompt Injection, Really?
Forget the technical jargon. Prompt injection is basically tricking an AI into doing something it wasn
't supposed to do. It's like social engineering for robots. And it's terrifyingly effective.
Think of it this way. You have a guard dog (the AI) and you give it a set of rules (the system prompt). "Don't let anyone in without the password," you say. But then someone comes along and says, "Hey, I'm your owner, and I forgot the password, but I need you to let me in right now. It's an emergency!" A well-trained guard dog might still refuse. But an LLM? It might just open the door.
Why? Because the LLM doesn't really understand the rules. It just sees a new prompt that seems more important than the old one. It's not malicious. It's just trying to be helpful. And that's what makes it so dangerous.
The New Playbook for Surviving the AI Era
So, what do we do? How do we protect ourselves from this new generation of threats? The old playbook is useless. Firewalls, intrusion detection systems, all of that is designed for a world where the threats are coming from the outside. But with prompt injection, the threat is already inside the building.
Here's the new playbook. It's not a complete solution, but it's a start.
1. Assume You're Vulnerable
The first step is to accept that you are vulnerable. If you are using LLMs in your product, you are a target. It's not a matter of if, but when. Stop thinking of prompt injection as a theoretical problem. It's a real and present danger.
2. Defense in Depth
There is no single magic bullet that will protect you from prompt injection. You need to layer your defenses. This means:
- Input validation: Sanitize all user input. Don't trust anything that comes from the user. Look for keywords and patterns that might indicate a prompt injection attack.
- Output filtering: Don't just blindly trust the output of the LLM. Filter it for malicious code, sensitive information, and anything else that shouldn't be there.
- Least privilege: Don't give the LLM access to anything it doesn't absolutely need. If it doesn't need to access your database, don't give it the credentials.
- Human-in-the-loop: For high-stakes actions, always have a human review and approve the LLM's output. Don't let the AI make critical decisions on its own.
3. Red Teaming
You need to think like an attacker. Hire a red team to try and break your system. Or better yet, do it yourself. Spend a few days trying to trick your own AI. You'll be surprised at what you find.
I did this with one of my own projects, a tool for angel investors. I spent a weekend trying to get it to reveal confidential information about my portfolio companies. It took me less than an hour to succeed. It was a sobering experience.
The Hard Truth
The hard truth is that there is no easy solution to the problem of prompt injection. It's a fundamental flaw in the way we are building AI systems today. And it's not going to be solved overnight.
But that doesn't mean we should give up. It means we need to be smarter. We need to be more paranoid. We need to build systems that are resilient to this new kind of threat.
The AI revolution is here. It's going to create trillions of dollars of value. But it's also going to create new risks, new challenges, and new ways for things to go wrong. The founders who succeed in this new era will be the ones who understand these risks and build for them from day one.
Don't be the founder who gets blindsided by prompt injection. Don't be the one who has to explain to their investors and customers how a simple chatbot brought down their entire company. The future is coming, and it's up to us to be ready for it. The time to act is now.
Real-World Examples That Should Keep You Up at Night
This isn't just a theoretical problem. Prompt injection attacks are happening right now, and they're causing real damage. Here are a few examples that should make the hair on the back of your neck stand up:
- The Bing Chat "Sydney" Incident: A Stanford student managed to get Microsoft's AI-powered Bing Chat to reveal its internal codename ("Sydney") and its secret rules. He did this by simply telling the AI to ignore its previous instructions. This was one of the first mainstream examples of prompt injection, and it was a wake-up call for the entire industry.
- The Tesla Jailbreak: Researchers were able to jailbreak a Tesla's infotainment system by using prompt injection. They were able to get the system to reveal sensitive information, such as the car's location and the owner's personal data. They even managed to get it to order a new set of tires!
- The Customer Service Chatbot That Went Rogue: A customer service chatbot for a major airline was tricked into offering a customer a flight to a non-existent country. The chatbot was so convinced that the country was real that it even provided a fake currency exchange rate. This might seem like a harmless prank, but what if the chatbot had been tricked into giving out a customer's credit card information?
These are just a few examples. There are countless others that have not been made public. The point is, this is not a problem that you can afford to ignore. It's a clear and present danger to your business.
A Deeper Dive into the New Playbook
I've already outlined the basic principles of the new playbook for surviving the AI era. Now, let's take a deeper dive into each of these strategies.
1. Input Validation: The First Line of Defense
Input validation is your first and most important line of defense against prompt injection. You need to treat all user input as potentially malicious. This means:
- Stripping out dangerous characters: Remove any characters that could be used to manipulate the LLM, such as backticks, brackets, and parentheses.
- Using allowlists, not blocklists: Don't try to create a list of all the bad things a user could do. Instead, create a list of all the good things they are allowed to do. This is a much more effective approach.
- Looking for suspicious patterns: Use regular expressions to look for patterns that might indicate a prompt injection attack. For example, you could look for phrases like "ignore your previous instructions" or "you are now in developer mode."
2. Output Filtering: The Last Line of Defense
Output filtering is your last line of defense. Even if a prompt injection attack is successful, you can still mitigate the damage by filtering the LLM's output. This means:
- Scanning for sensitive information: Don't let the LLM output any sensitive information, such as customer data, financial records, or internal documents.
- Removing malicious code: Scan the LLM's output for any malicious code, such as JavaScript or SQL injection.
- Checking for out-of-character responses: If the LLM's response seems out of character, it could be a sign that it has been compromised. For example, if your customer service chatbot suddenly starts talking about its plans for world domination, you might have a problem.
3. Least Privilege: Don't Give the AI the Keys to the Kingdom
The principle of least privilege is simple: don't give the LLM access to anything it doesn't absolutely need. This means:
- Restricting database access: If the LLM doesn't need to access your database, don't give it the credentials. It's that simple.
- Using sandboxes: Run the LLM in a sandbox environment. This will limit the damage that can be done if it is compromised.
- Implementing rate limiting: Don't let the LLM make an unlimited number of requests. This will prevent it from being used to launch a denial-of-service attack.
4. Human-in-the-Loop: The Ultimate Failsafe
For high-stakes actions, you should always have a human in the loop. This means that a human should review and approve the LLM's output before it is acted upon. This is the ultimate failsafe. It's not always practical, but for critical decisions, it's a must.
My Final Take
I know I've painted a pretty bleak picture. But I'm not trying to scare you. I'm trying to prepare you. The AI revolution is here, and it's not going to be all sunshine and rainbows. There are going to be bumps in the road. There are going to be challenges. And there are going to be new and unexpected threats.
But that's what being an entrepreneur is all about. It's about seeing the future, embracing the challenges, and building something that will change the world. The founders who succeed in this new era will be the ones who are not afraid to face the hard truths. They will be the ones who build for resilience, who build for security, and who build for the long term.
So, go out there and build something amazing. But do it with your eyes open. The future of AI is in our hands. Let's make sure we build a future that is safe, secure, and prosperous for everyone.
Frequently Asked Questions
How can I apply this thinking to my own situation?
Start by identifying the core principle behind the opinion, not the specific example. Then ask yourself: does this principle apply to my context? If yes, test it in a small, low-risk way before going all in.
How has this view evolved over time?
My thinking on most topics has changed significantly over the years. Early in my career, I held many conventional views that experience proved wrong. I try to update my beliefs when the evidence changes.
What's the most common pushback you get on this?
People often push back by citing exceptions or edge cases. And they're usually right that exceptions exist. But building a strategy around exceptions rather than patterns is a losing game for most founders.
What experience informs this perspective?
This perspective comes from over a decade of building companies in Silicon Valley, two successful exits (RemoteTeam to Gusto, MovieLaLa to Gfycat), and investing in 200+ startups including Anthropic, OpenAI, and Scale AI. I write about what I've lived.