I remember it like it was yesterday. I was sitting in a board meeting for a promising AI startup—I won’t name names—and the CEO was sweating bullets. Their burn rate was astronomical, and they were about to miss a critical product deadline. The reason? They’d bet the farm on a single GPU provider and were now staring down the barrel of a six-month waiting list for the hardware they needed to scale. The company folded three months later. A classic case of a brilliant idea killed by bad infrastructure choices.
Everyone in Silicon Valley is obsessed with AI models. They talk about parameter counts, training data, and the latest breakthroughs from OpenAI and Anthropic. But nobody wants to talk about the ugly, expensive, and brutal reality of the hardware that runs it all. After two exits and over 200 angel investments, I’ve seen this movie more times than I can count. Founders fall in love with a model, but they forget that the model is useless without the picks and shovels—the GPUs, the data centers, the networking—that bring it to life.
I’m writing this because I’m tired of seeing good startups die for bad reasons. I’m going to share the hard-won lessons I wish someone had told me when I was starting out. Here are the six AI infrastructure mistakes that are secretly killing your startup.
1. Blindly Following the NVIDIA Hype
Let’s get this out of the way. NVIDIA makes some incredible hardware. Their GPUs are the gold standard for a reason. But they are not the only game in town, and blindly committing your entire stack to them is a recipe for disaster. The demand for high-end NVIDIA chips is so insane that I’ve seen startups pay 3x the list price on the gray market just to get their hands on a few H100s. It’s a bloodbath.
This isn’t just about cost. It’s about vendor lock-in. When your entire software stack is built on CUDA, you’re at the mercy of NVIDIA’s roadmap, their pricing, and their supply chain. What happens when they decide to jack up prices? Or when the next GPU shortage hits? You’re stuck.
I’ve been advising my portfolio companies to diversify their hardware strategy for years. Look at Google’s TPUs. Look at what Cerebras is doing. There are options. Don’t be a lemming and follow the crowd off the NVIDIA cliff.
2. Underestimating the Power of Custom Silicon
Back at RemoteTeam, we hit a wall. We were processing millions of employee records and running complex payroll calculations. The off-the-shelf hardware just wasn’t cutting it. It was too slow and too expensive. So, we did something crazy. We built our own custom silicon. It was one of the hardest things I’ve ever done, but it was also one of the best decisions we ever made.
Building your own chips isn’t for everyone. It’s a massive undertaking that requires a specialized team and a lot of capital. But if you’re operating at scale, the benefits can be enormous. You can design a chip that is perfectly optimized for your specific workload, which can lead to a 10x or even 100x improvement in performance and efficiency. Just look at what Google did with the TPU. They saw the writing on the wall and invested in custom silicon early on. Now they have a massive competitive advantage.
3. Ignoring the Data Center
This one seems obvious, but you’d be surprised how many founders overlook it. They get so focused on the servers that they forget about the building that houses them. The data center is not a commodity. The location, the power, the cooling, the security—it all matters.
I once invested in a startup that had their entire infrastructure in a single data center in San Francisco. One day, there was a major power outage in the city. Their site was down for 12 hours. They lost millions in revenue and took a massive hit to their reputation. Don’t make the same mistake. You need a distributed, resilient data center strategy.
4. Choosing the Wrong Cloud Provider
AWS, GCP, Azure. They all want your business. And they’ll all tell you that they’re the best platform for AI. Don’t believe the hype. The right cloud provider for you depends on your specific needs.
AWS has the biggest market share and the broadest set of services. But they’re also the most expensive. GCP has the best AI and machine learning tools, thanks to their deep investment in things like TensorFlow and TPUs. But their market share is smaller, and their support can be hit or miss. Azure is a strong contender, especially if you’re already in the Microsoft ecosystem. But their AI offerings are not as mature as the competition.
Do your homework. Run benchmarks. Talk to other founders. Don’t just go with the default choice.
5. Not Having a GPU Shortage Plan
The GPU shortage is not a temporary problem. It’s the new normal. The demand for AI hardware is growing exponentially, and the supply chain can’t keep up. If you don’t have a plan for how to navigate the next shortage, you’re going to be left in the dust.
What does a GPU shortage plan look like? It means having a diversified hardware strategy, as I mentioned earlier. It means building relationships with multiple vendors. It means being smart about how you use your existing hardware. And it means being willing to get creative. I know one startup that bought a bunch of gaming PCs and retrofitted them for AI training. It wasn’t pretty, but it got the job done.
6. Focusing Only on Training, Not Inference
Everyone is obsessed with training. They want to build the biggest, most complex models. But they forget that training is a one-time cost. Inference—the cost of running your model in production—is a recurring cost that can eat you alive if you’re not careful.
I’ve seen startups spend millions of dollars training a model, only to realize that it’s too expensive to run in production. Don’t fall into this trap. You need to be thinking about inference from day one. This means designing your model with an eye towards efficiency. It means choosing the right hardware for your inference workload. And it means being ruthless about optimizing your code.
The Bottom Line
Building a successful AI startup is hard enough. Don’t make it harder by making these unforced errors. The infrastructure you choose will have a massive impact on your company’s trajectory. So, choose wisely. Don’t be afraid to be a contrarian. And for God’s sake, don’t bet the farm on a single vendor.
Related Investments
Sahin Boydas is an angel investor in these companies mentioned in this article:
Frequently Asked Questions
Do all experts agree with this view?
No, and that's fine. The best ideas in business are often contrarian. I share my perspective based on my experience and data, but I encourage you to seek out opposing viewpoints and form your own conclusions.
How can I apply this thinking to my own situation?
Start by identifying the core principle behind the opinion, not the specific example. Then ask yourself: does this principle apply to my context? If yes, test it in a small, low-risk way before going all in.
What's the most common pushback you get on this?
People often push back by citing exceptions or edge cases. And they're usually right that exceptions exist. But building a strategy around exceptions rather than patterns is a losing game for most founders.