Everyone talks about AI models, but nobody talks about the brutal reality of the hardware that runs them. Here's the unfiltered truth about what it really takes to build and scale AI infrastructure, based on my years in the trenches building and investing in AI companies.
The GPU Gold Rush
I remember the exact moment I knew the AI world had changed forever. It was 2022, and I was trying to help one of my portfolio companies, a brilliant team with a game-changing idea, secure a block of NVIDIA H100s. We had the money. We had the connections. It didn't matter. We were told the waitlist was over a year long. A year. In the startup world, that's an eternity.
We spent weeks, calling every contact we had, begging, borrowing, and nearly stealing to get our hands on the chips we needed. It felt less like building a technology company and more like being a concert scalper. That's the reality of the GPU shortage. It's a constant, gut-wrenching scramble that separates the winners from the losers before a single line of code is even written.
This absolute dependency on a single supplier is a nightmare. We've all become addicted to NVIDIA's ecosystem, and while you have to admire their execution, it's a dangerous position for the industry to be in. It stifles innovation and creates an artificial barrier to entry that has nothing to do with the quality of your ideas or your team.
The Cloud is a Trap
I'm going to say something controversial: for serious, large-scale AI, the public cloud is a trap. It's a fantastic place to start, to experiment, to run small-scale inference. But the moment you start training foundational models or serving millions of users, the costs become astronomical. I've seen it happen time and time again. A promising startup raises a solid seed round, only to burn through 70% of it in six months on AWS or GCP bills. It's a death sentence.
The illusion of infinite scalability is just that—an illusion. When you need a thousand H100s for a training run, you'll find yourself facing the same allocation issues and waitlists as everyone else, but now you're paying a massive premium for the privilege. The hyperscalers are buying up all the supply for their own AI services, and you're left fighting for the scraps.
Don't get me wrong, the cloud has its place. But you need to have a clear-eyed view of the economics. Once you reach a certain scale, the only path forward is to build your own infrastructure. It’s not a question of if, but when.
Building Your Own: Not for the Faint of Heart
So you've decided to take the plunge and build your own AI data center. Welcome to the real world. This is where the talk stops and the hard work begins. And I mean hard.
First, there's power. We were planning a 20-megawatt facility for one of my companies, and the conversations with the utility provider were some of the most complex negotiations I've ever been a part of. You're not just asking for a lot of power; you need it to be reliable, redundant, and clean. Any fluctuation can fry millions of dollars worth of equipment in an instant.
Then there's cooling. These GPUs run hot. Incredibly hot. You're essentially trying to cool a series of jet engines packed into a small room. We had to design a custom liquid cooling system, with miles of pipes and redundant pumps, just to keep the temperatures stable. A single failure in the cooling system could bring the entire operation to a halt.
And the supply chain. You think getting GPUs is hard? Try getting the high-speed networking switches, the power distribution units, the fiber optic cables, and all the other specialized equipment you need. It's a global scavenger hunt, and you're competing with every other AI company on the planet. It took us 18 months to go from breaking ground to having a fully operational data center. It was a brutal, exhausting process.
The Promise of Custom Silicon
This is why the holy grail for many of us in the AI space is custom silicon. Designing your own chips, optimized for your specific workloads, is the ultimate way to escape the tyranny of the GPU shortage and the crushing costs of the cloud. It's not easy, and it's not cheap. You're looking at a team of world-class chip designers and a nine-figure investment just to get started.
But the payoff can be enormous. I've invested in several companies that are going down this path, and the performance gains they are seeing are incredible. We're talking 10x improvements in performance per watt. That's a true game-changer. It unlocks new possibilities and allows you to build models and products that would be impossible with off-the-shelf hardware.
This is the future of AI. It's not just about bigger models and more data. It's about a vertically integrated stack, from the silicon up to the application. The companies that will dominate the next decade of AI will be the ones that control their own hardware destiny.
The Road Ahead
Building and scaling AI infrastructure is one of the hardest things you can do in technology right now. It's a high-stakes game of resource allocation, supply chain management, and deep technical expertise. There are no easy answers, and anyone who tells you otherwise is selling something.
But for those who are willing to brave the challenges, the rewards are immense. We are at the very beginning of a new era of computing, and the hardware we build today will be the foundation for the world of tomorrow. It's a daunting task, but it's also the most exciting work I've ever been a part of. The future isn't just written in code; it's forged in silicon and steel.
Frequently Asked Questions
Do all experts agree with this view?
No, and that's fine. The best ideas in business are often contrarian. I share my perspective based on my experience and data, but I encourage you to seek out opposing viewpoints and form your own conclusions.
What's the most common pushback you get on this?
People often push back by citing exceptions or edge cases. And they're usually right that exceptions exist. But building a strategy around exceptions rather than patterns is a losing game for most founders.
What experience informs this perspective?
This perspective comes from over a decade of building companies in Silicon Valley, two successful exits (RemoteTeam to Gusto, MovieLaLa to Gfycat), and investing in 200+ startups including Anthropic, OpenAI, and Scale AI. I write about what I've lived.