What It Really Takes to Build a Hyperscale AI Data Center

Published 2025-05-01 · Updated 2026-05-23 · 6 min read · AI Hardware and Infrastructure · By Sahin Boydas

After years working in Silicon Valley and investing in AI startups, I want to share the lessons I learned about the hardware side of AI — from dealing with GPU shortages to designing custom chips — that are crucial for success.

I still remember the call. It was 2 AM, and one of the founders I’d invested in was on the other end of the line, his voice a mix of panic and exhaustion. “Sahin,” he said, “we’re dead in the water. We can’t get the GPUs. We have the models, we have the customers, but we can’t get the hardware.”

This wasn’t a new conversation. For the past few years, it feels like I’ve had this same talk a hundred times. Everyone is talking about the magic of AI, the incredible models from OpenAI, Anthropic, and the others I’ve been fortunate to back. But very few people are talking about the brutal, unglamorous reality of what it takes to power all of it. The sheer, brute-force physics of it.

After two exits and over 200 angel investments, I’ve seen the sausage get made. I’ve been in the trenches with founders who are trying to build the future, and I can tell you that the future is built on a mountain of hardware. And right now, that mountain is incredibly difficult to climb.

The GPU chokehold is real

Let’s just get this out of the way. The GPU shortage is not some abstract supply chain issue you read about in the news. It’s a day-to-day street fight. When a startup needs to scale, they need thousands of high-end NVIDIA chips—H100s, A100s, you name it. And they need them yesterday.

I had a team that projected they’d need 2,000 H100s to train their next-gen model. They had a letter of intent from a major enterprise customer, a nine-figure deal on the line. They put in their order, and the lead time came back: 52 weeks. A year. In the AI world, a year is an eternity. That’s enough time for three new state-of-the-art models to be released and for your entire business to become irrelevant.

We had to pull every string I had. I called VCs, other founders, even people I knew at the big cloud providers. We ended up piecing together their compute from three different continents, renting smaller batches of GPUs from tier-2 cloud vendors at a 3x markup. It was a mess. It was expensive. But it was the only way to keep the company alive.

This is the reality that isn’t in the pitch decks. You’re not just competing on the quality of your algorithms; you’re competing for the physical resources to run them. It’s a game of who you know and how much you’re willing to pay. And the house—NVIDIA—always wins.

It’s not just about the GPUs

As painful as the GPU bottleneck is, it’s not the only hardware challenge. Building a hyperscale data center is a symphony of interconnected parts, and if one section is out of tune, the whole thing sounds terrible.

Think about networking. When you have tens of thousands of GPUs running in parallel, they need to talk to each other at insane speeds. We’re talking about petabits of data flying around every second. Your standard Ethernet connection isn’t going to cut it. You need specialized, high-bandwidth interconnects like NVIDIA’s NVLink or InfiniBand. And guess what? Those are also in short supply and come with a hefty price tag.

I saw a startup almost go under because they focused all their capital on securing GPUs but skimped on the networking. They had the horsepower, but their GPUs were starving for data. It was like owning a fleet of Ferraris but only having a dirt road to drive them on. Their training runs were so slow and inefficient that they were burning through their cash with nothing to show for it.

Then there’s power and cooling. An H100 GPU can draw up to 700 watts of power under full load. Now multiply that by 10,000. You’re suddenly talking about megawatts of power—enough to run a small town. You can’t just plug that into the wall. You need to build custom power substations. You need industrial-scale cooling systems, often involving liquid cooling, to keep the whole thing from melting.

One of the data centers I visited for a portfolio company was a marvel of engineering. It had its own dedicated power grid connection and a closed-loop liquid cooling system that looked like something out of a sci-fi movie. The monthly electricity bill was over a million dollars. That’s the kind of operational expense you’re dealing with at scale.

The rise of the custom chip

The pain of relying on a single supplier has pushed the industry to a breaking point. The smartest teams are realizing they can’t build their future on someone else’s roadmap. This has led to a Cambrian explosion in custom AI chips.

Google has its TPUs, Amazon has Inferentia and Trainium, and Microsoft is working on its own silicon. And it’s not just the giants. I’ve invested in several startups that are designing their own chips, tailored specifically for their models and workloads. It’s a high-stakes bet. Chip design is incredibly complex and expensive. A single tape-out—the final step of sending your design to the foundry—can cost millions of dollars.

But the payoff can be huge. A custom chip can be 10x more efficient for a specific task than a general-purpose GPU. That means lower power consumption, lower latency, and a lower cost per query. For a company operating at hyperscale, that can translate into hundreds of millions of dollars in savings.

The decision to build your own chip is one of the toughest a founder can make. It’s a multi-year journey with a high risk of failure. But for those who get it right, it can be the ultimate competitive advantage. It’s the difference between renting an apartment and owning the building.

Don’t forget the edge

While everyone is focused on these massive, centralized data centers, there’s another quiet revolution happening at the edge. As AI models get more efficient, it’s becoming possible to run them on smaller, specialized devices—phones, cars, sensors, and more.

This is a game-changer. Edge AI reduces latency, improves privacy, and lowers the reliance on massive data centers. I have an investment in a company that’s using AI for quality control in manufacturing. They have small, ruggedized AI cameras on the assembly line that can detect defects in real-time. They don’t need to send a video stream to the cloud; the inference happens right there on the device.

This creates a whole new set of hardware challenges. You need chips that are not only powerful but also incredibly power-efficient. You’re not plugging these into a power substation; you’re running them on a battery. This has opened the door for a new class of chip designers and hardware companies that are focused on the unique constraints of the edge.

The hard truth

Building in the world of AI right now is not for the faint of heart. It’s a constant battle against physical constraints. It requires a deep understanding of the full stack, from the silicon up to the application layer. The founders who succeed are the ones who are not just brilliant coders but also savvy supply chain managers and hardware strategists.

I’ve seen too many promising AI startups fail because they underestimated the hardware side of the equation. They had beautiful models and elegant code, but they couldn’t get the machines to run it on. It’s a hard lesson to learn, and it’s one that I try to impart to every founder I work with.

So, the next time you see a mind-blowing AI demo, take a moment to think about the invisible infrastructure that’s making it all possible. Think about the thousands of GPUs, the miles of high-speed cable, and the megawatts of power. That’s where the real war is being fought. And it’s a war that’s only just beginning.

The People Problem

I've talked a lot about hardware, but there's another, even scarcer resource: talent. You can't just hire a random IT team to build and manage a hyperscale AI data center. You need a very specific, very rare breed of engineer.

These are people who understand power engineering, advanced cooling systems, high-speed networking, and distributed systems. They can debug a faulty liquid cooling pump at 3 AM and then write a script to optimize GPU cluster scheduling. They are, in short, unicorns. And they are in incredibly high demand.

I watched one of my portfolio companies spend six months trying to hire a lead data center architect. They offered a ridiculous salary, equity that would make your eyes water, and a blank check to build the data center of their dreams. They were competing against Google, Meta, and Amazon for the same handful of qualified people. It was a brutal bidding war. They eventually found someone, but it delayed their roadmap by two quarters and cost them a fortune.

This is the hidden cost of building AI infrastructure. It's not just about the capital to buy the hardware; it's about the human capital to make it all work. And right now, the supply of that human capital is nowhere near the demand.

The Software Stack is Everything

Let's say you've done the impossible. You've secured the GPUs, you've built the data center, and you've hired the team. You're still not done. Now you have to make it all work together. And that's a software problem.

Running tens of thousands of GPUs in parallel is not a plug-and-play operation. You need a sophisticated software stack to manage the cluster, schedule jobs, and handle failures. This is where tools like Kubernetes, Slurm, and a whole ecosystem of custom orchestration software come in.

I've seen teams spend years building their own internal platforms to manage their AI infrastructure. It's a massive undertaking, but it's essential. Without a robust software stack, you're just left with a very expensive pile of silicon. You'll have GPUs sitting idle, training jobs failing for no reason, and engineers spending all their time firefighting instead of building models.

This is another area where the big cloud providers have a huge advantage. They've spent a decade and billions of dollars building their own internal software stacks. When you use their services, you're not just renting their hardware; you're renting their expertise. For many startups, that's a trade-off worth making.

The Uncomfortable Truth About AI's Environmental Impact

Finally, there's the topic that no one in the AI industry wants to talk about: the environmental impact. These hyperscale data centers are not just power-hungry; they're also incredibly water-intensive. The cooling systems that keep the GPUs from melting use millions of gallons of water every year.

I visited a data center in Arizona, and the facility manager told me that their biggest operational challenge wasn't power or networking; it was water. They were in the middle of a desert, and they were in a constant battle to secure enough water to keep their data center running. It was a stark reminder that the virtual world of AI has a very real, very physical footprint.

This is an uncomfortable truth for an industry that likes to see itself as clean and forward-thinking. But it's a truth we can't ignore. As we build more and more of these massive AI factories, we need to be thinking about how to do it sustainably. We need to be investing in more efficient cooling technologies, renewable energy sources, and data center designs that minimize their environmental impact.

This isn't just a matter of corporate responsibility; it's a matter of long-term survival. The AI industry can't thrive on a planet that's running out of resources. The sooner we confront that reality, the better.

Frequently Asked Questions

What's the most common pushback you get on this?

People often push back by citing exceptions or edge cases. And they're usually right that exceptions exist. But building a strategy around exceptions rather than patterns is a losing game for most founders.

How can I apply this thinking to my own situation?

Start by identifying the core principle behind the opinion, not the specific example. Then ask yourself: does this principle apply to my context? If yes, test it in a small, low-risk way before going all in.

Do all experts agree with this view?

No, and that's fine. The best ideas in business are often contrarian. I share my perspective based on my experience and data, but I encourage you to seek out opposing viewpoints and form your own conclusions.

More in AI Hardware and Infrastructure

  • From TPU to Your Own Custom Silicon: A Founder's Journey — After years in the trenches of Silicon Valley, I've seen firsthand how the right AI hardware can make or break a company. I'm sharing the hard-won lessons and contrarian insights I wish I had when I started, from navigating the GPU shortage to building our own custom silicon.
  • Surviving the GPU Apocalypse: A Founder's Guide to the Shortage — After years in the trenches of Silicon Valley, I've seen firsthand how the right AI hardware can make or break a company. I'm sharing the hard-won lessons and contrarian insights I wish I had when I started, from navigating the GPU shortage to building our own custom silicon.
  • The 6 AI Infrastructure Mistakes That Are Secretly Killing Your Startup — After years in the trenches of Silicon Valley, I've seen firsthand how the right AI hardware can make or break a company. I'm sharing the hard-won lessons and contrarian insights I wish I had when I started, from navigating the GPU shortage to building our own custom silicon.
  • Cerebras Systems — Portfolio Company | Angel Investment by Sahin Boydas — Building the world's largest AI chips for training and inference at unprecedented scale.
  • Why the Future of AI Is Not in the Cloud 338 — After years in the trenches of Silicon Valley, I've seen firsthand how the right AI hardware can make or break a company. I'm sharing the hard-won lessons and contrarian insights I wish I had when I started, from navigating the GPU shortage to building our own custom silicon.
  • The Counterintuitive Truth About AI Chip Design 869 — After years in the trenches of Silicon Valley, I've seen firsthand how the right AI hardware can make or break a company. I'm sharing the hard-won lessons and contrarian insights I wish I had when I started, from navigating the GPU shortage to building our own custom silicon.

All AI Hardware and Infrastructure articles · Sahin's angel investments · Startups he founded