I remember sitting in a boardroom in late 2022, watching a founder sweat through his Patagonia vest as he explained why his AI startup was burning $400,000 a month on AWS. He thought he had a model problem. I looked at his infrastructure bill and realized he had a physics problem.
Everyone talks about the models. We obsess over parameter counts, context windows, and benchmark scores. But nobody talks about the brutal reality of the hardware that runs them. As an angel investor in companies like Anthropic, OpenAI, Scale AI, and Hugging Face, I get to see the raw, unfiltered profit and loss statements of the fastest-growing companies on earth. I see the exact line items that keep founders awake at night.
And let me tell you a secret. The GPU is not your biggest cost. It is just the entry ticket.
When I was building RemoteTeam, before Gusto acquired us, our cloud bills were predictable. We scaled our web servers, optimized our databases, and paid a reasonable premium for managed services. The math was simple. You add a user, you pay a fraction of a cent in compute. Even earlier, when I built MovieLaLa and sold it to Gfycat, video processing was expensive, but it was a known quantity. We could model it in a spreadsheet and sleep soundly.
AI breaks that math completely.
We are living through a period where compute is the new oil, but most startups are treating it like water. They leave the tap running. They buy H100s at massive markups, assuming the venture capital will never dry up. But the music always stops. If you want to build a sustainable AI business, you need to understand the true economics of AI hardware.
The Illusion of the GPU Shortage
Let us get one thing straight. The GPU shortage is real, but it is also a convenient excuse for bad engineering.
Founders love to blame Nvidia for their slow shipping cycles. "We cannot get enough compute," they say. But when I look under the hood of their infrastructure, I see clusters running at 30% utilization. They are paying for Ferraris and driving them in school zones.
The actual cost of a GPU is not the purchase price. It is the total cost of ownership over its usable lifespan. An Nvidia H100 might cost $30,000 to buy, but that is just the beginning. You have to power it. You have to cool it. You have to connect it to thousands of other GPUs. And you have to pay the engineers who know how to keep it from crashing.
When you factor in power, cooling, networking, and talent, the silicon itself often represents less than half of your total infrastructure spend. I have seen pitch decks where founders project their costs based purely on the retail price of chips. I reject those pitches immediately. If you do not understand the total cost of ownership, you do not understand your own business.
The Physics of Power and Cooling
AI data centers are essentially massive heaters that happen to do math.
A standard server rack in a traditional data center might draw 5 to 10 kilowatts of power. A rack fully loaded with modern AI accelerators can draw 40 to 100 kilowatts. The power grid was not designed for this. The cooling systems were not designed for this.
I recently visited a facility being built specifically for AI workloads in the Nevada desert. The cooling infrastructure looked like something out of a nuclear submarine. They were using direct-to-chip liquid cooling because pushing cold air through the servers simply could not remove the heat fast enough. The pipes carrying the coolant were thicker than my arm.
Energy is becoming the ultimate bottleneck. You can print more money, but you cannot instantly print more megawatts. Securing power contracts is now a strategic advantage. The companies that win the next decade of AI will not just be the ones with the best algorithms. They will be the ones with the cheapest electricity.
If you are running your own hardware, your power bill will eventually eclipse your hardware depreciation. If you are renting from a cloud provider, that cost is baked into your hourly rate. Either way, you are paying for the physics of moving electrons and removing heat. I tell my founders to look at geography differently now. We used to build data centers near fiber optic hubs. Now we build them near hydroelectric dams and nuclear plants.
The Networking Tax
You can buy 10,000 GPUs, but if they cannot talk to each other fast enough, you just bought 10,000 very expensive space heaters.
Training large language models requires moving massive amounts of data between chips. If one GPU has to wait for data from another GPU, it sits idle. And an idle GPU is burning cash.
This brings us to the networking tax. To connect these chips, you need high-speed interconnects like InfiniBand. You need optical transceivers. You need specialized switches. The cost of the networking gear can easily add 20% to 30% to the total cost of your cluster.
I have seen startups try to cut corners here. They buy the best GPUs and connect them with standard Ethernet to save money. It is a fatal mistake. Their training runs take twice as long, which means they end up paying twice as much in compute time. In the world of AI infrastructure, being cheap is incredibly expensive.
When I look at the infrastructure of companies like OpenAI or Anthropic, the networking architecture is treated with the same reverence as the model architecture. They understand that the network is the computer. If your interconnects are slow, your entire multi-billion dollar cluster is slow.
The Memory Wall Nobody Mentions
There is a technical reality that most investors completely miss when they look at AI hardware. It is called the memory wall.
Processors have gotten exponentially faster over the last twenty years. Memory has not. We can do math incredibly quickly, but moving the data from the memory chips to the processor takes time. In modern AI workloads, especially large language models, the bottleneck is rarely the compute. The bottleneck is the memory bandwidth.
When you generate a token with a model like Llama 3, the GPU has to load the entire model weights from memory into the processor for every single word it generates. If you have a 70 billion parameter model, you are moving hundreds of gigabytes of data back and forth constantly.
This is why Nvidia's High Bandwidth Memory, or HBM, is such a massive competitive advantage. They package the memory directly next to the processor on the same piece of silicon. It is an engineering marvel. But it is also incredibly expensive and difficult to manufacture. The supply chain for HBM is just as constrained as the supply chain for the GPUs themselves.
I see founders trying to run massive models on consumer-grade GPUs like the RTX 4090. They look at the compute teraflops and think they found a loophole. Then they hit the memory wall. The consumer cards simply do not have the memory bandwidth to serve multiple users concurrently. Their latency spikes, their users churn, and their business model collapses.
Understanding the difference between compute-bound and memory-bound workloads is the difference between a profitable AI product and a money pit. You have to design your infrastructure around your specific bottleneck.
The Talent Premium
Finding someone who can write PyTorch is easy. Finding someone who can optimize CUDA kernels to squeeze 15% more utilization out of your cluster? That person costs $800,000 a year.
This is the hidden human cost of AI infrastructure. The software stack for AI hardware is notoriously difficult to master. Nvidia's moat is not just their silicon. It is CUDA. It is the decades of software optimization that makes their chips actually usable.
But even with CUDA, getting maximum performance out of a cluster requires deep, specialized knowledge. You need engineers who understand memory bandwidth, kernel fusion, and distributed training topologies. There are maybe a few thousand people in the world who are truly elite at this.
When you are spending $10 million a month on compute, paying an engineer $1 million a year to make that compute 20% more efficient is a massive bargain. The problem is finding them. The talent premium is a massive line item that most founders forget to include in their financial models. I have personally connected founders with these rare engineers, and the negotiations look more like professional sports contracts than standard tech hires.
The Cloud Provider Trap
When you are a young startup, the major cloud providers look like your best friends. They offer you hundreds of thousands of dollars in free credits. They give you white-glove support. They make it incredibly easy to spin up an instance and start training.
It is a trap.
Those credits run out faster than you think. And when they do, you are locked into an ecosystem with exorbitant egress fees and massive markups on compute. I have seen companies burn through a million dollars in AWS credits in three months, only to realize their actual run rate is completely unsustainable.
The cloud providers are playing a brilliant game. They are subsidizing the early days of AI development to ensure they own the infrastructure layer for the next decade. They are acting like venture capitalists, but instead of taking equity, they are taking your future cash flow.
If you are building a serious AI company, you need a multi-cloud strategy from day one. You need to be able to move your workloads to whichever provider offers the best spot pricing. You need to look at specialized AI clouds like CoreWeave or Lambda Labs. They do not have the massive ecosystem of AWS or GCP, but they offer bare-metal access to GPUs at a fraction of the cost.
Better yet, if your workloads are predictable, you should be buying your own hardware. The payback period for a server rack of GPUs is often less than a year compared to cloud rental prices. Yes, it requires capital expenditure. Yes, it requires hiring infrastructure engineers. But if you are spending millions a year on compute, renting is financial malpractice.
The Edge AI Reality
So, how do we escape this trap? How do we build AI products that do not require burning a pile of cash every time a user hits "submit"?
The answer is moving inference to the edge.
Right now, almost all AI inference happens in the cloud. You send a request from your phone, it travels to a massive data center, a GPU processes it, and the answer comes back. This is incredibly inefficient. It introduces latency, it costs a fortune in cloud compute, and it raises massive privacy concerns.
We need to push the compute to where the data is generated. Your phone, your laptop, your car.
Apple is already doing this. They are quietly building neural engines into every device they ship. They understand that the only sustainable way to run AI at a global scale is to make the user pay for the hardware and the electricity.
Edge AI is not just a technical optimization. It is a fundamental shift in the business model. If you can run your model locally on the user's device, your marginal cost of inference drops to zero. That is how you build a software business with 90% gross margins in the age of AI. I am actively investing in startups that are compressing models to run on consumer hardware because that is where the real scale will happen.
The Rise of Custom Silicon
For the massive models that must remain in the cloud, the endgame is custom silicon.
Google saw this coming a decade ago. They realized that running neural networks on general-purpose GPUs was inefficient. So they built the Tensor Processing Unit, or TPU. A TPU is an application-specific integrated circuit designed to do one thing incredibly well: matrix multiplication.
By stripping away all the graphics processing logic that a GPU needs, Google created a chip that is faster and more power-efficient for AI workloads.
Now, everyone is following suit. Amazon has Trainium and Inferentia. Microsoft has Maia. Meta is building their own chips. Even OpenAI is reportedly exploring custom silicon.
If you are spending billions of dollars on compute, you cannot afford to pay Nvidia's 70% gross margins forever. You have to build your own chips. It is a massive upfront capital expenditure, but at a certain scale, it is the only way to make the math work. I tell founders that if they are planning to be a foundational model company, they need a silicon strategy. You cannot just be a software company anymore.
The Data Center Real Estate Boom
There is another layer to this infrastructure puzzle that most people ignore entirely. Real estate.
You cannot just put an AI data center anywhere. You need proximity to massive power substations. You need access to vast amounts of water for cooling, or at least a climate that supports advanced cooling techniques. You need fiber optic backbones.
I have friends in commercial real estate who are making fortunes right now just flipping land that happens to have the right power permits. The physical footprint of AI is expanding rapidly. We are seeing old factories and abandoned shopping malls being retrofitted into compute hubs.
This physical reality creates a massive barrier to entry. A software startup used to need a laptop and a coffee shop. An AI infrastructure startup needs zoning board approvals and environmental impact studies. The capital requirements are staggering. This is why we see sovereign wealth funds getting involved. The scale of investment required to build the next generation of AI data centers looks more like national infrastructure projects than traditional venture capital rounds.
The Open Source Hardware Movement
While the giants are building custom silicon, there is a fascinating counter-movement happening in the open-source world.
We are seeing the rise of open hardware architectures like RISC-V. Engineers are collaborating globally to design chips that anyone can manufacture without paying massive licensing fees. It is still early days, but the parallels to the early days of Linux are striking.
If open-source hardware can achieve even a fraction of the success of open-source software, it will radically alter the economics of AI. Imagine a world where the designs for highly efficient AI accelerators are freely available, and you just pay a foundry to print them.
I am watching this space closely. The companies that figure out how to monetize the open-source hardware ecosystem will build massive businesses. It is a high-risk bet, but the potential payoff is astronomical.
What This Means for Founders
If you are building an AI startup today, you need to think like a CFO as much as a CTO.
You cannot just build a cool demo and assume the unit economics will sort themselves out. You need to know exactly how much it costs to serve a single request. You need to know your GPU utilization rates. You need to understand the difference between memory-bound and compute-bound workloads.
Here are the rules I give to my portfolio companies:
1. Optimize before you scale. Do not buy more compute until you have squeezed every drop of performance out of what you already have. Use quantization. Implement speculative decoding. Profile your code. I have seen companies cut their compute bills in half just by spending a week optimizing their inference pipeline.
2. Match the hardware to the task. You do not need an H100 to run a simple classification model. Use cheaper, older GPUs for less demanding tasks. Look into alternative providers. Sometimes an A10G is exactly what you need. Stop buying the most expensive hardware just because it is the newest.
3. Plan for the edge. If your product can run locally, make it run locally. The open-source community is doing incredible work making models smaller and more efficient. Take advantage of it. The future belongs to hybrid architectures where the heavy lifting happens in the cloud and the fast, personalized inference happens on the device.
4. Track your unit economics obsessively. AI is the first software category where your cost of goods sold can scale faster than your revenue. If you are losing money on every request, you cannot make it up in volume. You need a clear path to profitability on a per-user basis from day one.
The Hard Truth
I have been through enough hype cycles in Silicon Valley to know how this plays out. I saw it with mobile. I saw it with crypto. Now I am seeing it with AI.
The companies that survive the inevitable correction will not be the ones with the most hype. They will be the ones with the best margins.
We are building the most important technology in human history. But the laws of physics and the laws of economics still apply. The GPU is just a tool. The real competitive advantage is knowing how to use it efficiently.
Stop obsessing over the GPU shortage. Start obsessing over your unit economics. That is how you build a company that lasts. That is how you become the top 1%.
Frequently Asked Questions
Do all experts agree with this view?
No, and that's fine. The best ideas in business are often contrarian. I share my perspective based on my experience and data, but I encourage you to seek out opposing viewpoints and form your own conclusions.
How has this view evolved over time?
My thinking on most topics has changed significantly over the years. Early in my career, I held many conventional views that experience proved wrong. I try to update my beliefs when the evidence changes.
What experience informs this perspective?
This perspective comes from over a decade of building companies in Silicon Valley, two successful exits (RemoteTeam to Gusto, MovieLaLa to Gfycat), and investing in 200+ startups including Anthropic, OpenAI, and Scale AI. I write about what I've lived.
How can I apply this thinking to my own situation?
Start by identifying the core principle behind the opinion, not the specific example. Then ask yourself: does this principle apply to my context? If yes, test it in a small, low-risk way before going all in.