''' I see it everywhere. Founders, investors, even engineers, all obsessing over one thing: GPUs. How many H100s can we get? What's our NVIDIA bill look like? It's a frenzy. And it's completely missing the point.
Everyone thinks the biggest line item for running an AI company is the raw compute. The sexy, headline-grabbing GPUs. I’m here to tell you that’s a dangerously simplistic view. I’ve built and sold two companies, RemoteTeam and MovieLaLa, and I’ve invested in over 200 startups, including some of the biggest names in AI like Anthropic, OpenAI, and Scale AI. I’ve seen the P&Ls. I know where the money really goes.
Focusing only on GPUs is like looking at the tip of an iceberg. It’s the most visible part, but the real danger, the bulk of the mass, is lurking below the surface. The truth is, your AI infrastructure cost is a multi-headed beast. And if you’re not watching all the heads, you’re going to get eaten.
The GPU Illusion
Don't get me wrong, GPUs are expensive. Absurdly so. When we were scaling up MovieLaLa, which was a heavy user of recommendation algorithms, we felt the pain. We were constantly trying to optimize our models to squeeze more performance out of every chip. It was a battle. But it was a battle we could understand, a variable we could model.
But here’s the kicker: the GPU costs were predictable. They were a known quantity. The real surprises, the budget-busting, "where did all our money go?" moments came from somewhere else entirely.
They came from the hidden costs. The ones nobody wants to talk about because they’re not as glamorous as a rack full of blinking NVIDIA cards.
Beyond the Chips: The Four Horsemen of Hidden AI Costs
There are four main areas that will silently drain your bank account if you’re not careful: Data, Networking, Storage, and People.
1. Data: The Insatiable Beast
This is the big one. The single most underestimated cost in AI. Your models are nothing without data. And data is expensive. Not just in terms of acquisition, but in cleaning, labeling, and processing.
At one of my portfolio companies, a promising AI-powered drug discovery platform, they were spending more on data curation than on their entire engineering team. They had to license specialized biological datasets, a process that made our enterprise SaaS negotiations look like a lemonade stand. Then came the cleaning. The data was a mess. Full of inconsistencies, missing values, and weird formatting. They had to hire a team of PhDs just to make it usable.
And that’s before you even get to the labeling. If you’re doing any kind of supervised learning, you need labeled data. And that means humans. Lots of them. Services like Scale AI and Hugging Face have built massive businesses on this fact alone. It’s a huge, ongoing operational expense.
2. Networking: The Silent Killer
Ever heard the phrase "don't underestimate the bandwidth of a station wagon full of tapes"? It’s an old saying, but it’s more relevant than ever in the age of AI. Moving data around is expensive. Really expensive.
When you’re training a large model, you’re constantly shuffling petabytes of data between storage and your compute cluster. Every time that data crosses a network boundary, you’re paying. Egress fees from cloud providers are notorious for this. They are the silent killer of AI budgets.
I remember a founder coming to me, completely bewildered. Their AWS bill had tripled in a month. They thought it was a bug. It wasn’t. They had inadvertently set up their training pipeline in a way that was constantly moving data between regions. A simple architectural mistake had cost them hundreds of thousands of dollars.
3. Storage: The Slow Bleed
Storage is cheap, right? That’s the common wisdom. And on a per-gigabyte basis, it is. But we’re not talking about gigabytes anymore. We’re talking about petabytes. Exabytes, even.
Your raw data, your processed data, your model checkpoints, your experiment logs… it all adds up. And you have to store it somewhere. And back it up. And secure it. It’s a slow bleed, a constant drip-drip-drip that can turn into a torrent if you’re not paying attention.
And it’s not just the cost of the storage itself. It’s the cost of managing it. Of making sure it’s accessible, reliable, and compliant. It’s a full-time job. Or, more accurately, a full-time team.
4. People: The Most Expensive Resource of All
This is the most important, and most expensive, piece of the puzzle. You can have all the data, GPUs, and storage in the world, but if you don’t have the right people, you have nothing.
And the right people are expensive. And rare. Good AI/ML engineers are some of the most sought-after talent on the planet. They command massive salaries, and they’re notoriously difficult to retain. You’re not just competing with other startups; you’re competing with Google, Meta, and OpenAI.
I’ve seen companies spend more on recruiting a single top-tier AI researcher than on their entire GPU budget for a year. And it was worth it. Because that one person could come in and 10x the performance of their models. But you have to be prepared for that level of investment. You have to be prepared to build a culture that attracts and retains that kind of talent.
A New Framework for AI Infrastructure
So, what’s the solution? It’s not to stop using GPUs. It’s to stop only thinking about GPUs. It’s to adopt a more holistic, more realistic view of AI infrastructure costs.
It starts with a simple mindset shift: Your AI infrastructure is not a cost center. It’s a product. You need to treat it with the same level of rigor and discipline as you would your customer-facing application.
That means:
- Obsess over your data pipeline. It’s the foundation of everything you do. Invest in tools and processes that make it as efficient and automated as possible.
- Design for network efficiency. Think about data locality. Minimize data movement. Understand the cost implications of your architecture.
- Have a storage strategy. Don’t just throw everything in S3 and hope for the best. Think about data lifecycle management. What do you need to keep hot? What can be moved to cold storage?
- Invest in your people. They are your biggest asset. Create a culture of learning and experimentation. Give them the tools and resources they need to be successful.
The Future is Serverless AI
I believe the long-term solution to this problem is the continued rise of serverless AI platforms. Services that abstract away the complexity of infrastructure and allow you to focus on what you do best: building great models and products.
Companies like Replicate, Anyscale, and Modal are at the forefront of this movement. They are building the platforms that will power the next generation of AI applications. They are the ones who are truly solving the AI infrastructure problem, not by making GPUs cheaper, but by making them invisible.
So next time you hear someone bragging about their GPU cluster, ask them about their data pipeline. Ask them about their networking costs. Ask them about their people. Because that’s where the real conversation is. That’s the truth about AI infrastructure costs. '''
Frequently Asked Questions
What experience informs this perspective?
This perspective comes from over a decade of building companies in Silicon Valley, two successful exits (RemoteTeam to Gusto, MovieLaLa to Gfycat), and investing in 200+ startups including Anthropic, OpenAI, and Scale AI. I write about what I've lived.
What's the most common pushback you get on this?
People often push back by citing exceptions or edge cases. And they're usually right that exceptions exist. But building a strategy around exceptions rather than patterns is a losing game for most founders.
How has this view evolved over time?
My thinking on most topics has changed significantly over the years. Early in my career, I held many conventional views that experience proved wrong. I try to update my beliefs when the evidence changes.
Do all experts agree with this view?
No, and that's fine. The best ideas in business are often contrarian. I share my perspective based on my experience and data, but I encourage you to seek out opposing viewpoints and form your own conclusions.