The invoice arrived in my inbox with a thud. $101,428.73. That was our AI infrastructure bill for last month. For a moment, I just stared at it. It was a terrifying and clarifying number.
I’m Sahin Boydas. I’ve built and sold a couple of companies, RemoteTeam and MovieLaLa, and now I spend a lot of my time as an angel investor in over 200 startups, including some you might know like OpenAI, Anthropic, and Scale AI. But I’m still an operator at heart. I’m in the trenches building my next thing, and that means I’m dealing with the same messy realities as every other founder. Realities like a six-figure cloud bill.
When you’re moving fast, it’s easy to just keep swiping the corporate card on new services. A little OpenAI here, a little Anthropic there, a vector database, some serverless GPUs. It all seems reasonable in isolation. But it adds up. Fast.
I believe in transparency. So I’m going to do something that makes my finance person nervous. I’m going to give you a line-by-line breakdown of where every single one of those hundred thousand dollars went. No fluff. Just the raw numbers, what we learned, and how we’re fighting to get that number down.
The Anatomy of a Six-Figure AI Bill
So, where did the money go? It wasn’t one single thing. It was a death by a thousand papercuts, a complex mix of large language model APIs, specialized services, and the underlying cloud infrastructure to stitch it all together.
Here’s the high-level breakdown:
Large Language Model APIs: $62,000
- OpenAI (GPT-4 & Turbo): $35,000 - This is our workhorse. We use it for a ton of complex reasoning and generation tasks. The quality is amazing, but that power comes at a premium price. One of our features, a real-time code generation assistant, was responsible for about half of this. A single inefficient query being run in a loop cost us nearly $10k before we caught it.
- Anthropic (Claude 3 Opus & Sonnet): $27,000 - We’ve been testing Opus for high-stakes, long-context summarization and it’s been phenomenal. Sonnet is our go-to for faster, more cost-effective tasks like chat and content moderation. The cost crept up as we shifted more production traffic to it.
Cloud Infrastructure (AWS & GCP): $28,500
- GPU Instances (AWS P4/P5s): $18,000 - We do some fine-tuning in-house. Spinning up these beefy instances is like leaving the meter running on a taxi in rush hour. We had a few instances that were left running over a weekend by mistake. That was a fun discovery on a Monday morning.
- Vector Databases & Storage (Pinecone, S3): $7,500 - Embeddings gotta live somewhere, right? The cost here wasn’t just storage, but the compute needed for indexing and querying. As our dataset grew, so did this bill.
- Standard Compute & Networking: $3,000 - This is all the other stuff. The API gateways, the load balancers, the basic EC2 instances running our application logic. It’s the boring plumbing that you forget about until the bill comes.
Specialized AI Services: $10,928
- Voice & Speech-to-Text (ElevenLabs): $4,500 - We have a feature that allows users to interact via voice. The quality is incredible, but real-time transcription and generation for thousands of users isn’t cheap.
- Image Generation (Midjourney API): $3,428 - We use this for generating custom assets and illustrations. It’s a powerful tool, but every image is a new charge.
- Other Tooling (LangSmith, etc.): $3,000 - Monitoring, logging, and observability for AI are critical. You need to see what’s going on, but these tools have their own usage-based pricing. It’s turtles all the way down.
Seeing it all laid out like this was a wake-up call. The dream of serverless AI is that you only pay for what you use. The nightmare is when you don’t realize just how much you’re using.
My Biggest Takeaways (The Expensive Lessons)
This $100k bill taught me a few things. Expensive lessons, but valuable ones.
1. There is no such thing as "set it and forget it."
Usage-based pricing is a double-edged sword. It’s fantastic for getting started because the barrier to entry is low. But it’s a disaster if you don’t have rigorous monitoring in place. You absolutely need real-time dashboards and alerts. Not just for your overall spend, but for cost per user, cost per feature, cost per API call. We now have a dedicated channel in Slack where cost anomalies are posted automatically. If a specific user’s costs spike 3x in an hour, we know immediately.
2. Caching is not optional.
This was our biggest single mistake. We were making duplicate calls to LLMs for identical prompts. A user asking the same question twice would result in two separate API calls, each costing us money. We implemented a simple caching layer using Redis. If we’ve seen the exact same prompt in the last 24 hours, we just serve the cached response. This alone cut our API costs by nearly 20%. It’s a simple fix, but it’s one you have to be intentional about.
3. The right model for the right job.
Not every task needs a sledgehammer like GPT-4 or Claude Opus. We were using the most powerful, expensive models for everything out of convenience. It was lazy. Now, we have a routing system. Simple classification or formatting tasks? Use a cheap, fast model like Claude 3 Haiku or GPT-3.5-Turbo. A complex, multi-step reasoning task? Okay, now you can bring out the big guns. This model-tiering strategy is saving us an estimated $15-20k per month going forward.
My Unpopular Opinion: Vertical is the Future
Everyone in Silicon Valley is chasing the holy grail of AGI. They want to build the biggest, most powerful foundation model. I’m an investor in some of those companies, and I’m rooting for them. But as an operator, I think the real opportunity is in vertical AI.
Instead of a massive, general-purpose model, I want a model that is an expert in one specific domain. A model for legal contract review. A model for SaaS financial forecasting. A model for biotech research. These models can be smaller, cheaper to run, and more accurate for their specific task. They are trained on proprietary data and workflows that a general model will never have.
This is where we’re focusing our efforts. We’re fine-tuning smaller, open-source models on our own data to create a specialized expert system. The upfront cost in time and GPU hours is higher, but the long-term unit economics are an order of magnitude better. We control our own destiny, and we’re not at the mercy of a pricing change from a big provider.
The Plan to Not Spend $100k Again
So, what now? We’re not turning off the AI. That’s not an option. But we are getting smarter.
First, we assigned a single owner for our cloud bill. One person whose job it is to review it every single day. They have the authority to question any service and shut down any experiment that’s running too hot.
Second, we’re implementing a “cost-first” development culture. Before we build any new AI feature, the first question is: “What is the estimated cost per 1,000 users?” If we can’t answer that, we don’t build it.
Third, we’re doubling down on our vertical AI strategy. The future isn’t just about using AI, it’s about building defensible AI assets that you own.
That $100,000 invoice was painful. It was a month’s worth of salaries for a team of engineers. But it was also the best investment we could have made. It forced us to confront the true cost of building with AI and to get serious about our strategy. Don’t wait for your own six-figure bill to learn these lessons. Start today.
Frequently Asked Questions
What experience informs this perspective?
This perspective comes from over a decade of building companies in Silicon Valley, two successful exits (RemoteTeam to Gusto, MovieLaLa to Gfycat), and investing in 200+ startups including Anthropic, OpenAI, and Scale AI. I write about what I've lived.
What's the most common pushback you get on this?
People often push back by citing exceptions or edge cases. And they're usually right that exceptions exist. But building a strategy around exceptions rather than patterns is a losing game for most founders.
Do all experts agree with this view?
No, and that's fine. The best ideas in business are often contrarian. I share my perspective based on my experience and data, but I encourage you to seek out opposing viewpoints and form your own conclusions.
How can I apply this thinking to my own situation?
Start by identifying the core principle behind the opinion, not the specific example. Then ask yourself: does this principle apply to my context? If yes, test it in a small, low-risk way before going all in.