The cloud is a trap.
I’ve seen it dozens of times. A promising startup raises a solid seed round, gets amazing traction, and then slams headfirst into a wall. That wall isn’t the competition. It’s their AWS bill. They’re spending hundreds of thousands, sometimes millions, a year just on inference costs for a large language model. It’s a cash bonfire. And for what? For an experience that’s often slow, generic, and raises all sorts of privacy red flags.
After two successful exits—RemoteTeam to Gusto and MovieLaLa to Gfycat—and investing in over 200 companies, including foundational players in AI like Anthropic, OpenAI, and Scale AI, I’ve spent the last two years obsessed with the solution: on-device AI. It's a core theme in my book, 'Becoming Top 1%', because it’s about finding an unfair advantage. It’s not just a niche, it’s the future. It’s about moving from a world of slow, expensive, centralized AI to one of fast, private, and deeply personalized intelligence that lives right on your user's hardware.
This isn’t just theory. At RemoteTeam, which was acquired by Gusto, we saw the power of putting tools directly in the hands of the user. We cut down on so much backend complexity. On-device AI is that principle on steroids. This is the playbook for building the next generation of AI-powered products.
Why On-Device AI is the Only Path Forward
For the last few years, the race was all about size. Who could build the biggest model? Who had the most GPUs? That was the old game. The new game is about efficiency and user experience. The advantages of running AI on the edge are so massive that in a few years, we’ll look back at cloud-only AI the same way we look at dial-up internet.
1. Your Data Stays Your Data (Privacy)
This is the big one. Every time you send a prompt to a cloud API, you’re sending your data, or your users' data, to a third party. Do you know how it’s being stored? Who has access? Are they training on it? With on-device AI, the data never leaves the phone or laptop. For applications in healthcare, finance, or anything handling sensitive information, this isn’t a feature—it’s a requirement. It’s the ultimate form of privacy.
2. Instantaneous Speed
Network latency is a killer for user experience. Waiting a second or two for a response from a cloud LLM feels like an eternity. On-device models respond instantly. The interaction feels fluid, natural, and real. Imagine a creative tool that provides feedback as you type, or a coding assistant that suggests completions with zero lag. That’s the kind of magic that on-device AI enables.
3. It Just Works (Offline Access)
I remember the early days of my second company, MovieLaLa. We were trying to build the ultimate movie recommendation engine. We had this grand vision of giving you the perfect movie to watch, right when you needed it. But we kept hitting the same wall: what if you're on a flight with spotty Wi-Fi? The cloud can't help you there. On-device AI works completely offline. Your app’s core intelligence is always available, regardless of internet connectivity. This is a huge win for utility and reliability.
4. The End of Insane Cloud Bills
Let’s talk numbers. A popular cloud LLM might cost you $20 for every million tokens of output. If you have a million users, and they each generate just 1,000 tokens a day, you’re looking at a bill of $600,000 a month. That’s not sustainable. It’s a tax on innovation. By moving inference to the user's device, you eliminate that cost entirely. The compute is free, paid for by the user when they bought their device.
The Practical Guide: How to Actually Do It
This isn't a distant dream. You can start building with on-device AI today. The hardware and software are finally ready for primetime.
The Rise of the Small Language Model (SLM)
The breakthrough has been the development of powerful, yet small, language models. You don’t need a 175-billion parameter model to summarize an email or power a chatbot. Models like Phi-3 Mini, Llama 3 8B, and Gemma 2B are incredibly capable and designed to run efficiently on consumer hardware.
These models are the result of better training data, improved architectures, and new distillation techniques. They punch way above their weight class, delivering performance that was state-of-the-art just a year or two ago.
Your Toolkit for Local LLMs
Getting started is easier than you think. Here are the tools I recommend to my portfolio companies:
LM Studio: If you’re just getting started, this is the place to begin. It’s a desktop app that lets you download, manage, and chat with open-source models with a simple graphical interface. No command line needed. It’s the easiest way to get a feel for what these models can do.
Ollama: For the more technical folks, Ollama is fantastic. It’s a command-line tool that lets you run and manage LLMs with a single command. It handles all the complexity of model weights, quantization, and serving. You can get a powerful model like Llama 3 running in minutes.
Model Quantization: This is the secret sauce. Quantization is a process that shrinks models by reducing the precision of their weights (think of it like compressing a file). A 13-billion parameter model might take 26GB of RAM, but a quantized version might only need 6GB, with a minimal loss in quality. Tools like
llama.cpphandle this automatically, using formats like GGUF.
The Hardware is Already in Your Pocket
For years, the excuse was that consumer hardware wasn’t powerful enough. That’s no longer true. Apple’s M-series chips in their Macs are AI powerhouses with a built-in Neural Engine. New PCs are shipping with NPUs (Neural Processing Units) specifically for AI tasks. And the latest smartphones from Apple and Google have more AI processing power than supercomputers from a decade ago.
The hardware is there. It’s sitting dormant, waiting for developers to unlock its potential.
The Future is Personal
I didn’t invest in 200+ companies to watch them all become slaves to a few cloud providers. I invested to see them build new, category-defining experiences. The biggest opportunity in AI right now isn’t in the cloud; it’s on the edge.
It’s about creating AI that is truly personal—an assistant that knows you, a creative partner that understands your style, a tool that adapts to your workflow. And the only way to do that safely and privately is to run it on the user's own device.
Stop burning your cash on cloud APIs. The future is local. Start building it.
Frequently Asked Questions
How often is this guide updated?
I revisit and update my guides regularly as I learn new things and as the market evolves. The core principles tend to stay stable, but specific tactics and tools get refreshed based on what's working right now.
Is this guide based on real experience?
Every recommendation in this guide comes from direct experience, either from building and selling my own companies, or from patterns I've observed across 200+ angel investments. I don't write about things I haven't personally tested.
What if I disagree with some of the advice?
Good. That means you're thinking critically, which is exactly what a good founder should do. Take what resonates, test it, and discard what doesn't work for your specific situation. No advice is universal.
Who is this guide designed for?
This guide is written for founders and operators who want practical, actionable advice rather than theoretical frameworks. Whether you're just starting out or scaling an existing business, the principles here apply across stages.