The Rise of Small Language Models: Why Bigger Isn't Always Better #38

Published 2025-09-13 · Updated 2026-05-23 · 7 min read · Large Language Models · By Sahin Boydas

Everyone is obsessed with massive, billion-parameter models, but they're missing the bigger picture. I'll show you why small language models are the future of AI and how they're quietly powering a revolution in efficiency and accessibility.

I’ve seen a lot of hype in Silicon Valley. I’ve seen bubbles inflate and burst. But the current obsession with massive, billion-parameter language models feels different. It feels like we’re all staring at the sky, mesmerized by the moon, and completely missing the rocket ship taking off right behind us.

Everyone is chasing bigger. More parameters, more data, more compute. We’re in an arms race where the only metric that seems to matter is size. As someone who has backed over 200 companies, including some of the biggest names in AI like Anthropic, OpenAI, Scale AI, and Hugging Face, I can tell you this is a mistake. A big one.

I’m not saying large models are useless. They’re incredible research tools. They push the boundaries of what’s possible. But for 99% of real-world applications? They’re overkill. They’re a sledgehammer to crack a nut.

The future of AI isn't bigger models, it's smarter ones. And right now, “smarter” means smaller.

The Aha! Moment: My RemoteTeam Story

When we were building RemoteTeam, which was later acquired by Gusto, we were obsessed with efficiency. We had to be. We were a small, scrappy team trying to build a global product. Every dollar, every line of code, every server cycle mattered.

We were early adopters of AI, using it to automate everything from payroll summaries to compliance checks. We tried using some of the big, early models, and the experience was painful. They were slow, expensive, and a nightmare to integrate. We’d spend weeks trying to get a simple feature to work, only to have it break with the next API update.

Out of sheer necessity, we started experimenting with smaller, more focused models. We trained models specifically for our use case, on our own data. And you know what? They worked. Not only did they work, they worked better. They were faster, cheaper, and more reliable than their bigger, more famous cousins.

That was my “aha!” moment. I realized that the future of AI in business wasn’t about having one giant model that could do everything. It was about having a whole army of small, specialized models, each one a master of its own domain.

The Unseen Revolution: On-Device AI

One of the most exciting things about small language models is that they can run on-device. This is a huge deal. It means you don’t need a massive server farm to run your AI applications. You can run them directly on your phone, your laptop, even your car.

Think about the implications. Imagine a world where your phone’s virtual assistant is truly personal, running a model trained on your own data, without ever sending a single byte to the cloud. Imagine a car that can learn your driving habits and adjust its performance in real-time. Imagine a doctor’s office where medical transcription is done instantly and securely on a local device, with no risk of a data breach.

This isn’t science fiction. This is happening right now. And it’s all thanks to small language models.

The Magic of Model Distillation

So how do we create these small, powerful models? One of the key techniques is called model distillation. It’s a bit like a master craftsman teaching an apprentice. You take a large, powerful “teacher” model and use it to train a smaller “student” model.

The student model learns to mimic the teacher’s output, but with a fraction of the parameters. It’s a way of compressing all the knowledge and capabilities of a massive model into a much smaller, more efficient package.

I’ve seen this firsthand with some of my portfolio companies. They’re using model distillation to create incredible products that would have been impossible just a few years ago. They’re building everything from on-device translation apps to real-time code completion tools, all powered by small, distilled models.

The Bottom Line: Efficiency and Accessibility

At the end of the day, it all comes down to two things: efficiency and accessibility.

Small language models are more efficient. They’re faster, cheaper, and require less energy to run. In a world where we’re all trying to reduce our carbon footprint and build more sustainable businesses, this is a massive advantage.

And they’re more accessible. They lower the barrier to entry for building AI-powered products. You don’t need to be a giant corporation with a massive budget to get in the game. You can be a small startup, a solo developer, or even a student with a laptop. If you have a good idea and the right skills, you can build something amazing.

This is why I’m so bullish on small language models. They’re not just a technical curiosity. They’re a democratizing force. They’re putting the power of AI into the hands of more people than ever before.

So next time you read a headline about the latest, greatest, biggest language model, take it with a grain of salt. The real revolution isn’t happening in the cloud. It’s happening on the edge, on our devices, in our pockets. And it’s being powered by the small, the smart, and the efficient.

This is where I’m placing my bets. And I’m confident it’s a bet that will pay off.

Frequently Asked Questions

How can I apply this thinking to my own situation?

Start by identifying the core principle behind the opinion, not the specific example. Then ask yourself: does this principle apply to my context? If yes, test it in a small, low-risk way before going all in.

What's the most common pushback you get on this?

People often push back by citing exceptions or edge cases. And they're usually right that exceptions exist. But building a strategy around exceptions rather than patterns is a losing game for most founders.

What experience informs this perspective?

This perspective comes from over a decade of building companies in Silicon Valley, two successful exits (RemoteTeam to Gusto, MovieLaLa to Gfycat), and investing in 200+ startups including Anthropic, OpenAI, and Scale AI. I write about what I've lived.

Do all experts agree with this view?

No, and that's fine. The best ideas in business are often contrarian. I share my perspective based on my experience and data, but I encourage you to seek out opposing viewpoints and form your own conclusions.

More in Large Language Models

All Large Language Models articles · Sahin's angel investments · Startups he founded