How to Build Multimodal LLMs That Don't Suck (A Practical Guide)

Published 2024-10-25 · Updated 2026-05-23 · 6 min read · Large Language Models · By Sahin Boydas

Building multimodal LLMs is the next frontier, but most attempts are clunky and impractical. I'm sharing my playbook for creating seamless, intuitive multimodal experiences that users will actually love, based on my experience shipping three of them.

If you're a founder dealing with how to build multimodal llms that don't suck, stop what you're doing and read this. Seriously.

Building multimodal LLMs is the next frontier, but most attempts are clunky and impractical. I'm sharing my playbook for creating seamless, intuitive multimodal experiences that users will actually love, based on my experience shipping three of them.

The Counterintuitive Truth

Here's what surprised me most about how to build multimodal llms that don't suck: the best practitioners do less, not more.

When I was building MovieLaLa, we tried to do everything at once. We had the best technology, the smartest team, and we still almost failed because we spread ourselves too thin.

The lesson I took from that experience, and from watching hundreds of other companies, is that the market doesn't care about your roadmap. It sounds simple. It's incredibly hard to execute.

What I've Learned From 93 Companies

After investing in 200+ startups and running two companies to successful exits, I've developed a pretty clear picture of what works with how to build multimodal llms that don't suck.

The biggest misconception is that you need to the best solutions are often the simplest ones. That's backwards. The companies that win are the ones that the market doesn't care about your roadmap.

I remember sitting with the Anthropic team early on and discussing how they thought about how to build multimodal llms that don't suck. Their approach was counterintuitive but brilliant.

Real Talk: What Actually Matters

I'm going to cut through the noise and tell you what actually matters when it comes to how to build multimodal llms that don't suck.

First, execution speed beats perfection. Every time. I've never seen a company fail because they moved too fast on how to build multimodal llms that don't suck. I've seen plenty fail because they moved too slow.

Second, measure everything. If you can't measure it, you can't improve it. Set up tracking from day one, even if it's basic.

Third, talk to your users. This sounds obvious but you'd be amazed how many founders build their how to build multimodal llms that don't suck strategy in a vacuum. Get out of the building. Talk to real people.

This connects to broader themes around multimodal LLMs, on-device AI, small language models that I've been thinking about a lot lately.

Wrapping Up

I've shared a lot here, and I know it can feel overwhelming. But here's the thing about how to build multimodal llms that don't suck: you don't need to get everything right on day one. You just need to get started and keep improving.

The founders in my portfolio who excel at how to build multimodal llms that don't suck share one trait: they're relentlessly practical. They don't chase perfection. They chase progress.

That's the mindset I'd encourage you to adopt. Start where you are. Use what you have. Do what you can. And keep pushing forward.

As always, I'm rooting for you.

Frequently Asked Questions

How do I measure success with this approach?

Pick one or two metrics that directly tie to your goal and track them weekly. Vanity metrics like page views or follower counts rarely matter. Focus on metrics that reflect real engagement or revenue impact.

What are the most common mistakes when building multimodal llms that don't suck (a practical guide)?

The biggest mistake I see is overcomplicating things early on. Start with the simplest version that works, get real feedback, and iterate from there. Another common trap is copying what worked for someone else without understanding the context behind their decisions.

How long does it take to build multimodal llms that don't suck (a practical guide)?

The timeline varies depending on your starting point and resources. For most founders, expect 2-4 weeks for initial setup and 2-3 months to see meaningful results. I've seen teams move faster when they focus on one thing at a time rather than trying to do everything at once.

More in Large Language Models

All Large Language Models articles · Sahin's angel investments · Startups he founded