When we were building RemoteTeam, why most founders get rlhf completely wrong (and fix it) nearly killed us before we figured it out.
I've seen countless startups burn through cash trying to implement RLHF without understanding the fundamentals. Here's the painful truth about why it fails and a simple framework for getting it right from day one.
What I've Learned From 110 Companies
After investing in 200+ startups and running two companies to successful exits, I've developed a pretty clear picture of what works with why most founders get rlhf completely wrong (and fix it).
The biggest misconception is that you need to most founders overthink this and underspend on execution. That's backwards. The companies that win are the ones that timing is everything in this game.
I remember sitting with the Anthropic team early on and discussing how they thought about why most founders get rlhf completely wrong (and fix it). Their approach was counterintuitive but brilliant.
Why Most Approaches Fail
Let me be direct: about 70% of the approaches I see to why most founders get rlhf completely wrong (and fix it) are fundamentally flawed. Not slightly off. Fundamentally flawed.
The root cause is usually one of three things:
- Copying what big companies do without understanding why they do it. What works for Google doesn't work for a 10-person startup.
- Over-engineering the solution when a simple approach would work better. I've seen teams spend six months building something that could have been done in two weeks.
- Ignoring the human element. Technology is the easy part. Getting people to actually use it is where the real challenge lives.
What I Tell Founders
When a founder in my portfolio asks me about why most founders get rlhf completely wrong (and fix it), I usually start with three questions:
- What's your timeline? Because the right approach for a company with 6 months of runway is very different from one with 3 years.
- What have you already tried? Most founders have tried something. Understanding what didn't work is often more valuable than knowing what might.
- Who on your team owns this? If the answer is "everyone" or "no one," that's your first problem to solve.
These questions seem simple but they reveal a lot about where a company actually stands.
This connects to broader themes around RLHF, small language models, DPO that I've been thinking about a lot lately.
Wrapping Up
I've shared a lot here, and I know it can feel overwhelming. But here's the thing about why most founders get rlhf completely wrong (and fix it): you don't need to get everything right on day one. You just need to get started and keep improving.
The founders in my portfolio who excel at why most founders get rlhf completely wrong (and fix it) share one trait: they're relentlessly practical. They don't chase perfection. They chase progress.
That's the mindset I'd encourage you to adopt. Start where you are. Use what you have. Do what you can. And keep pushing forward.
As always, I'm rooting for you.
Frequently Asked Questions
What experience informs this perspective?
This perspective comes from over a decade of building companies in Silicon Valley, two successful exits (RemoteTeam to Gusto, MovieLaLa to Gfycat), and investing in 200+ startups including Anthropic, OpenAI, and Scale AI. I write about what I've lived.
What's the most common pushback you get on this?
People often push back by citing exceptions or edge cases. And they're usually right that exceptions exist. But building a strategy around exceptions rather than patterns is a losing game for most founders.
How has this view evolved over time?
My thinking on most topics has changed significantly over the years. Early in my career, I held many conventional views that experience proved wrong. I try to update my beliefs when the evidence changes.