I’ve seen more AI pitches than I can count. As an investor in over 200 startups, including foundational companies like Anthropic, OpenAI, and Scale AI, my inbox is a constant stream of “groundbreaking” models and “revolutionary” algorithms. Everyone wants to show me their shiny new toy.
But here’s a secret. After two successful exits and years of watching companies either take off or nosedive, I’ve learned that the tech itself is rarely the deciding factor. Especially in AI.
This got me thinking. What really separates the winners from the losers in the AI race? Is it the team with the most Kaggle grandmasters? The one with the biggest compute budget? The fanciest new model architecture?
We decided to find out. My team and I spent the last six months digging into the data from over 500 AI teams—some from my own portfolio, some from major tech companies, and others from scrappy startups. We looked at their projects, their team composition, their workflows, and their outcomes.
The results were not what we expected. There was one factor, above all others, that consistently predicted success. And it’s something most teams completely ignore.
It’s not the model. It’s the data.
More specifically, the single most important factor is a relentless, obsessive, and systematic approach to improving data quality.
The Great Deception: Why We All Chase Models
It’s easy to fall in love with models. They’re the glamorous part of AI. Reading a new paper on a novel transformer architecture feels like you’re touching the future. It’s exciting. It’s what gets written about in the press.
I get it. When I was building my first company, MovieLaLa, we were obsessed with having the best recommendation engine. We’d spend weeks trying to tweak the algorithm, hoping for a small percentage point improvement in our metrics. We thought a better model was the key to unlocking growth.
This is what I call the “Model-Centric Trap.” It’s the belief that the path to success is paved with more complex algorithms and bigger neural networks. Teams caught in this trap spend 80% of their time on the model and 20% on the data. They’re constantly chasing the state-of-the-art, hoping the next big architecture will solve all their problems.
I saw a startup last year that had raised $10 million to build an AI for legal contract analysis. They had a team of brilliant PhDs from top universities. They spent a year and a fortune in GPU costs trying to build a massive, custom language model from scratch. But the product was consistently unreliable. Why? Their training data was a mess. It was full of inconsistencies, mislabeled examples, and noise. They were trying to build a skyscraper on a foundation of sand.
They focused on the 1% of the problem (the model) and ignored the 99% (the data).
The Unsexy Truth: Success is Built on Better Data
The successful teams we analyzed did the exact opposite. They were data-centric. They understood that the model is a commodity, but the data is the real moat. These teams spend 80% of their time on the data and 20% on the model.
What does a data-centric approach look like in practice? It’s not just about having a lot of data. It’s about having the right data, and systematically making it better.
Here’s what the winning teams do differently:
They Treat Data Like a Product: Data isn’t just a file you feed into a model. It’s a product in itself. It needs a product manager. It needs a roadmap. It needs versioning. The best teams have engineers whose entire job is to build tools and workflows for labeling, cleaning, augmenting, and managing data. They build a “data factory.”
They Obsess Over Labeling Consistency: In one of my portfolio companies, they were building a system to detect manufacturing defects. The model’s performance was stuck at around 80% accuracy, which wasn’t good enough to be useful. The team spent a month tweaking the model with no luck. Finally, a new engineer joined and decided to audit the training data. He found that the labels were a disaster. Two different labelers would look at the same image of a tiny scratch and call one a “defect” and the other “normal.” There was no clear, objective definition. They paused all model development and spent two weeks creating a detailed labeling guide with dozens of examples. They retrained the labelers. Then they retrained the same, simple model on the newly consistent data. The accuracy jumped to 99.5%. Problem solved.
They Embrace Active Learning: Instead of labeling millions of examples at random, successful teams are smart about it. They use the model itself to find the most valuable data to label next. This is called active learning. The model essentially says, “Here are the 1,000 examples I’m most confused about. If you label these for me, I’ll learn the most.” This creates a powerful feedback loop where the model gets smarter with a fraction of the data.
They Focus on the Long Tail: Real-world data is messy. There are always rare, unexpected edge cases. A data-centric team doesn’t see these as an annoyance; they see them as an opportunity. They build systems to actively find and fix the model’s mistakes on these “long tail” examples. This is how you get from a model that works 95% of the time to one that you can actually trust in production.
How to Become a Data-Centric Leader
Shifting from a model-centric to a data-centric culture is one of the hardest—and most important—things a leader can do in an AI company. It requires a fundamental change in mindset.
First, you have to change what you celebrate. Stop praising the team for shipping a new model architecture. Instead, celebrate them for improving the data. Did someone find and fix a major labeling inconsistency? Did they build a tool that doubled the speed of data cleaning? Shout it from the rooftops. Make heroes out of the data engineers, not just the research scientists.
Second, change how you allocate resources. Your data infrastructure is not a cost center; it’s your most important R&D investment. Are you spending more on GPUs than on your labeling and data management tools? If so, you’re probably doing it wrong. I’ve advised companies to literally cut their compute budget in half and reinvest that money into their data pipeline. The ROI is almost always higher.
Third, get your AI team as close to the customer as possible. The people building the model need to feel the pain of its failures. At RemoteTeam, which was acquired by Gusto, we had our engineers sit in on customer support calls. When they heard a customer complain about a bug, it wasn’t an abstract problem anymore. It was real. This direct feedback is the fastest way to identify the weaknesses in your data and prioritize what to fix.
I’m not saying models don’t matter at all. Of course they do. But in 2025, we have an abundance of great models. What we have a scarcity of is high-quality, well-curated data and the teams who know how to build with it.
Stop chasing the shiny new algorithm. The boring, unsexy, and incredibly powerful truth is that the path to building successful AI products is paved with better data. Focus your team there, and you’ll leave the model-chasers in the dust.
Frequently Asked Questions
How has this view evolved over time?
My thinking on most topics has changed significantly over the years. Early in my career, I held many conventional views that experience proved wrong. I try to update my beliefs when the evidence changes.
What experience informs this perspective?
This perspective comes from over a decade of building companies in Silicon Valley, two successful exits (RemoteTeam to Gusto, MovieLaLa to Gfycat), and investing in 200+ startups including Anthropic, OpenAI, and Scale AI. I write about what I've lived.
Do all experts agree with this view?
No, and that's fine. The best ideas in business are often contrarian. I share my perspective based on my experience and data, but I encourage you to seek out opposing viewpoints and form your own conclusions.
What's the most common pushback you get on this?
People often push back by citing exceptions or edge cases. And they're usually right that exceptions exist. But building a strategy around exceptions rather than patterns is a losing game for most founders.