Nobody Talks About the Data Moat Problem in AI Diagnostics—Here's How to Solve It

Published 2026-03-10 · Updated 2026-05-23 · 5 min read · AI in Healthcare · By Sahin Boydas

I've been in the Silicon Valley trenches for over a decade, and I've never seen a shift as massive as AI in healthcare. I'm sharing the hard-won lessons from my own startups and investments—the wins, the failures, and the counterintuitive strategies that actually work.

Everyone in Silicon Valley is chasing the same thing: the “moat.” That magical, defensible advantage that keeps competitors at bay. For the last five years, the consensus in AI diagnostics has been that the ultimate moat is data. A massive, proprietary dataset of medical images or records that nobody else can get.

I’ve seen hundreds of pitch decks that say the same thing. “We have 10 million X-rays.” “We have an exclusive partnership with a hospital for their pathology slides.” Founders and VCs alike have worshipped at the altar of big data. I’m here to tell you it’s mostly a fantasy.

I’ve built and sold two companies, one to Gusto and another to Gfycat. I’ve put my own money into over 200 startups, including AI pioneers like Anthropic, OpenAI, and Scale AI. I’ve seen the inside of what works and what gets burned. And in healthcare AI, the idea of a static data moat is a trap that will kill your startup.

The Myth of the Data Pile

Most founders think building a great AI model is the main event. They’re wrong. They spend years and millions of dollars assembling a giant dataset, thinking it’s their golden ticket. They believe that once they have more data than anyone else, they’ve won.

Here’s the uncomfortable truth: your giant pile of data is probably worth a lot less than you think. In the brutal, regulated world of healthcare AI, it’s not the size of your dataset that matters. It’s the quality, the cost of annotation, and your ability to navigate the regulatory labyrinth.

I remember meeting a team a few years back. Brilliant PhDs, a fantastic algorithm for detecting a rare type of cancer from MRI scans. They had a deal with a major university hospital and had amassed over 500,000 scans. They’d raised $10 million and were convinced they were untouchable. Twelve months later, they were nearly bankrupt. Why? The cost of getting those scans labeled by expert radiologists was astronomical—we’re talking $50 to $100 per scan. Their data was also noisy, full of edge cases and inconsistencies from different machines and protocols. Their model worked great on data from that one hospital, but it failed miserably when they tried to use it on data from a different health system. The FDA process was a nightmare they never even got to.

Their “moat” was a swamp. And they drowned in it.

The Real Moat is a Data Engine

The winning strategy isn’t to build a static pile of data. It’s to build a Data Engine. It’s not a noun; it’s a verb. It’s a system, a process, a machine that continuously and cheaply acquires, cleans, and annotates high-quality data that is purpose-built for the regulatory gauntlet.

Think of it like this: a data moat is a stagnant lake. A Data Engine is a powerful, ever-flowing river. The river constantly brings in new water (data), purifies it (cleans and annotates), and generates power (the model and its regulatory approval).

This is the single biggest lesson I’ve learned from watching companies like Scale AI succeed. They didn’t just sell data; they built the infrastructure for creating and managing it. In healthcare, this is even more critical. So, how do you build this engine?

Step 1: Integrate into the Clinical Workflow

Stop trying to buy data. It’s too expensive and it’s not tailored to your needs. The smartest companies I’ve seen build tools that clinicians—radiologists, pathologists, oncologists—actually want to use in their daily work.

Instead of building a standalone diagnostic tool, build a better viewing software. Build a reporting tool that saves them 30 minutes a day. Build something that makes their miserable, click-heavy workflow a little less painful. And as they use your tool, you get access to a clean, steady stream of data as a natural byproduct. The data acquisition cost plummets to near zero.

This is how you get your foot in the door. You’re not selling “AI” to a skeptical hospital administrator. You’re selling a workflow tool that saves them time and money. The AI is the Trojan horse.

Step 2: Build a Human-in-the-Loop Annotation System

Fully automated annotation is a pipe dream. The edge cases in medicine will kill you. At the same time, paying board-certified specialists to label every pixel is a recipe for bankruptcy.

The solution is a human-in-the-loop system. Use your AI model to do the first pass at labeling. Let it highlight potential areas of interest and make a preliminary diagnosis. Then, have a human expert quickly review and correct it. Your AI does 90% of the grunt work, and the expensive human expert just does the final 10% validation. This makes your annotation process 10x faster and cheaper.

This system is the core of your Data Engine. It’s a flywheel. The more data you process, the better your AI gets. The better your AI gets, the faster your human experts become. The faster they are, the more data you can process. This is how you build a cost advantage that competitors can’t touch.

Step 3: Design for the Regulatory Flywheel

Most founders treat the FDA like an afterthought. They build a model, then they go looking for a regulatory consultant. This is backward. You need to design your entire company, from the product to the data engine itself, around the regulatory process.

Start with an extremely narrow and specific use case. Don’t try to build a model that detects 50 different diseases. Build one that does one simple thing incredibly well, like measuring the size of a specific type of tumor. This is called a “de novo” submission.

Why? Because it’s a much easier path to your first FDA clearance. That first clearance is everything. It proves to the market, to investors, and to your own team that you can navigate the system. It’s a massive de-risking event.

Once you have that first clearance, you can start expanding. You can add a second feature, then a third. Each new clearance builds on the last one, creating a regulatory flywheel. Your initial product becomes a platform. By the time your competitors are just starting their first FDA submission, you’re on your fourth or fifth. That’s a real moat.

This Isn't a Tech Problem

Here’s the final, hard truth. Building a successful AI diagnostics company isn’t a machine learning problem. It’s an operations, sales, and regulatory problem. It requires a founder who is willing to live in the messy, frustrating reality of the healthcare system.

I see too many technical founders who just want to sit in a lab and tweak their models. That’s not the job. The job is getting on a plane and talking to doctors. It’s understanding the billing codes. It’s hiring a team that knows how to run a clinical trial.

So forget the fantasy of the data moat. It’s a distraction. Focus on building a machine. A machine that integrates into the workflow, produces clean data cheaply, and is designed from day one to conquer the regulatory beast.

That’s the only moat that matters.

Frequently Asked Questions

Do all experts agree with this view?

No, and that's fine. The best ideas in business are often contrarian. I share my perspective based on my experience and data, but I encourage you to seek out opposing viewpoints and form your own conclusions.

How has this view evolved over time?

My thinking on most topics has changed significantly over the years. Early in my career, I held many conventional views that experience proved wrong. I try to update my beliefs when the evidence changes.

What experience informs this perspective?

This perspective comes from over a decade of building companies in Silicon Valley, two successful exits (RemoteTeam to Gusto, MovieLaLa to Gfycat), and investing in 200+ startups including Anthropic, OpenAI, and Scale AI. I write about what I've lived.

What's the most common pushback you get on this?

People often push back by citing exceptions or edge cases. And they're usually right that exceptions exist. But building a strategy around exceptions rather than patterns is a losing game for most founders.

More in AI in Healthcare

All AI in Healthcare articles · Sahin's angel investments · Startups he founded