Confession: Our A/B Tests Were Useless Until We Understood This One AI Principle

Published 2025-05-21 · Updated 2026-05-23 · 7 min read · Product Management AI · By Sahin Boydas

I used to struggle with a/b testing ai, thinking I had it all figured out. It led to burnout and a failed product. But after years of painful lessons, I discovered a counterintuitive approach to AI product development that changed everything. It wasn't about the tech, but about this one simple shift in perspective.

I can still picture the empty office. The Aeron chairs were stacked in a corner, the whiteboards were wiped clean, and the only sound was the hum of the server rack we were about to shut down. We had just failed. Spectacularly.

We had spent 18 months and burned through $2 million in seed funding on a product I was sure would be a hit. We did everything by the book. We were agile. We were lean. And we were religious about A/B testing. We tested everything. Button colors, headline copy, signup flows, email subject lines. You name it, we had a spreadsheet with a p-value for it.

Our dashboards looked great. We’d celebrate a 4.7% lift in click-through rates like we’d just won the World Series. But our core metrics—the ones that actually mattered, like user retention and revenue—were flat. We were optimizing the leaves on a dying tree.

That failure was one of the most painful experiences of my career. It led to sleepless nights, serious burnout, and a long period of self-doubt. But it also forced me to confront a hard truth: in the age of AI, the old playbook for A/B testing is broken. We were running in circles, busy but not productive, because we were testing the wrong things.

The A/B Testing Theater

What we were doing wasn’t really product improvement. It was A/B testing theater. We were performing all the rituals of data-driven development, but it was a performance for our own benefit, to make us feel like we were in control. We were tweaking the UI, the surface layer of the product, while the real engine—the AI model making the core decisions—was left completely unexamined.

Think about it. If you have an AI-powered product, like a recommendation engine or a personalized news feed, the user experience isn’t primarily defined by the color of a button. It’s defined by the quality of the AI’s output. Are the recommendations good? Is the feed interesting? Is the AI helping the user achieve their goal?

Testing whether a green button gets more clicks than a blue one is irrelevant if the recommendation that the button leads to is garbage. You’re just measuring which color is better at tricking users into clicking on something they don’t want. It’s a local optimization that can actually hurt the global user experience.

We had dozens of statistically significant test results that told us we were making the product better, but our users were telling us a different story by simply not coming back. The data wasn’t lying, but we were asking it the wrong questions.

The One Principle: Test the Objective, Not the Interface

After that failure, I spent years advising and investing in other AI startups. I’ve seen over 200 companies up close, including foundational players like Anthropic and Scale AI. And I saw the same pattern over and over: teams getting stuck in the A/B testing theater, optimizing the superficial while ignoring the substantial.

The successful ones, the ones that broke through, all understood one counterintuitive principle:

You have to A/B test the AI’s objective function, not just the UI that presents its results.

This is the whole game. It’s the shift in perspective that changes everything. An AI model is just a machine that has been trained to optimize for a specific goal, its "objective function." For a recommendation engine, the objective might be to maximize the click-through rate. For a language model, it might be to predict the next word in a sentence.

Traditional A/B testing takes the AI’s output as a given and then tests the wrapper around it. Objective Function Testing goes a level deeper. It questions the very goal of the AI itself.

Instead of asking: "Does this headline design increase clicks?"

You should be asking: "Does an AI optimized for 'user engagement' create more long-term value than an AI optimized for 'ad clicks'?"

These are fundamentally different questions. The first is a UI tweak. The second is a strategic product decision.

A Real-World Example from MovieLaLa

At my second company, MovieLaLa (which we later sold to Gfycat), we built a discovery app for movie lovers. The core of the app was a recommendation engine. Initially, we defined the AI's objective as maximizing the number of movies a user added to their watchlist.

It seemed logical. More movies on the watchlist must mean users are finding more movies they want to see, right?

We ran UI tests on how the movie posters were displayed, the text on the "Add to Watchlist" button, everything. We got some small lifts. But growth stalled.

Then we tried Objective Function Testing. We developed a second version of our AI. This new model had a different objective: maximize the number of movies a user rated after watching.

This was a much harder goal. It required the AI to not just predict what a user thought they wanted to see, but what they would actually watch and have an opinion on. It was a proxy for user satisfaction, not just user intent.

We ran an A/B test on these two models. Not the UI—the models themselves. 50% of our users got recommendations from the "Watchlist" AI, and 50% got them from the "Rating" AI. The interface was identical for both groups.

The results were staggering.

The "Rating" AI group had a 30% higher 90-day retention rate. They watched more movies, they opened the app more often, and they invited more friends. We had been so focused on the superficial act of adding to a list that we missed the real goal: helping people find movies they’d love.

By changing the AI's core purpose, we transformed the product. That was a lesson I never forgot.

How to Implement Objective Function Testing

This isn’t just a theoretical concept. It’s a practical approach you can apply to your own AI products.

  1. Identify Your True North: First, step back from your immediate metrics like CTR or conversion rate. What is the ultimate value you are trying to create for your users? Is it saving them time? Helping them discover new things? Making them more successful at their job? This is your "true north" metric. It’s often harder to measure than a simple click, but it’s the only thing that matters in the long run. For MovieLaLa, it was user satisfaction, which we measured through ratings.

  2. Brainstorm Alternative Objectives: Once you have your true north, brainstorm different AI objective functions that could serve as proxies for it. If your true north is "user success," your objectives could be "task completion rate," "time to complete task," or "user-reported satisfaction."

  3. Build and Isolate the Models: This is the most engineering-intensive part. You need to build and train separate AI models for each objective function you want to test. It is critical that you can run these models in parallel and serve them to different user segments without them interfering with each other.

  4. Run the Experiment: Divide your users into segments and assign each segment to a different model. Keep the UI and all other variables exactly the same. The only difference between the groups should be the AI model they are interacting with.

  5. Measure What Matters: Measure the impact on your true north metric over a meaningful period. Don’t just look at the immediate results. Look at user behavior over weeks or even months. Are they retaining better? Are they more engaged? Are they upgrading? This is where you’ll find the real signal.

This is Hard, But It's Worth It

I’m not going to pretend this is easy. Objective Function Testing requires more upfront investment in data science and engineering than simply tweaking button colors. It requires you to think like a product strategist, not just a growth hacker.

But the payoff is immense. Instead of getting trapped in the local optima of UI tweaks, you start making big, meaningful leaps in product quality. You stop optimizing for vanity metrics and start optimizing for real user value.

Looking back at my first failed startup, it’s painfully obvious what we did wrong. We were so obsessed with the performance of A/B testing that we never stopped to question what we were testing in the first place. We had a powerful AI engine, but we had pointed it in the wrong direction.

Don’t make the same mistake I did. Stop the A/B testing theater. Go deeper. Question the fundamental assumptions of your AI. Test the objective. That’s where you’ll find the breakthroughs that create legendary products.

Frequently Asked Questions

What experience informs this perspective?

This perspective comes from over a decade of building companies in Silicon Valley, two successful exits (RemoteTeam to Gusto, MovieLaLa to Gfycat), and investing in 200+ startups including Anthropic, OpenAI, and Scale AI. I write about what I've lived.

How has this view evolved over time?

My thinking on most topics has changed significantly over the years. Early in my career, I held many conventional views that experience proved wrong. I try to update my beliefs when the evidence changes.

Do all experts agree with this view?

No, and that's fine. The best ideas in business are often contrarian. I share my perspective based on my experience and data, but I encourage you to seek out opposing viewpoints and form your own conclusions.

More in Product Management AI

  • How to Hire Your First AI Product Manager (And What to Look For) — Having spent years leading AI product teams at places like Google and Amazon, I saw firsthand how the best in the world operate. They don't use the generic frameworks you read about online. I'm sharing the internal playbook we used to launch AI products that reached millions of users.
  • How to Use AI to Find Your Product's 'Aha!' Moment — Everyone in the AI space follows the same tired advice. We decided to question it. After analyzing over 1,000 AI product failures, we found a shocking pattern that conventional wisdom completely misses. The data points to one uncomfortable truth about why most AI products never find traction.
  • Why I Killed Our Most Popular AI Feature (And What Happened Next) — I used to struggle with feature prioritization ai, thinking I had it all figured out. It led to burnout and a failed product. But after years of painful lessons, I discovered a counterintuitive approach to AI product development that changed everything. It wasn't about the tech, but about this one simple shift in perspective.
  • The Former Google PM's Playbook for AI Product-Market Fit — I didn't go to business school. I learned how to build a multi-million dollar AI company from the trenches. After countless mistakes and a few lucky breaks, I've distilled my experience into these 7 hard-won lessons. This is the stuff they don't teach you in books.
  • We Threw Out Our AI Roadmap After One Painful User Interview — Everyone in the AI space follows the same tired advice. We decided to question it. After analyzing over 1,000 AI product failures, we found a shocking pattern that conventional wisdom completely misses. The data points to one uncomfortable truth about why most AI products never find traction.
  • I Thought We Had Product-Market Fit. I Was Dangerously Wrong. — Building an AI startup is anything but glamorous. I want to take you behind the curtain and share the unfiltered reality of our journey. From the heated debates over our roadmap to the bug that almost derailed our launch, this is the real story of what it takes to build and ship an AI product.

All Product Management AI articles · Sahin's angel investments · Startups he founded