''' I remember one of the most confusing A/B test results I ever saw. At MovieLaLa, we were testing a new recommendation algorithm. The metrics were all over the place. Engagement was up, but retention was down. Some users loved it, others hated it. The data was telling us two different stories, and we had to make a decision that could make or break the company.
This is the kind of situation that founders of AI products face all the time. The standard A/B testing frameworks they teach you in school, or that you read about in blog posts, just don't cut it when you're dealing with the complexities of AI. Having spent years leading AI product teams at places like Google and Amazon, and investing in over 200 AI startups, I saw firsthand how the best in the world operate. They don't use the generic frameworks you read about online. I'm sharing the internal playbook we used to launch AI products that reached millions of users.
Why A/B Testing AI Products is a Different Beast
Before we dive into the playbook, let's get one thing straight: A/B testing for AI products is not the same as testing a button color change. The systems are dynamic, the user responses are more complex, and the data can be downright misleading. Here's why:
- Non-determinism: Many AI models have a degree of randomness. You can run the same test twice and get different results. This makes it hard to isolate the true impact of your changes.
- Feedback loops: AI products learn from user behavior. This creates feedback loops that can amplify small changes over time, making it difficult to interpret short-term test results.
- Novelty effect: Users might initially react positively to a new feature simply because it's new, not because it's actually better. This can lead to false positives in your A/B tests.
- Multiple, conflicting metrics: In AI products, you're often optimizing for multiple metrics at once. For example, you might want to increase engagement, but not at the expense of user satisfaction. These metrics can often be in conflict, making it hard to declare a clear winner.
The Founder's Playbook for Interpreting Confusing A/B Test Results
So, how do you navigate this minefield? Here's the playbook I've developed over the years. It's not a magic formula, but it's a systematic way to think through confusing A/B test results and make better decisions.
Step 1: Segmentation is Your Superpower
When your overall A/B test results are confusing, the first thing you should do is segment your users. Don't just look at the average user. Look at how different groups of users are reacting to the change. Here are some of the segments I always look at:
- New vs. returning users: New users might be more open to change, while returning users might be more resistant.
- Power users vs. casual users: Power users might be more sensitive to small changes in the product, while casual users might not even notice.
- Demographics: Depending on your product, you might want to segment by age, gender, location, or other demographic factors.
At RemoteTeam, we once tested a new feature that was a huge hit with our power users but actually hurt engagement for new users. If we had only looked at the overall numbers, we would have missed this crucial insight. By segmenting our users, we were able to roll out the feature to our power users while we worked on a better onboarding experience for new users.
Step 2: Hunt for the Novelty Effect
The novelty effect is one of the most common traps in A/B testing. It's the tendency for users to initially engage with a new feature simply because it's new. This can lead to a temporary spike in your metrics, which can be misleading. To identify the novelty effect, you need to run your tests for a longer period of time. I usually recommend at least two weeks, but for some products, you might need to run them for a month or more. If you see a big spike in engagement in the first few days, followed by a gradual decline, you're likely seeing the novelty effect.
Step 3: Go Beyond the Numbers - Watch Your Users
Metrics can only tell you what is happening. They can't tell you why it's happening. To get the full story, you need to combine your quantitative data with qualitative data. This means watching user sessions, running user interviews, and reading customer support tickets. At MovieLaLa, we once had a test where the metrics were completely flat. But when we watched the user sessions, we saw that users were getting stuck on a particular screen. This was a huge insight that we would have missed if we had only looked at the numbers.
Step 4: The "Time-Shift" Analysis
This is a technique I developed to get a better read on the long-term impact of a change. The idea is simple: instead of just comparing the treatment group to the control group during the test period, you also compare them to a cohort of users from before the test started. This allows you to control for seasonality and other external factors that might be affecting your results. It's a bit more work to set up, but it can give you a much more accurate picture of the true impact of your change.
A Real-World Example: The Recommendation Engine That Almost Wasn't
I want to end with a story from my time at MovieLaLa. We were testing a new recommendation algorithm that we had been working on for months. The team was convinced it was going to be a huge win. But when we launched the A/B test, the results were a disaster. Engagement was down, and users were complaining. The data was telling us to kill the project.
But I had a gut feeling that something was wrong. So, we dug deeper. We segmented our users and found that while the new algorithm was performing poorly for our casual users, it was a huge hit with our power users. We then watched the user sessions and saw that the casual users were getting overwhelmed by the new recommendations. They didn't know what to do with all the new choices.
Based on this insight, we decided to roll out the new algorithm to our power users only. We then went back to the drawing board and designed a new onboarding experience for the casual users. A few months later, we launched the new algorithm to all our users, and it was a huge success. It ended up being one of the key features that led to our acquisition by Gfycat.
Stop Chasing P-Values
The lesson here is that you can't just blindly follow the data. You need to be a detective. You need to dig deep, ask the right questions, and use a combination of quantitative and qualitative data to get to the truth. A/B testing is a powerful tool, but it's just one tool in your arsenal. Don't let it become a crutch. Stop chasing p-values and start thinking like a founder. ''')) HBox(children=(FloatProgress(value=0.0, max=1.0), HTML(value=
Frequently Asked Questions
Is this guide based on real experience?
Every recommendation in this guide comes from direct experience, either from building and selling my own companies, or from patterns I've observed across 200+ angel investments. I don't write about things I haven't personally tested.
How should I work through this guide?
Don't try to absorb everything in one sitting. Read through once to get the big picture, then go back and work through each section as it becomes relevant to your current challenges. Bookmark it and return to it regularly.
How often is this guide updated?
I revisit and update my guides regularly as I learn new things and as the market evolves. The core principles tend to stay stable, but specific tactics and tools get refreshed based on what's working right now.