I’ve seen it happen more times than I can count. A company invests millions in a new AI system, expecting it to revolutionize their business. They set it up, feed it data, and wait for the magic to happen. But when it comes time to review the AI’s performance, the results are… underwhelming. In fact, data shows that a staggering 82% of companies fail to see a significant return on their AI investments. That’s not just a statistic to me; I’ve had a front-row seat to this disaster movie for years.
As the founder of RemoteTeam, which was acquired by Gusto, I oversaw $8M in remote team payroll. I’ve also been an angel investor in over 200 companies, including some of the biggest names in AI like Anthropic, OpenAI, and Scale AI. I’ve seen the good, the bad, and the ugly when it comes to implementing and evaluating AI. And I can tell you this: most companies are doing it wrong.
They treat AI performance reviews like a traditional employee review, using vague metrics and subjective feedback. Or they go to the other extreme, relying solely on automated outputs without any human oversight. Both approaches are doomed to fail. Here’s what really goes wrong—and how to fix it.
The “Garbage In, Garbage Out” Problem
The oldest saying in computer science is still the most relevant. If you feed your AI system bad data, you’re going to get bad results. It’s that simple. I remember a time at RemoteTeam when we were developing an AI to help with project management. We thought we had a brilliant system. It was tracking deadlines, monitoring progress, and flagging potential issues. But the initial performance review showed that it was actually slowing down our team.
We were baffled. The AI was doing everything we asked it to. The problem wasn’t the AI; it was our instructions. We were tracking the wrong metrics. We were so focused on deadlines and progress that we forgot to account for the complexity and interdependence of tasks. The AI was flagging things as “late” that were actually just waiting on another, more critical task to be completed. We were feeding it garbage, and it was spitting garbage back out.
To avoid this, you need to be crystal clear about what you want your AI to achieve. Don’t just throw data at it and hope for the best. Define specific, measurable goals. And make sure your data is clean, relevant, and unbiased. This requires a lot of upfront work, but it’s the only way to get meaningful results.
The “Black Box” Fallacy
Another huge mistake I see is leaders treating AI as a “black box.” They don’t understand how it works, and they don’t think they need to. They just want to see the results. This is a dangerous mindset. If you don’t understand the underlying logic of your AI, you can’t possibly evaluate its performance accurately.
I was once considering an investment in a promising young startup. They had a slick presentation and some impressive-looking data. But when I asked the founders to explain the core algorithm of their AI, they couldn’t do it. They just kept saying, “It’s very complex, but it works.” I passed on the investment. If the founders don’t understand their own product, how can they expect to lead their company to success?
As a leader, you don’t need to be a machine learning expert. But you do need to have a fundamental understanding of how your AI systems work. You need to know what data they’re using, what assumptions they’re making, and what their limitations are. This is the only way to have a real conversation about performance and to make informed decisions about how to improve it.
The Human Element is Missing
Perhaps the biggest mistake of all is forgetting the human element. AI is a tool. It’s meant to augment human intelligence, not replace it. But so many companies get caught up in the technology that they forget about the people who have to use it.
One of my most successful investments has been in a company that builds AI-powered tools for writers. Their secret? They didn’t try to create an AI that could write better than a human. Instead, they built a tool that helps humans write better. It suggests ideas, catches grammatical errors, and provides feedback on tone and style. It’s a collaborator, not a competitor.
When you’re evaluating your AI’s performance, don’t just look at the numbers. Talk to the people who are working with it every day. How is it affecting their workflow? Is it making their jobs easier or harder? Is it helping them to be more creative and productive? The answers to these questions are just as important as any automated metric.
Stop the Insanity
Let’s be honest: most AI performance reviews are a waste of time. They’re based on flawed data, a lack of understanding, and a disregard for the human element. But it doesn’t have to be this way. By focusing on clear goals, promoting transparency, and putting people first, you can turn your AI performance reviews into a powerful tool for driving real business results.
It’s not rocket science. It’s just good management. So stop making these common mistakes and start getting the most out of your AI investments. The future of your company may depend on it.
Frequently Asked Questions
What experience informs this perspective?
This perspective comes from over a decade of building companies in Silicon Valley, two successful exits (RemoteTeam to Gusto, MovieLaLa to Gfycat), and investing in 200+ startups including Anthropic, OpenAI, and Scale AI. I write about what I've lived.
How can I apply this thinking to my own situation?
Start by identifying the core principle behind the opinion, not the specific example. Then ask yourself: does this principle apply to my context? If yes, test it in a small, low-risk way before going all in.
Do all experts agree with this view?
No, and that's fine. The best ideas in business are often contrarian. I share my perspective based on my experience and data, but I encourage you to seek out opposing viewpoints and form your own conclusions.
What's the most common pushback you get on this?
People often push back by citing exceptions or edge cases. And they're usually right that exceptions exist. But building a strategy around exceptions rather than patterns is a losing game for most founders.