I Wasted 5 Years on AI Ethics Frameworks. Here's What Actually Works.

Published 2025-06-13 · Updated 2026-05-23 · 5 min read · AI Ethics and Regulation · By Sahin Boydas

I chased complex AI ethics frameworks for half a decade, getting it all wrong. I'm sharing my painful journey from buzzword-chasing to building responsible AI that ships. This is the stuff nobody tells you about the gap between theory and reality.

I’m going to tell you about my biggest professional failure. It wasn’t a startup that crashed and burned. It wasn’t a missed investment. It was spending five years of my life chasing a ghost.

That ghost was the perfect, all-encompassing AI ethics framework. I thought if I could just find the right set of rules, the right academic paper, the right 100-page PDF, I could solve the problem of building responsible AI. I was wrong. Dead wrong.

For years, I was convinced I was on the right track. At my last company, RemoteTeam, we spent months trying to build a comprehensive “Ethical AI” governance structure. We had committees. We had documents. We had endless meetings discussing fairness, accountability, and transparency. We brought in consultants who charged us a fortune to tell us things we already knew.

And what did we ship? A lot of documents. Our engineers, the people actually building the products, saw it as a tax. A bureaucratic hurdle to be cleared, not a tool to be used. The documents got filed away in a Google Drive folder that nobody ever opened again. The gap between our beautiful, theoretical framework and the reality of shipping code was a canyon.

I saw the same pattern play out after we were acquired by Gusto, and I saw it again in the dozens of AI companies I invested in. Everyone was talking about ethics, but it felt like a performance. A pantomime to reassure board members and regulators. The real work of building safe and reliable AI wasn't happening in those committee meetings.

My turning point came during a board meeting for a portfolio company. They were building a diagnostic tool using computer vision. The CEO presented a beautiful, 50-slide deck on their ethics framework. It was full of buzzwords—value alignment, stakeholder engagement, algorithmic accountability. Then a junior engineer in the room, probably terrified, raised his hand and asked, “But what do I do when the model shows a 15% higher error rate for patients with darker skin tones?”

Silence. The CEO, for all his talk, had no answer. The framework had no answer. It was all just words.

That’s when I realized I had wasted five years. I had been focusing on the wrong thing entirely. I was chasing the appearance of doing good, not the practice of building good things. So I threw it all out. I started from scratch, and I focused on one question: What actually works?

Here’s what I learned.

Stop Writing Novels, Start Writing Tests

The single biggest problem with most AI ethics frameworks is that they are not testable. They are collections of noble-sounding principles like “Be Fair” or “Promote Human Flourishing.” How do you write a unit test for “Promote Human Flourishing”? You can’t.

Instead of abstract principles, you need simple, falsifiable rules that an engineer can actually test. A rule isn’t “The model should be fair.” A rule is, “The model’s loan approval rate for applicants from postal code 90210 shall not be more than 5% higher than for applicants from postal code 90211, assuming all other variables are equal.”

That’s a rule you can build a test for. It either passes or it fails. It’s binary. It’s clear. It forces you to define what you mean by “fair” in a concrete, measurable way.

  • Bad: Our AI will not perpetuate historical biases.
  • Good: We will audit our training data for demographic representation and re-weight the data if any group is underrepresented by more than 10%.
  • Good: The model’s performance on our internal test set must show less than a 1% variance in accuracy across all identified demographic groups.

This is how you bridge the gap between theory and reality. You turn your principles into code.

Your AI Is Only as Good as Your Data

We love to talk about fancy algorithms and model architectures. We spend far less time on the boring, unglamorous work of data hygiene. This is a huge mistake. All the ethical frameworks in the world can’t save a model trained on biased, garbage data.

For one of my investments, a hiring AI startup, the model was supposed to identify top engineering candidates. But it kept ranking candidates from certain universities much higher than others. Why? Because the initial training data was a decade of the founders’ own hiring decisions, and they had a strong preference for their alma maters. The model wasn’t discovering a secret truth about talent; it was just laundering their personal bias.

The solution wasn’t an ethics committee. The solution was to go back to the data. They spent months painstakingly building a new dataset, ensuring it represented a wide range of backgrounds, universities, and career paths. It was expensive and time-consuming. But it worked. The model became dramatically more accurate and, more importantly, fairer.

If you’re not spending at least half your time on data sourcing, cleaning, and auditing, you’re not serious about responsible AI. You’re just playing pretend.

Red Teaming Is Not Optional

For years, the security world has known that the only way to build a secure system is to pay people to try and break it. We call this penetration testing or red teaming. For some reason, the AI world has been slow to adopt this mindset.

We build our models, test them on a clean, sanitized evaluation set, and are shocked when they fall apart in the real world. You have to assume your model will be abused. You have to actively try to make it fail. That’s how you find the weaknesses before your users do.

I’m an investor in Anthropic, and one of the things that impressed me most was their early and deep commitment to red teaming. They are constantly attacking their own models, trying to get them to generate harmful content, reveal private information, or go off the rails. They don’t just wait for problems to emerge; they hunt for them.

This needs to be a standard part of the development lifecycle for any serious AI product. What happens if a user inputs gibberish? What if they try a prompt injection attack? What are the weird edge cases you haven’t thought of? You need a dedicated team whose only job is to break things. Their success is your success.

Governance Is About People, Not Paper

Finally, even with testable rules, clean data, and aggressive red teaming, you need clear lines of responsibility. This is where the EU AI Act is actually getting something right. It’s forcing companies to ask a simple question: Who is on the hook if something goes wrong?

If your self-driving car model has a critical failure, who is responsible? The engineer who wrote the code? Their manager? The VP of Engineering? The CEO? If you don’t have a clear answer to that question, you don’t have a governance plan.

A real governance plan is a simple document. It maps specific risks to specific people. For example:

Risk Category Description Owner Mitigation / Test
Model Bias Model shows statistically significant performance degradation for a protected class. VP of Engineering Run demographic bias audit (Test #2.1) before every deployment.
Harmful Content Model generates unsafe, violent, or hateful content. Head of Safety Red team attack suite (Test #4.5) must pass with 99.9% success.
Data Privacy Model leaks personally identifiable information (PII) from the training set. Chief Legal Officer Run PII detection scan on training data and model outputs (Test #3.2).

This isn’t a 100-page document. It’s a spreadsheet. But it’s a hundred times more useful than a vague ethics manifesto because it’s actionable. It creates accountability.

The Only Thing That Matters

I threw away five years of my career chasing complexity and academic theories. I learned the hard way that when it comes to building responsible AI, the only thing that matters is what you do, not what you say.

Stop talking about fairness and start writing tests for it. Stop admiring your algorithm and start auditing your data. Stop hoping for the best and start paying people to find the worst. And for God’s sake, stop writing documents nobody reads.

Build a simple, testable, and accountable process. That’s it. That’s the framework that actually works. It’s not as glamorous as a 50-page ethics charter, but it has one big advantage: it might actually prevent a disaster.

Frequently Asked Questions

Do all experts agree with this view?

No, and that's fine. The best ideas in business are often contrarian. I share my perspective based on my experience and data, but I encourage you to seek out opposing viewpoints and form your own conclusions.

What's the most common pushback you get on this?

People often push back by citing exceptions or edge cases. And they're usually right that exceptions exist. But building a strategy around exceptions rather than patterns is a losing game for most founders.

How has this view evolved over time?

My thinking on most topics has changed significantly over the years. Early in my career, I held many conventional views that experience proved wrong. I try to update my beliefs when the evidence changes.

How can I apply this thinking to my own situation?

Start by identifying the core principle behind the opinion, not the specific example. Then ask yourself: does this principle apply to my context? If yes, test it in a small, low-risk way before going all in.

More in AI Ethics and Regulation

  • AI Regulation in 2027: 3 Predictions From a Serial Entrepreneur — Having lived through the dot-com bust, the mobile revolution, and now the AI explosion, I've learned to see around corners. The current AI regulation is just the beginning. I'm sharing my 3 bold predictions for the 2027 regulatory landscape and how to prepare now.
  • How to Conduct an AI Alignment Audit (The Counterintuitive Guide) — Forget the standard AI alignment checklists. They don't work. After auditing dozens of models, I've developed a counterintuitive method that actually surfaces deep alignment issues. I'll walk you through my exact 3-step process for finding what others miss.
  • The Truth About AI Bias: 7 Shocking Stats from Our 2026 Audit — We just completed a massive audit of 100+ production AI models, and the results on bias are staggering. I'm pulling back the curtain on the real numbers—not the sanitized corporate reports. This is what hidden bias actually looks like in the wild.
  • Nobody Talks About the Real Cost of AI Safety. Until Now. — As a Silicon Valley veteran who has built and sold two AI companies, I'm breaking the code of silence. The true cost of implementing robust AI safety isn't in the tech—it's in the human capital and culture. I'll reveal the numbers and strategies you need to know.
  • The Truth About AI Bias: 7 Shocking Stats from Our 2026 Audit — We just completed a massive audit of 100+ production AI models, and the results on bias are staggering. I'm pulling back the curtain on the real numbers—not the sanitized corporate reports. This is what hidden bias actually looks like in the wild.
  • I Wasted 5 Years on AI Ethics Frameworks. Here's What Actually Works. — I chased complex AI ethics frameworks for half a decade, getting it all wrong. I'm sharing my painful journey from buzzword-chasing to building responsible AI that ships. This is the stuff nobody tells you about the gap between theory and reality.

All AI Ethics and Regulation articles · Sahin's angel investments · Startups he founded