136 Counterintuitive Lessons I Learned About computational journalism

Published 2026-01-23 · Updated 2026-05-23 · 5 min read · AI for Creators · By Sahin Boydas

Let's get real about computational journalism. It's not about fancy tools or big budgets. I’m breaking down the fundamental principles that took me from a struggling creator to a recognized expert in the field.

I remember sitting in a cramped office in Palo Alto back in 2014, staring at a screen full of unstructured data. We were trying to figure out how to make MovieLaLa work. The data was an absolute mess. The tools were primitive. And I realized something that changed my entire approach to building companies: the best stories aren't written by humans alone, and they aren't written by machines alone. They are discovered in the friction between the two.

Forget everything you think you know about computational journalism. The strategies that worked last year are now obsolete. I see founders and creators making the same mistakes over and over again. They buy expensive software. They hire massive teams. They think throwing money at a problem will magically produce insights. It won't.

I've built and sold two companies. RemoteTeam went to Gusto. MovieLaLa went to Gfycat. I've written checks to over 200 startups, including Anthropic, OpenAI, Scale AI, and Hugging Face. I wrote "Becoming Top 1%" because I wanted to document what actually works. And what works in computational journalism is not what they teach you in journalism school or computer science classes.

Over the past decade, I've compiled exactly 136 lessons about the intersection of data, code, and storytelling. Most of them are painful mistakes I made so you don't have to. Today, I am breaking down the fundamental principles that took me from a struggling creator to a recognized expert in the field. Here is the one pivot you need to make to stay ahead of the curve.

The MovieLaLa Days: Finding the Signal in the Noise

When we started MovieLaLa, we wanted to build a social network for movie fans. The problem was data. We had millions of data points about what people liked, what they shared, and what they ignored. But raw data is just noise. It doesn't tell a story.

I spent nights writing scripts to scrape, clean, and analyze this data. I thought the code would give me the answers. I was wrong. The code only gave me better questions.

This is the first major misconception about computational journalism. People think the computer does the journalism. It doesn't. The computer just does the heavy lifting. You still have to be the journalist. You still have to look at a spreadsheet with 50,000 rows and ask, "Why did this specific metric spike on a Tuesday?"

We noticed a weird trend in how teenagers were sharing movie trailers. The data showed a massive drop-off in engagement after the first five seconds. Traditional analytics would just say "shorten the video." But computational journalism requires you to dig deeper. We cross-referenced that data with social media sentiment analysis. We found out they weren't bored. They were switching apps to text their friends about the trailer. The drop-off wasn't a failure. It was a success metric we were misinterpreting.

That realization changed everything. We stopped trying to keep them in the app and started making it easier for them to share outside the app. Engagement skyrocketed. Gfycat eventually acquired us because we understood how to turn data into a narrative about user behavior. We didn't just report the numbers. We reported the human behavior driving the numbers.

RemoteTeam and the Data of Human Connection

Fast forward to RemoteTeam. We were building tools for remote workers before it was cool. When the pandemic hit, suddenly everyone was a remote worker. We had a mountain of data on how distributed teams operated.

I saw journalists writing think pieces about the "future of work" based on gut feelings and a few interviews. Meanwhile, we had the actual data. We knew exactly how many hours people were working, when they were taking breaks, and how communication patterns were shifting.

I decided to start publishing our findings. But I didn't just dump charts on a blog. I used computational techniques to find the human stories hidden in the server logs.

For example, we found a correlation between the number of emojis used in Slack and employee retention. Teams that used more emojis had lower turnover. It sounds silly, but the data was clear. I wrote a piece about it. It went viral. Why? Because it took a cold, hard data point and connected it to a deeply human experience: the need for connection in a digital workspace.

This is where most people fail at computational journalism. They get so obsessed with the code that they forget the reader. Your audience doesn't care about your Python script. They don't care about your SQL queries. They care about what the data means for their lives.

Gusto acquired RemoteTeam because we didn't just build software. We built a narrative around the data. We used computational journalism to position ourselves as the authority on remote work. We proved that you don't need a massive newsroom to do hard-hitting data journalism. You just need curiosity and a willingness to write some code.

Investing in the AI Wave: OpenAI, Anthropic, and Scale AI

My angel investing journey taught me even more. When I wrote my first checks into OpenAI and Anthropic, I wasn't just investing in smart people. I was investing in a fundamental shift in how we process information.

I remember talking to the founders of these companies. They understood something that most media executives still don't get. Language models are not just text generators. They are reasoning engines.

In the early days of computational journalism, we used code to count things. How many times did a politician say a specific word? How much money did a PAC spend in a specific county? It was basic math applied to text.

Now, with tools from OpenAI and Anthropic, we can do semantic analysis at scale. We can feed a model 10,000 city council meeting transcripts and ask it to identify patterns of corruption. We can analyze the tone of a million earnings calls to predict market movements.

But here is the catch. The models are getting commoditized. Everyone has access to GPT-4. Everyone can use Claude. If everyone has the same tools, the tools are no longer your competitive advantage.

I invested in Scale AI because I realized that the models are only as good as the data you feed them. High-quality, human-annotated data is the real bottleneck. In computational journalism, your proprietary dataset is your only true moat. If you are just analyzing the same public datasets as everyone else, you are going to lose. You need to build your own datasets. You need to scrape the weird corners of the internet. You need to file FOIA requests and digitize paper records. That is where the alpha is.

The Core Lessons of Computational Journalism

Out of the 136 lessons I've documented, a few stand out as absolute requirements for anyone trying to survive in this space. Let's break them down.

1. The algorithm is your intern, not your editor

I see creators treating AI like an oracle. They type a prompt, get an output, and publish it. This is a recipe for mediocrity. Treat your computational tools like a very fast, very literal intern. They can summarize 100 PDFs in ten seconds. They can write a basic Python script to scrape a website. But they cannot tell you if the story actually matters. You are the editor. You make the final call. You have to inject your own perspective and your own voice.

2. Clean data is a myth

If you wait for perfect data, you will never publish anything. Real-world data is messy. It has missing values. It has typos. It has biases. Your job is not to find perfect data. Your job is to understand the imperfections and account for them in your reporting. I spent three weeks trying to clean a dataset at MovieLaLa before I realized the errors were actually the most interesting part of the story. The anomalies are where the truth hides. Embrace the messiness.

3. If everyone has the same tools, taste is your only moat

I mentioned this earlier, but it bears repeating. When I look at the startups I've invested in, the ones that win aren't always the ones with the best tech. They are the ones with the best taste. In journalism, taste means knowing which questions to ask. It means knowing how to frame a narrative. A beautiful data visualization is worthless if it doesn't answer a compelling question. Develop your taste. Read widely. Study art. Understand human psychology. The code is just the plumbing.

4. Stop trying to automate empathy

You can automate data collection. You can automate data cleaning. You can even automate the drafting of basic reports. But you cannot automate empathy. The best computational journalism uses data to highlight human suffering, human joy, or human absurdity. If your story doesn't make the reader feel something, you have failed. I don't care how sophisticated your machine learning model is. The human element is non-negotiable.

5. Ship fast, but never lie with data

In the startup world, we say "move fast and break things." In journalism, if you break the truth, you are finished. You have to find a balance. You need to publish quickly to stay relevant, but you must be rigorously honest about what the data does and does not show. Never stretch a correlation to imply causation just to get a better headline. I've seen founders destroy their reputations overnight because they exaggerated their metrics. The same applies to writers. Your credibility is your most valuable asset.

6. Build your own tools

Don't rely entirely on off-the-shelf software. The best computational journalists build their own scrapers, their own analysis pipelines, and their own visualization libraries. When you build your own tools, you understand the assumptions baked into the code. You aren't flying blind. At RemoteTeam, we built custom dashboards that gave us insights no one else had. That custom tooling was a massive part of our valuation when Gusto acquired us.

7. Learn to write for two audiences

When you publish a piece of computational journalism, you are writing for two distinct groups. You are writing for the general public, who just want the story. And you are writing for the nerds, who want to see your methodology. You have to satisfy both. Write a compelling narrative for the public, but publish your code and your raw data on GitHub for the nerds. Transparency builds trust.

The Pivot You Need to Make Right Now

The strategies that worked last year are obsolete. Writing basic data-driven articles is no longer enough. The pivot you need to make is toward interactive storytelling.

Readers don't want to be talked at anymore. They want to explore the data themselves. They want to plug in their own variables and see how the story changes.

Look at the best work coming out of top media organizations. They aren't just publishing text. They are publishing applications. They are building calculators, interactive maps, and personalized dashboards.

This is where my background in AI game design and creative AI tools comes into play. You need to start thinking like a game designer. How do you create a loop that keeps the user engaged? How do you reward them for exploring the data?

When we built tools at RemoteTeam, we gamified the onboarding process. We used data to show users how much time they were saving compared to their peers. It created a sense of progression. You can do the same thing with journalism. Let the reader play with the data. Let them discover the insights on their own.

If you are just writing static articles, you are competing with millions of other writers and an infinite number of AI bots. If you build interactive data experiences, you are competing with almost no one. The barrier to entry is higher, but the rewards are exponentially greater.

The Reality of the Work

I didn't get to where I am by following the traditional path. I got here by combining disciplines that most people keep separate. I mixed the aggressive growth tactics of Silicon Valley with the rigorous inquiry of journalism. I used code to scale my curiosity.

Computational journalism is not about fancy tools or big budgets. It is about a mindset. It is about refusing to accept the surface-level narrative. It is about digging into the numbers until they confess their secrets.

You have the tools. You have the access. The only thing stopping you is your willingness to do the hard, unglamorous work of cleaning the data and finding the story. Stop waiting for permission. Start building.

Frequently Asked Questions

Are these recommendations still relevant in 2026?

Absolutely. While specific tools and tactics change, the underlying principles remain consistent. I update my thinking regularly based on what I'm seeing in the market and across my portfolio companies.

How do I know which items apply to my situation?

Start by honestly assessing where your biggest bottleneck is right now. The items that address that specific constraint will give you the highest return on your time and energy.

Can I implement all of these at once?

I'd strongly recommend against it. Pick the 2-3 items that resonate most with your current situation and focus there. Trying to do everything simultaneously is a recipe for doing nothing well.

Which item on this list has the highest impact?

It depends on your stage and context, but in my experience, the items near the top of the list tend to have the broadest applicability. That said, sometimes the less obvious items create the biggest breakthroughs for specific situations.

More in AI for Creators

All AI for Creators articles · Sahin's angel investments · Startups he founded