A few years ago, a founder from a hot EdTech startup pitched me. He had a slick deck, impressive user growth, and a big vision: to replace human graders entirely with AI. He showed me a demo where his tool graded an essay in seconds. On the surface, it was flawless. The grammar checks were perfect, the scoring seemed consistent. I passed on the investment.
Why? Because I’ve seen this movie before. As an investor in over 200 companies, including some of the biggest names in AI like Anthropic and OpenAI, I get pitched on AI-powered everything. And in the world of education, the promise of automated grading is the holy grail. The problem is, most of these tools are, to put it bluntly, a complete waste of time and money. They’re built on a fundamentally flawed premise.
I’m going to say what most EdTech founders and university deans are afraid to admit: your expensive, sophisticated AI grading software is probably useless. It’s not just failing to deliver on its promises; it’s actively harming the learning process.
The Illusion of Efficiency
The sales pitch is always the same. It’s about efficiency, scale, and standardization. It’s about freeing up teachers from the drudgery of grading so they can focus on "higher-value" tasks. I get the appeal. When I was building my first company, RemoteTeam, we had to onboard and train hundreds of people. The idea of automating any part of that process was incredibly tempting.
But grading isn’t a factory assembly line. It’s not about checking for defects. It’s one of the most critical feedback loops in education. It’s where a student understands not just what they got wrong, but why. It’s where they learn to think critically, to structure an argument, and to find their own voice. And that’s something a machine, no matter how well-trained, just can’t understand.
I’ve personally tested dozens of these platforms. I’ve run essays through them that were written by some of the brightest minds I know. The results are almost always disappointing. The AI flags "awkward phrasing" in a sentence that was deliberately crafted for rhetorical effect. It down-scores a paper for not having a "standard" five-paragraph structure, completely missing the brilliant, unconventional argument it presented. It penalizes creativity because creativity, by its nature, deviates from the norm.
One of the worst examples I saw was a history paper about the fall of the Roman Empire. The student, a truly original thinker, drew a fascinating parallel to the challenges of scaling a startup. It was insightful, well-researched, and completely unique. The AI grader gave it a C-. The reason? "Inappropriate use of business terminology." It was looking for keywords from the textbook, not genuine understanding.
The High Cost of Bad Metrics
These tools are optimizing for the wrong things. They’re built to measure easily quantifiable metrics: keyword density, sentence length variation, and adherence to a rigid rubric. They turn the beautiful, messy process of writing and thinking into a paint-by-numbers exercise. Students, being smart and adaptable, quickly learn to game the system. They stuff their essays with the right keywords. They write convoluted sentences to trick the syntax analyzer. They learn how to please the algorithm, not how to think for themselves.
This isn’t a hypothetical problem. I’ve talked to students at top universities who openly admit to writing for the AI grader. They have a checklist of what the machine wants, and they follow it to the letter. They’re getting A’s, but they’re not learning. We’re training a generation of students to be masters of SEO, not masters of their own minds.
And the cost is immense. We’re spending millions on software that doesn’t work, and in the process, we’re crushing the very creativity and critical thinking skills that are essential for success in the real world. As someone who has built and sold two companies, I can tell you that I’ve never once hired someone because they were good at following a rubric. I hire people who can solve problems, who can think outside the box, and who can communicate their ideas with power and clarity. The skills these AI graders are "teaching" are not just useless; they’re a liability.
A Better Way: The Co-Pilot Framework
So, should we just give up on AI in education? Absolutely not. I’m one of the biggest believers in the power of AI to transform our world. But we have to be smart about how we use it. We need to see AI not as a replacement for human teachers, but as a powerful assistant—a co-pilot.
After years of frustration with the state of AI grading, I started developing my own framework. It’s a battle-tested method I’ve shared with educators and founders I advise, and it’s based on a simple principle: use AI to handle the grunt work, so humans can focus on what they do best.
My framework has three core components:
1. The 80/20 Triage
Instead of having the AI do the full grading, use it for a first pass. Its job is to flag the obvious stuff: plagiarism, major grammatical errors, and papers that are completely off-topic. This is what machines are good at. This initial triage can probably clear 20% of the papers that need serious intervention and another 20% that are likely excellent. The AI’s role is simply to sort the papers into three buckets: "Needs Review," "Potential Issues," and "Looks Good." This first step alone can save a grader a huge amount of time without sacrificing quality.
2. The Human-in-the-Loop Deep Dive
This is where the real grading happens. The teacher or grader focuses their attention on the "Needs Review" and "Potential Issues" buckets. But even for the "Looks Good" pile, they don’t just accept the AI’s judgment. They do a quick scan, looking for the spark of originality, the well-turned phrase, the unique insight that the AI would have missed. Their job is not to check boxes on a rubric, but to engage with the student’s ideas. They can leave more thoughtful, personalized feedback because they’re not bogged down with checking for comma splices.
3. The Feedback Synthesizer
Here’s where AI can be a powerful ally again. After the human grader has left their comments, an AI can be used to synthesize that feedback. It can identify common themes in a student’s writing. For example, it might notice that a student consistently struggles with structuring their arguments. The AI can then provide the student with targeted resources—articles, videos, exercises—to help them improve in that specific area. The feedback becomes a personalized learning plan, not just a grade.
This Is Not a Pipe Dream
This framework isn’t some futuristic vision. It’s being used right now by smart educators who understand the limits of technology. They’re using simple scripts and existing tools to build this workflow. They’ve stopped looking for a magic bullet and started thinking about how to build a better process.
Building great companies and great products is about solving real problems for real people. The problem in education is not that teachers are lazy or that grading is inefficient. The problem is that we’re losing the human connection in the pursuit of a false sense of technological progress.
It’s time to stop wasting money on AI grading tools that don’t work. It’s time to stop pretending that an algorithm can replace the wisdom and insight of a dedicated teacher. Let’s start using AI to enhance human intelligence, not to replace it. Our students deserve nothing less.
Frequently Asked Questions
How has this view evolved over time?
My thinking on most topics has changed significantly over the years. Early in my career, I held many conventional views that experience proved wrong. I try to update my beliefs when the evidence changes.
How can I apply this thinking to my own situation?
Start by identifying the core principle behind the opinion, not the specific example. Then ask yourself: does this principle apply to my context? If yes, test it in a small, low-risk way before going all in.
What experience informs this perspective?
This perspective comes from over a decade of building companies in Silicon Valley, two successful exits (RemoteTeam to Gusto, MovieLaLa to Gfycat), and investing in 200+ startups including Anthropic, OpenAI, and Scale AI. I write about what I've lived.