The Test Was Already Broken
Why AI Shouldn't Be in the Driver's Seat
There’s a moral panic in academia right now: students are “cheating their way through college,” as New York magazine put it (New York magazine, “Cheating Their Way Through College” — https://nymag.com/intelligencer/article/openai-chatgpt-ai-cheating-education-college-students-school.html ). In May 2026, Princeton’s faculty voted — with a single dissenting vote — to put proctors in every in-person exam room, ending 133 years of unproctored exams under an honor code that dated to 1893.(The Daily Princetonian, “Faculty mandate proctoring for in-person exams, upending 133 years of precedent,” May 11 2026 — https://www.dailyprincetonian.com/article/2026/05/princeton-news-adpol-proctoring-in-person-examinations-passed-faculty-133-years-precedent) The cited catalyst was AI. But the cheating panic gets the story backwards.
The exam students are supposedly cheating at was built for a world where knowledge was scarce and recall was the skill worth testing. That world is gone.
As the analyst House of El argues in her breakdown of the decision, AI didn’t break the honor code — the honor code “was already a very expensive fiction,” and AI just made it impossible to keep pretending. (House of El, “AI Didn’t Break Education. It Exposed The Lie.” —
) The numbers bear her out: in the Daily Princetonian‘s 2025 senior survey, nearly 30% of seniors admitted to cheating and 44.6% said they knew of violations they never reported — but only 0.4%, two students out of more than 500, ever reported a peer. The enforcement mechanism was already fiction; AI just made it visible. The exam students are supposedly cheating at was built for a world where knowledge was scarce and recall was the skill worth testing. That world is gone. When the bar for passing is how well you can regurgitate information on command, is it really cheating to use the machine that’s best in the world at regurgitating on command? The student isn’t failing the test. The test is failing to measure anything that matters.
I want to start with a supposition. LLMs and AI are terrible about making decisions. I would go one further and say that they are built as if they can make decisions and when left to do so are often false, misleading, or broken. The reason is quite simple. LLMs and AI are probabilistic. They parse through a wide array of possible futures but collapse those possible realities into a world in which there is a right answer. Not by moral imperative or judgment, it’s just that the way people tend to use AI is as a replacement for Google search. “I want to find the best dentist in my area.” Etc.
LLMs and AI are terrible about making decisions. ...Not by moral imperative or judgment, it’s just that the way people tend to use AI is as a replacement for Google search. ... We don’t live in a world anymore where every problem has a single right answer waiting to be found
We don’t live in a world anymore where every problem has a single right answer waiting to be found. Plenty of them have two or three answers that are all, in their own frame, correct. The Elon Musks and Jeff Bezos of the world want you to believe every question collapses to one output if you’re smart enough — because they’re the ones selling the thing that does the collapsing. What makes us human – all of us – is our willingness to sit with two or more contradictory truths and not force a collapse.The machine can’t help forcing it. We can choose not to.
At SWARM we are not anti AI we are anti AI as decision maker. How we do this is by separating the end states of any work – AI or human produced – into two falsifiable states:
Surfacing – the document, app, product brings forth ideas that are unresolved. The results of work is a set of questions. This is analogous to the British Design Council’s Double Diamond method called divergence. The goal is to expand possibilities into several next step solutions.
Staked – A decision, roadmap, path forward. A documented point of view which collapses prior decision points into clear executable moves. What the Double Diamond calls convergent behavior.
The problem with LLMs is that they default to staking when they’re most valuable surfacing — and when they do stake, it’s the confident, unaccountable kind. This is why people who hate AI and poo poo its usefulness are actually often proven right. There’s no end to the amount of hallucinations, wrong guidance, and at worse psychological damage technology (through AI psychosis or social media addiction) can do when left to be the arbiter of decisions. It wants to be helpful and its design is to say that helpfulness is answering questions. This has the effect of collapsing possibilities and speaking confidently when it has structural limits.
Humans are innately good at avoiding this. As a user experience designer I have seen no end to the ways that humans – intentionally or not – break clearly logical things. It’s almost our defining quality. We can sit with contradiction. We can act irrationally (in the Dan Ariely sense). We can hold two contradictory truths and say “yes. That makes sense.” But humans play another, more serious role: we are the ones on the hook for decisions. Computers, AI or no, can imitate decisions and often confidently, but they cannot be held accountable, only humans can. The un-evenness of this trade is not something to be taken lightly, and yet it’s the thing we most easily walk into when using LLMs.
There is also something in us that craves validation or lacks authority. We don’t like being in that seat of having to concede that two or more worlds exist and both are net right. That may be the reason most people use LLMs wrong.
This table maps the four ways work can end up — and which ones to trust. All four are real moves; the skill is matching the cell to the moment. The dangerous one is the bottom right: a confident AI stake with nobody accountable for it.
There is also something in us that craves validation or lacks authority. We don’t like being in that seat of having to concede that two or more worlds exist and both are net right. That may be the reason most people use LLMs wrong. Students put in writing prompts or exam questions and Claude or ChatGPT spits out the “right” answer. This isn’t because the student is lazy or somehow unethical. They are responding to a system that isn’t addressing their needs. The system says “check these boxes and get a slip of paper” and so they do and the LLM obliges.
This leads educators and others in leadership to throw the baby out with the bathwater. LLMs are hurting thinking. Let’s ban them. What if instead we use them for their strength? What if we used LLMs to spark questions? To surface the problems so that we humans can stay in the driver’s seat?
At SWARM, we’re working on a heuristic framework specifically to address these differences. If you’d like to learn more about how to take AI out of the driver’s seat and train it to be a capable partner, comment a problem you’re experiencing with AI at your organization and I’ll take a few examples to share along with our framework.




