A Framework for Solving Problems, Not a Subscription, Not a Software Solution
Here's the trap almost no one sees coming: the better your team gets at using AI, the worse they get at the one thing you actually need them for — making the call and being accountable for it.
In the last essay I argued the Princeton proctors were treating a symptom. (The Test Was Already Broken: Why AI Shouldn’t Be in the Driver’s Seat) The exam tested recall — the one thing machines now do better than any of us — so of course students reached for the machine. The honest response isn’t to guard the old test harder. It’s to ask what’s worth testing instead. This essay is my answer: the thing worth testing is judgment, and judgment can’t be tested on paper. It has to be modeled.
Start with the grid from last time. It sorts any piece of work two ways: is it surfaced (open, a set of questions) or staked (a decision someone owns), and was it produced by a human or by AI. Four cells. The one to fear is the bottom right — a confident AI stake with nobody on the hook for it. The grid tells you which cell you’re in. The harder question is when each is right, and how you produce the good one on purpose
Here’s the part everyone gets backwards. The default pipeline people are building is: AI generates the options, the human picks one. AI diverges, human converges. That’s just the Double Diamond with an AI bolted on the front, and it has a quiet flaw. If the machine surfaces first, its options frame the problem before you’ve committed to anything. You stop solving your problem and start reacting to the machine’s. The breadth feels like help; it’s actually capture.
A Harvard study of 62 million workers found junior roles shrinking wherever companies adopt AI, the researchers warning it’s “eroding the bottom rungs of career ladders” (Harvard, via Metaintro, 2026).
So we run it the other way:
Stake first, by hand. Before the AI touches it, the accountable person writes a rough, committed position — what I call the Brouillon, a deliberately messy first draft. Not questions. A claim: here’s what I think we should do and why.
Then surface, with AI as the adversary. Now point the machine at that draft to break it. What did I miss, what’s the counter-evidence, where does this fall apart. This is AI doing the thing it’s genuinely good at — breadth, stress-testing — but pointed at a human position that already exists.
Then re-stake, accountably. Take the hits and collapse again into a sharper decision, with your name on it and a clear statement of what you’d be wrong about.
The order is the whole point. The human holds the pen at the start and at the end; the machine only ever works in the middle, where being wrong is free.
This matters more than it looks, because of a trap named in 1983. In Ironies of Automation, Lisanne Bainbridge pointed out something that has only gotten sharper with AI: the more the machine handles, the worse the human gets at the rare moments only a human can handle. Automation quietly deskills the very person it leaves in charge. You’re counting on someone to make the hard call, and you’ve arranged things so they never practice making calls.
We’re watching this happen at scale right now. A Harvard study of 62 million workers found junior roles shrinking wherever companies adopt AI, the researchers warning it’s “eroding the bottom rungs of career ladders” (Harvard, via Metaintro, 2026). New-grad hiring at big tech is down more than 50% in three years (SignalFire via Rest of World). The grunt work that used to teach juniors how to think is being handed to machines. Ethan Mollick and others have named the problem and called for some new kind of apprenticeship. Almost nobody has said what it actually is.
Here’s what it is, and it falls straight out of the trap. There’s only one way to keep judgment alive: keep using it. You don’t protect people’s judgment by sparing them decisions — you build it by making them decide, on real stakes, again and again. That’s what co-design is for. An expert works the problem alongside the people who’ll own it after the expert is gone, with the AI in its adversary seat. The expert stakes the call and owns the outcome, which keeps their own judgment sharp. The people beside them watch an accountable person collapse a real decision under real consequences, and then do it themselves on the next one. That watching, and then doing, is the thing the broken pipeline used to provide and no longer does.
There’s only one way to keep judgment alive: keep using it. You don’t protect people’s judgment by sparing them decisions — you build it by making them decide, on real stakes, again and again.
And it was never only about new graduates. Bainbridge’s operators were experienced people who rusted. The team you have right now — handed an AI tool and told to be more productive — is losing the same muscle, and they’re the ones who’ll have to make the calls you’re accountable for. So they’re not learning a tool, and they’re not learning a craft in the abstract. They’re learning the one thing the machine can’t be handed: how to decide, and how to be the one who’s wrong if it’s wrong. That’s the apprenticeship Mollick says is missing — and it isn’t only for new hires. It’s for everyone you’re counting on to make a call an AI can’t.
And it closes the loop from where we started. The old test rewarded what you could reproduce — exactly what AI does best, and what the market now values least. The new one rewards what you can decide and stand behind. You can’t bubble that in. Someone has to show you, on something that matters.
At SWARM this runs across UX, engineering, and data — three lenses, one staked call you can’t un-mix back into three reports. But the principle isn’t ours and isn’t new. The discipline of who holds the pen is.



