“Intelligence is whatever machines haven’t done yet.” Larry Tesler, Tesler’s Theorem, in CV: Adages and Coinages, nomodes.com (ca. 1970)
The support queue used to eat a junior’s whole day: read the ticket, pull the logs, check the runbook, draft the reply, repeat until dark. Then one morning it was not a full-time job anymore. I cannot hand you the before-and-after in hours, because I never wrote it down, so take that as memory, not a metric, and notice that I, a guy who preaches instrumentation, missed the one number I would most want back. I remember it as most of a day before and a glance at a dashboard after. I never pulled the query, so treat it as a war story, not a benchmark. Nobody sent a memo. Nobody threw a parade. I spent that same week arguing on the internet about whether any of it counted as real intelligence, with the same energy I once spent racing for first post on Slashdot, which tells you how much attention I was paying to the thing already at the desk.
Bottom line: We didn’t reach AGI. We blew past the useful definition of it and then moved the goalpost so we wouldn’t have to admit it. No ceremony, no memo, no parade. Just agents in prod doing work that used to be somebody’s job while Twitter argues it doesn’t count because the model can’t fold laundry.
Every decade the definition of “real AI” gets rewritten the minute a machine does the thing:
The objection I hear most, in my mentions and once from a guy at a wedding, is “that’s not AGI because it’s not sentient.” Fine. That is also not what AGI means. The G stands for general, not for feelings. AGI is a capability bar, human-level across most knowledge work, and nobody put an inner life on the spec sheet. Have that argument with a philosopher, on your own time, after the tickets are closed.
Notice the trick. The 2010s definition, most knowledge work at or above a junior, is my own summary of what people meant then, not a citation from a committee. It got met, so we promoted it to superintelligence-plus-embodiment and declared victory for the skeptics.
The freshest example of the rename has a date on it. Securing.AI, September 7, 2026, spent 7,000-plus words proving GPT-6 Astra is not AGI, running it past five definitions so it could fail all five. The interesting number in that piece is a gap, not a score: 62.7% on ARC Prize under a neutral test rig versus 99.9% with OpenAI’s own adapter. That spread is a fact about scaffolding, not about a brain, and scaffolding is an engineering problem you get paid to solve. The one argument that bites is price, roughly $360 a game in compute against a human’s fraction of a cent. Fair hit, and an argument about the bill, not the intelligence, and bills move. Nowhere in it is the question I answer Monday: person or machine on the next batch of tickets.
I moved that goalpost myself, more than once. My personal version was “sure, but it can’t do MY job,” a bar I quietly relocated every time something in prod did a piece of my job, which is the rigor of a guy losing at darts and redrawing the board.
After years of watching people argue definitions the way we argued Netscape versus IE, here is mine, on a bar napkin:
If I asked a human to do this task, would I bet my life the human does a better job?
That’s the whole test. Yes means the machine isn’t there yet on that task. No means stop arguing and staff it. The one edge it does not settle is the task where you are not buying output at all, you are buying accountability: a name on the sign-off, a license to pull, someone a regulator or a jury can hold. The machine can still do the work there. It just cannot answer for it, so the agent drafts and a human signs.
It works because it doesn’t ask whether the thing “really understands.” It asks who I would put my actual life on, the one bar I have never talked myself out of at 2 a.m. Run it on the support queue from the top of this chapter: would I bet my life a junior beats the agent at reading a ticket, pulling the log line, and quoting the runbook, on a Tuesday, after lunch, on ticket number 60 of the day? No. Run it on rewiring my house: yes, and it is not close, and please do not let the model near my breaker box.
It also kills the strawman fight in both directions. Nobody gets to call the bar rigged low, because I’m betting my life on it, and nobody gets to raise it to god-with-hands, because a human plumber can’t fold spacetime and I still pick him over me.
I ran all three in one afternoon, in order, out loud, in a meeting I called. Complete n00b move, and I sent the invite myself.
The quiet part: the people moving the goalpost are usually selling something, a book, a lab, a doomer brand, their own job security. The people staffing the intern are shipping.
The big version got a press release. On 8 September 2026 OpenAI said an internal model, pointed at the open Millennium Prize problems with about 10,000 agents at once, had resolved Navier-Stokes existence and smoothness in roughly 88 hours, plus 17 hours of Lean formalization, at about 130 billion output tokens. The claimed answer is a disproof: a smooth fluid with smooth forcing and finite energy that blows up in finite time. That is OpenAI’s account of its own run. The Clay Institute still listed the problem unsolved on 9 September, independent verification was pending, and OpenAI says it will not claim the million dollars. Tristan Buckmaster and Levent Alpoge posted their own forced-Euler blowup results in Lean on 7 September, and OpenAI acknowledges their priority there. If it holds, a machine closed a problem the field left open since 1934. If it does not, the thesis survives anyway, because the thesis lives in the small version, which got no press release.
That one is on my own machines. The week I stopped arguing was 12 August 2026. Ninety-odd agents across three airank workflows built and shipped an API and a toolbar in a day, and the part that earned the hire was not the code. A side research run, about 1.2 million tokens and half an hour, killed two of the three assumptions the product was built on before we wrote it: the “3.2x more AI citations” pitch collapsed to a measured 0 to 2.4% lift, and the paste-one-tag JS delivery was worthless because the GPT, Claude, and Perplexity crawlers do not execute JavaScript. It also flagged that the rating markup I wanted to ship was a fake-review policy landmine. A junior who told me that on day one would get a raise. Mine cost a half-hour and no benefits.
Here is the spec, because the spec is the transferable part. The support agent gets a tiny fixed toolset and nothing else, scoped to the ticket, the logs, and the runbook. Small on purpose, because every extra tool is another way to be confidently wrong. The stop rule is one line and it is enforced in code, not in the prompt: cite the log line or escalate. No cited line, no reply goes out. Escalation is not a mood, it is a threshold, and the threshold lives in code where you can read it. Escalations go to a senior with the ticket, the log lines it did find, and the runbook section it thought applied, so the human starts at minute three, not minute zero.
That stop rule exists because the first version did not have one, and what changed is the whole lesson. Version one had the same three tools and no gate, and it produced beautiful, confident, fully formatted replies citing runbook steps that did not exist. It was not wrong occasionally. It was wrong in a format that looked more correct than the correct answers, and I had built no way to notice. The fix took an afternoon: make the citation a required field, fail closed when it is empty, route the failure to a person. Two weeks of me arguing about whether the model “understood” anything, one afternoon of wiring the thing that made the question irrelevant. That is the 2010s AGI definition sitting in one queue. Not conscious. Not general. Hired.
The loud failure is believing the hype deck that says the god-mind arrives next quarter. The quiet failure is the opposite:
You wait for a consciousness certificate while the capability eats your market.
Consciousness is a philosophy problem. Cost-per-task is a business problem, and it doesn’t care about the first. While you argue about whether it “really understands,” a competitor with worse taste and better tooling ships straight to prod, ICQ “uh-oh” and all, and halves their ops cost with a loop you could have built. Bad plan shipped beats good plan discussed.
Second, you grant it god-status and stop checking its work. Surpassing junior-level doesn’t mean trustworthy. It means productive and wrong at scale unless Parts V through VII are wired in. Ask me why those are three whole parts and not a paragraph. I did not write them because I am careful. I wrote them because I was not. AGI-or-not is the wrong audit. The audit is: does it cite the log line, and what happens when it can’t.
Do
Don’t
Ch. 1 was the history, this one is the capability verdict, and next come the corpus it learned from, the money around it, and the attachment wave, before Ch. 6 pins down what an agent actually is.
Argument is Jeremy’s (surpassed, not reached), kept as thesis. The “2010s definition” is the author’s summary of the era’s working bar, not a cited standard.
Verified
Reported claims of the moment (September 2026), not settled results
What I could not verify: