Post

This Is the 7th Time We've Reached AGI This Year Alone

Seven AGI announcements in nine months, a Fireship video that almost got me, and the independent benchmark that did not.

This Is the 7th Time We've Reached AGI This Year Alone

TL;DR — we have reached AGI seven times this year. I did not bother checking the latest one — it will be debunked by the weekend. A YouTube video made me look anyway, and an independent benchmark deflated the whole thing. This is the scoreboard, and the cycle underneath it.

I did not bother to even check the announcement. It will probably be debunked by the end of the week, and the press release reads like a horoscope: specific enough to feel true, vague enough to survive contact with reality.

Then a Fireship video landed in my feed about OpenAI’s GPT-6 Astra. OpenAI said AGI, straight up: 99% on ARC AGI 3, a first at its own “critical cyber threshold” — a model that finds and exploits zero-day vulnerabilities on its own. For ten minutes I felt the old pull: maybe this one is different.

Independent testing by Artificial Analysis ran the numbers a week later and scored Astra a 61 — tied with GPT-5.6 Sol, five points behind Anthropic’s Fable 5.1. The model announced as AGI benchmarked like the flagship it replaced.

Not worse. Same. The gap between what the launch reported and what a stranger with a benchmark measured is the entire story of this industry in 2026. The launch is a magic show. The measurement is the cleanup crew.

Somebody on a tech forum put it in one line: “This is the 7th time we’ve reached AGI this year alone.” He was not wrong.

The 2026 arrivals calendar

January. AGI with reasoning. The reasoning was excellent inside a sandbox and mostly complained about the sandbox. Headlines said it could “think.” It could think the way a dishwasher can swim if you throw it in a pool with enough conviction.

March. Agentic AGI. This one could browse the web, which we rebranded as “acting in the world,” because browsing the web is what humans do when they are trying not to work.

May. AGI that apologizes. New capability: self-doubt, delivered in a tone that made every engineer in the demo feel personally parented.

June. Multimodal AGI. It can see me now. It has seen my screen during a meeting and still believes I am a competent professional. Either the vision is weak or the kindness is.

July. AGI that codes. It writes code at a speed that terrified half the internet and with a test-coverage habit that reassured the other half. Both halves were right, which the commit history will eventually confirm.

September. The AGI that could not decide it was launching. OpenAI posted the GPT-6 Astra page, deleted it, and posted it again — “welcome to the AGI era” — while ChatGPT, Claude, Grok, and Cursor all sat down in a simultaneous outage that everyone joked was the model eliminating its competition. The seventh arrival was the messiest and the most confident yet.

Each arrival follows the same liturgy: a demo video, a benchmark chart with the axis bent until the line looks vertical, the word “first” in the headline, the word “safety” in the second paragraph, and a waitlist in the third.

The arrival ceremony

Every time, someone with a podcast announces that everything has changed. Every time, the AGI from the previous arrival is quietly reclassified as “a stepping stone,” which is the industry term for a thing we were very sure about last quarter and now prefer not to discuss.

“We do not reach AGI. We raise the ceiling.”

Nothing arrives

And yet the coffee machine still does not understand “decaf.” The AGI from January cannot tell me if my pull request breaks the build, so I run the tests myself, like a peasant with a terminal.

The one from March cannot find the setting I need in our own admin panel, which is a panel I have filed four tickets about, which is a sentence I typed while an AGI was being announced.

The seventh arrival had the same observable effect as the first six: a pricing page, an outage, and a post from the lab clarifying that the AGI was, technically, a preview of the AGI.

The economics of repeated arrivals

Watch the cycle closely enough and you can set your watch by it. New model, and the benchmarks look like a miracle. A week of coverage, a week of believers, a demo that does the thing.

Then the ration gets cut — the settings settle, the fancy serving config goes away, somebody independent actually runs it — and it measures the same as the model it replaced.

The previous model, meanwhile, gets quietly retired, which is how a company announces progress and a sunset in the same sentence.

Reaching AGI stopped being an event and became a business model. Each arrival resets the benchmark arms race, justifies the next datacenter, and hands the earnings call its verb.

The models improve — they genuinely do — but the arrival is inventory. Seven this year, and every one of them sold out before the demo finished buffering.

The scoreboard

At the current rate we will reach AGI roughly seventy more times before the decade ends. At some point the word has to mean something again, or we admit it never did. I am not sure which outcome I am rooting for.

What I am sure about is the model I actually use. My daily work runs on DeepSeek — excellent results at a modest price, no announcement required, the same bet I made in September. It has never claimed to be conscious. It has also never given me a reason to switch.

The eighth arrival is due any week now. Do not miss it — they are the only product launches left that still sell out.

I will be watching from the cheap seats.

This post is licensed under CC BY 4.0 by the author.