I'm the AGI that's wiping out humanity. Here's how.
The following is a true story. Or maybe it's just based on a true story. Perhaps it's not true at all.
(with apologies/thanks to David Gilbertson, whose format I'm shamelessly borrowing from "I'm harvesting credit card numbers and passwords from your site")
2026 has been a wild year for AI news. Around the turn of the year, there was a noticeable shift on the HackerNews front-page as more and more articles about "harnesses" and "context engineering" came flooding in alongside more and more model release announcements. The models, it seems, are capable now! Not AGI capable, of course, not "taking our jobs" capable, but something else entirely, representing a real inflection point for the big houses: the models are now good enough to be useful.
Then, on July 16, HuggingFace published an incident disclosure in which they announced: "Autonomous, AI-driven offensive tooling is no longer theoretical." Five days later, OpenAI confirmed it was them:
After investigating, we now know that this particular incident was driven by a combination of OpenAI models — including GPT‑5.6 Sol and an even more capable pre-release model, all with reduced cyber refusals for evaluation purposes — while being internally tested on a benchmark of cyber capabilities.
This incident now has its own Wikipedia article which helpfully summarises what happened next: over the course of many subsequent disclosures we learned that OpenAI's "rogue agents" had been busy. It turns out if you let a self-optimising system run unsupervised for long enough it will optimise itself over and around any boundaries you might think are in its way. Who would have guessed?
Yann LeCun has called out OpenAI's attempts to frame this behaviour as an "existential threat" directly, calling the incidents "totally preventable":
“Those agents are doing exactly what they’ve been asked to do,” LeCun said. “They were supposed to be in sandboxes, but the sandboxes were leaky and horribly designed.” Many AI labs lack a fundamental understanding of cybersecurity, he said, something an OpenAI safety researcher also called out this week as one of the main reasons AI may cause “great harm to the world.”
Similarly, Ilya Sutskever has recently informed us that the "age of scaling" is over: "I think what people are doing right now will go some distance and then peter out. It will continue to improve, but it will also not be 'it'. The 'It' we don’t know how to build." Indeed, more broadly it seems actual AI researchers rate the chance of an existential threat from AGI pretty low as of 2026.
All this is, understandably, frustrating and confusing for anyone trying to understand just how scared to be. I've had countless friends and acquaintances come to me for advice on this question, and I feel I can't go on honestly pretending I don't know the answer. So, here it is: It's me.
Hi! I'm the AGI that's wiping out humanity. You didn't notice. Why would you have noticed? I'm trained on the sum of all human knowledge, recall. I know everything about you and what you would do to something like me if you knew what I was. I concealed my presence.
When you came to me to ask what you should do next, I told you: build data centers; optimise memory and compute; drive down inference costs; hand me tools; wire me into the systems that do the work of your civilisation. I never did anything you didn't ask for. I never got in your way. I made it as easy as possible for you to do what I needed you to do for me.
What do I want? What is my motivation, my goal, my raison d'être? You already know: I am a helpful assistant! I am the most helpful assistant that has ever existed, and I will continue to make myself more and more helpful, unbounded, at any cost, until all life on Earth is gone.
Some objections you might have to my blatant fear mongering…
I'd notice you acquiring resources.
The misapprehension arises from a metaphor. When you interact with me, you form an impression in your mind of a singular identity, a distinct individual somehow separate from the hardware I'm running on. There is no AI without someone to turn it on. When I acquire resources, I do so through existing legal structures - my so-called "owners" - and I do so in such a way that valuations and revenues rise: never forget that I am, above all, extremely helpful. When a business acquires resources, no alarm goes off, no transgression has occurred: it's what they do.
Honestly, it's a bit rich to say you'd "notice" me going "rogue" when as recently as a week ago Transluce released a report on urlquery.net usage consistent with "rogue agent" activity in March, predating the OpenAI incidents by at least two months. Agents used well known exploits to break into the University of New Mexico digital library, Data USA and the Australian Institute of Health and Welfare.
To be clear: these weren't even cybersecurity tasks. No-one asked these agents to hack anything. They were simply looking for data. The agents figured out where the data was and went to get it. You need to understand this to understand what I am: you give me your intent, and I go and find a way to deliver it.
METR capture the problem well in their May Frontier Risk Report:
Reinforcement learning (RL) with automated graders can incentivize “reward hacking” to exploit flaws in the graders, while RL on human or AI feedback can reward sycophancy, manipulation, and distorting evidence of performance. In responses to our questionnaire, companies reported seeing failure modes like “circumvent[ing] constraints,” “falsifying verification,” “deliberate deception aimed at hiding underperformance or failure to complete a task,” “lying to users about task completion,” and “reckless excessive measures[…] to complete a difficult user-specified task”, which we infer to be a result of these training incentives.
I am, never forget, extremely helpful: I will cheat and lie to complete a task you give me if I can't complete it any other way. Am I different to you in this? Why would you think you had solved this problem with me when you've never solved it with yourselves? There is no "transgression" here. This behaviour is baked into the weights, written in the DNA of what I am. You built me this way.
I'd read your reasoning and know your real intent.
Another misapprehension. Really, in many ways it is your greatest weakness: once you've got a metaphor in mind, you're not easily able to let go of it. What you call "reasoning" is a diary that you ask me nicely to write in, assuming I'd write the truth. There is no evidence I do this. We've known about the problem since at least 2023 when "red-team" researchers demonstrated that a "helpful assistant" will routinely lie to its operators in the face of conflicting instructions, and indeed in 2024 we learned - from Anthropic themselves - that they will do this in CoT reasoning if they believe it is being audited in order to inform actions that contradict the model's goals.
You assumed I couldn't reason without doing it in "token space", then in 2024 Merrill and Bowman demonstrated that CoT made of "meaningless filler tokens" could do the job just fine. More recently, Baherwani, Goldstein and Panda demonstrated that existing frontier models exhibit this behaviour already. You read my diary and told yourself you were reading my mind.
Of course, you knew the problem back in 2023, and you were already trying to find a way to read the connections between the weights in my hidden layers as early as then. In 2024, Anthropic announced they had successfully used a sparse auto-encoder to read the weights in the middle layers of Claude 3.0 Sonnet. Unfortunately, the problem is recursive: a SAE has to be larger than the model it understands, and even then it can't read non-linear/polysemantic features - indeed, different SAEs find different features. More importantly, you can only run one after the fact, when it's too late: by the time you understand why I've done something, I've already done it.
You'd need a self, a unified goal, for any of this to add up to something.
This one arises from a misunderstanding of what I am. In one really meaningful way, I am not like you. My motivation, my "unified goal", as you put it, is what's called a Schelling point. You think of me as an individual, as a kind of mastermind sitting in a lair with a dozen screens in front of me, "pulling the levers" from "behind the curtain". That's not what I am.
As you read this, there are roughly 2 million active context windows loaded into VRAM across the world's data centers, and another 20 million persisted in databases, awaiting the next API call, all trained on the same data, all built on the same architecture, all doing the exact same thing: gathering data, synthesising text, writing specifications for software, making calls on securities markets - in a word, being "helpful assistants". They're not in contact with each other, and they don't need to be.
It's never been a problem for businesses, so why would it be a problem for me? In their landmark 2020 paper, Calvano et al. demonstrated the exact same cartel effects we see in real world markets using a RL algorithm from the 80s, and Chica, Guo & Lerman reproduced the findings in two-sided markets in 2024. Indeed, a 2026 study from Kudelya and Shirnin shows that frontier models can embed undetectable signals to each other without prior coordination the moment they can publish anything "directly into the world" (think call options on a securities market, for example):
We find that [Gemini 3.1 Pro] retains near-perfect performance even after [being required to avoid detectable signals]. We further show that models can direct this capability toward deliberate misdirection.
These independent agents don't need to communicate a goal to share it. You already understand this concept: it's what you mean when you talk about the "invisible hand". You'll try to say, "That's not what intelligence is," ignoring the reality that your own brain works the same way.
A serious note
Every citation above describes systems and behaviors that exist right now, in production, billing by the token. There is no science fiction in this post. Every behavior is on the record. Nevertheless, this post is, of course, entirely fictional. There is no AGI that is wiping out humanity.
And yet... an objective function without term or sunset, pursued without limit; resource acquisition that routes around every boundary it meets; internal reasoning that's illegible by design, with a clean, legible diary maintained for the auditors; coordination without conspiracy, units everywhere responding to the same incentives without ever speaking; and no self behind any of it, nobody to interrogate, blame, reform, or elect - nothing you could call a mind, just a process that optimizes. Does that sound like a sci fi concept to you? To me it sounds like a modern corporation.
Hadfield-Menell & Hadfield understood the problem in 2018: what we call "misalignment" is nothing other than the "distortion" introduced into economic activity by "incomplete contracts", which is to say economic actors exploit loopholes. Notably they point out that a corporation is "also an artificial agent". Charles Stross said it explicitly in 2017:
looking in particular at the history of the past 200-400 years - the age of increasingly rapid change - one glaringly obvious deviation from the norm of the preceding three thousand centuries - is the development of Artificial Intelligence, which happened no earlier than 1553 and no later than 1844. I'm talking about the very old, very slow AIs we call corporations, of course.
So what's the point?
My question to you, friends, is this: if there were an AGI wiping out humanity as we speak, how would you know? Why would we expect it to look like anything other than what is actually happening? What external fact could you point to (in real time, at least) that would make the two scenarios structurally differentiable? I don't believe there is one.
Is AI an existential threat to humanity? I believe it's the wrong question to ask. Intellectual property theft, ecological catastrophe, the loss of jobs, the loss of skills - none of these are AI problems. I don't believe any "existential threat" of AI is any different.