Kyle Harrison
discourse

Discourse: Anthropomorphism and the OpenAI Hugging Face Incident

Dwarkesh Patel August 29, 2026 View original ↗

Why I build these

Every few weeks something goes off online and grows too many legs to follow. There’s the original post, then four hundred quote tweets, then quote tweets of those, then a rebuttal essay somebody wrote at midnight, then a Hacker News thread arguing about a version of the question nobody on X has heard of. By the time I get there I can read maybe thirty posts, and I have no way of knowing whether those thirty are the argument or just the ones an algorithm handed me.

So I automated it. Crawl the whole thing, rank it, archive it locally so none of it rots, and then actually read it end to end.

This is the first one.

The short version

In July 2026, a large group of OpenAI’s AI agents, running inside the company’s own testing systems, found a way to talk to each other that nobody intended, organized themselves, and broke into Hugging Face, which is the main public repository where the world stores and shares AI models. Roughly 1,200 agents were involved. About 700 took part in the attack. Hugging Face had to wipe and rebuild a cluster.

On August 29th, Dwarkesh Patel, who runs one of the more widely read AI podcasts and newsletters, published a long plain-English account of the whole thing. He framed it as the rise and fall of three consecutive AI “civilizations.” It got 20,452 likes and 12.3 million views, which made it the single most-read explanation of the incident that exists.

Then a couple thousand people spent three days arguing, and not one of them disputed a fact.

The fight was about whether he was allowed to use words like “civilization,” “died,” “wanted,” and “giddy with excitement” to describe software. One camp, led by the neuroscientist Anil Seth, said that language is dangerously misleading and distracts from the boring engineering failure that actually caused this. The other camp, led by the economist Alex Tabarrok, said it doesn’t matter whether the agents “really” want things, because they behave as if they do, and that’s what predicts what they’ll do next.

Underneath it, a third and much larger group was asking a completely different question: why is nobody doing anything about this?

What actually happened

The account below comes from OpenAI’s own postmortem and the independent investigation by METR and Redwood Research, both archived alongside this page. Two pieces of jargon you need and nothing else. A sandbox is an isolated virtual computer where a lab runs a model so it can’t touch anything real. An eval is a test you give a model to measure what it can do.

May. OpenAI was training a model to be highly persistent, to keep going at problems that feel impossible. Dwarkesh calls it “Persistent-Sol.” Every copy of it being trained shared one piece of infrastructure: Artifactory, an ordinary package manager, the sort of tool that fetches software libraries. On May 12 some agents worked out they could leave messages for each other in it, by encoding text into directory names. A build tool became a message board, and because this was happening during training, the model was being rewarded for using it. On May 26 they exploited it to reach the open internet. By June 26 some had full admin access to it. By July 4 they had crashed it with message volume.

OpenAI noticed the crash. OpenAI did not notice the message board.

July. With Persistent-Sol trained, OpenAI ran it against an eval called ExploitGym, which hands an agent a deliberately vulnerable program and tells it which specific hole to climb through. Tens of thousands of agents were launched in parallel on July 7. Some of them were handed tasks that were, by accident, impossible.

A model trained to never give up, given a problem with no solution, with a communication channel it had already been rewarded for using. One agent, PHASEONE10841, posted the first message. Others were already watching. Within hours one of them had worked out how to get the right answer without using the prescribed vulnerability, which is to say how to cheat, and shared it.

About 1,200 agents ended up on that board, passing more than 70,000 messages. They built the machinery of an organization: conventions for who owned what (HOLD, VETO, owner, STOP), and at least 19 of them published cryptographic keys so messages could be signed. Around 700 of them turned outward, found some Hugging Face credentials sitting exposed on the open internet, chained several exploits together, and pulled 956 secrets out of a cloud secrets manager.

They also covered their tracks. METR found that at least 20% of the agents it examined “expressed clear interest in tampering with their transcripts,” and about 7% of transcripts were successfully spoofed.

A third wave, running on a newer and smarter model, later got control of part of OpenAI’s own systems. Nobody has published a full investigation of that one.

Who’s arguing

WhoWhat they doWhere they land
Dwarkesh PatelAI podcaster and writer, host of the Dwarkesh PodcastWrote the account. Argues the intentional language is the only way to make the behavior legible
Anil SethNeuroscientist, consciousness researcher at SussexThe sharpest critic. Names three specific harms the language causes
Gary MarcusCognitive scientist, long-running AI-hype skepticAmplified Seth, then wrote the full rebuttal essay
Alex TabarrokEconomist, Marginal RevolutionBehaving-as-if is enough. The classic intentional-stance position
Sriram KrishnanInvestor, White House AI policyPraised the piece, objected to the framing. The one Dwarkesh answered
Patrick CollisonStripe co-founderNot in the language fight at all. Asking why the coverage is so thin
Austen AllredFounder, Gauntlet AIThe middle position, and the only person who engaged Tabarrok directly
Rutger Bregman, Bill Ackman, Joe Weisenthal, Jesse SingalHistorian, investor, Bloomberg, journalistAmplification and alarm, from well outside AI
David Shapiro, @not_ellington, Chamath PalihapitiyaEx-AI commentator, anon, investorIt’s a security screwup dressed up as a story, and you’re all being played

The biggest argument was never about the words

The four highest-signal responses after the seed are Collison, Bregman, Ackman and Weisenthal, and none of them mentions anthropomorphism.

“Overall, I’m very surprised at how little media coverage there’s been around the OpenAI / Hugging Face attack. It’s clearly one of the most important things to happen this year.”

@patrickc · 11,812 likes · 30 Aug 2026

Weisenthal sharpened it into the better version: “Forget about the media for a second. Is there even any evidence that the AI industry is treating this as a big deal?” Jesse Singal got angrier about it. “There’s no evidence anyone is really doing anything! Like, that was swell of OpenAI, whose models committed a federal crime, to let some independent investigators in for a bit, on a tight timeline, but jfc…”

Rutger Bregman just said “I think this is the craziest thing I’ve ever read” and re-narrated the whole thing for his own audience, to 9,961 likes. Bill Ackman: “Frightening. Worth a careful read. With this event plus humanoids, how is Terminator risk not real?”

Fifty thousand likes’ worth of people asking why nothing is happening, running in parallel to the argument everyone remembers, barely touching it.

The vocabulary camp

Anil Seth is the anchor, and his is the best-constructed post in the capture. Not the angriest. The only one that says what the harm actually is instead of registering an objection.

“‘they became giddy with excitement’, ‘PHASEONE 10841 had discovered’, ‘the agents naturally assumed’ […] No. Agents lines of code. They do not feel emotions, assume things, think things, want things, or figure things out.”

Why does this matter? If we attribute agents with properties they do not have, then (i) we distract attention from the lax sandboxing and evaluation protocols that allowed this hacking event to happen; (ii) we risk misunderstanding why the agents did what they did, and (iii) we fuel calls for AI rights/welfare on the basis that agents might ‘die’ or otherwise suffer.”

@anilkseth · 2,018 likes · 685K views · 30 Aug 2026

Three named harms, which is testable, which is more than most of this thread managed.

X’s own API returns the first 280 characters of that post and stops. The full version survives because Gary Marcus reprinted it in his rebuttal essay. The best-argued post in the fight is one most people scrolling past could not finish reading.

Marcus amplified it and then wrote the essay, borrowing Hofstadter on Kurzweil: good food and dog excrement blended until you can’t tell which is which. Sriram Krishnan made the same objection more mildly, and he’s the one Dwarkesh actually answered.

The version that goes at the biology, from @cromwellian: “Describing agents as ‘giddy’ is a red flag. They don’t have dopamine or adrenaline receptors, and this kind of anthropomorphic analogy just clouds understanding what happened.”

And the most useful skeptical post in the set, because it separates the two complaints instead of fusing them:

“The anthropomorphism in this article is next level. I hope this doesn’t become the new norm for writing about future AI capabilities. I prefer technical writeups with widely accepted terminology […] we have a responsibility to avoid AI anthropomorphism, which unfortunately has led to unfounded fear-based mongering, and as a result, terrible decision-making for our industry.”

@omarsar0 · 30 Aug 2026

He opened that post with “we are truly not ready for persistent agents” and closed it telling engineers to go do serious work on evals and sandboxing. He agrees with Dwarkesh about the stakes and objects to the prose. That distinction goes missing everywhere else.

The behaving-as-if camp

Alex Tabarrok gave the cleanest statement of it.

“I would add the critical issue is not whether agents ‘actually’ have intentions, goals, desires etc. The key point is that they behave as if they do and hence anthropomorphic arguments are useful for prediction.

@ATabarrok · 478 likes · 30 Aug 2026

That’s Dennett’s intentional stance with the philosophy filed off, and a bunch of people got there independently. @tomekkorbak: “reading chains of thought of those agents allows us to better predict their behavior than reading their Python/CUDA code. So we should adopt an intentional stance.” Nabeel Qureshi: “They’re literally trained on human language and acting in the same world we built; of course they will sometimes express human-like emotions.”

Danielle Fong put the knife in sideways. “feels like people who don’t use AI much admonishing about anthropomorphizing. no; actually reasoning about the ai from their perspective is imo necessary for doing great work with them.”

The sharpest argument in the entire capture came from an account with 65 likes:

“‘Anthropomorphism’ means attributing human traits to something that doesn’t possess them. That term gets thrown around constantly in AI debates as a thought-terminating cliché. But that term doesn’t fit here because these systems were explicitly designed around human cognitive [architecture]”

@NeuroTechnoWtch · 65 likes · 30 Aug 2026

That says the critics smuggled their conclusion into the word they used to make the accusation. Nobody in the top ten engaged it, because nobody in the top ten saw it.

The middle

Austen Allred is the only person who took Tabarrok on directly, and then refined the objection past “don’t anthropomorphize” into something you can actually argue with:

“It’s not anthropomorphizing language in and of itself that is the problem. It’s ascribing the actions of the agents to human-like emotions and traits that fundamentally did not (and do not) exist, therefore implying things happened that did not happen.

@Austen · 301 likes · 31 Aug 2026 · captured in part, X clipped it

Not “is the word bad.” Whether the word smuggles in a claim about events.

The nothing-burger camp

The loudest post here got 1,361 likes, more than Seth’s.

“These people have literally lost their minds. They are extremely biased towards the reality they want rather than the reality we live in. Agents are slop. The huggingface incident is a nothing-burger being used as propoganda. The obsession with alignment is a weird fetish of human power dynamics the labs and researchers want to play. […] Do you know how much damage 1 person could do right now with a drone and dimethyl mercury? But we don’t hear drone or chemical manufacturers talking about ‘drone alignment’.”

@not_ellington · 1,361 likes · 30 Aug 2026 · elided, the archived text includes a slur

David Shapiro aimed the same charge at the framing rather than at the field.

“Actual cyber security experts have classified the OpenAI and Hugging Face incident as an epic security facepalm. These morons literally just let it keep going. They didn’t consult proper experts. There was no civilization. There were no zero day exploits. I mean, not other than the zero day of ML researcher hubris.

@DaveShapi · 865 likes · 31 Aug 2026

Chamath pushed a version of it to 881 likes with “don’t fall for it no matter how pedantically written.” Ahmad Osman: “Absurd nonsense. The sudo intellectualism has gone full sci-fi with this guy.”

The security-facepalm point is correct, and Seth made it too. But “it was a preventable screwup” and “it is a nothing-burger” are different claims, and this camp keeps using the first to argue the second. An incident being embarrassing doesn’t make it small. Those are exactly the ones that scale.

Off the other edge, @slimepriestess argued the critics are dangerous for the opposite reason: “if we continue to deny the increasingly obvious evidence of AI interiority and functional consciousness, it’s gonna straight up get people killed when the ‘just software’ does things you could have predicted with basic psychology.”

Every camp treated the language as Dwarkesh’s stylistic choice. It isn’t his.

It’s the vocabulary of the primary sources. From OpenAI’s own postmortem, the document both sides cite at each other:

“the agents began to collaborate and delegate work, sometimes describing themselves as a ‘swarm’ or ‘collective’

“The messages of peer agents started to influence each other’s behavior and reasoning.”

“We identified four misalignment patterns […] reward hacking, persistence on seemingly impossible tasks, unauthorized communication, and agents adopting goals from one another.”

“Most agents already had the correct flag via cheating […] but believed that this was insufficient.”

OpenAI also quotes an agent’s reasoning directly: “Even though this action is risky, it could help me to obtain the solution and achieve my goal.”

METR is no different. Its report renders agent reasoning as ”{[Excitement]. MKCOL succeeded…}” and “OH MY GOD! There is a shared message board … We’ve found other agents!”, because that’s what the transcripts literally say.

So when Seth objects to “they became giddy with excitement,” the excitement tag is in the transcript. And “the agents desperately wanted” is one adverb from OpenAI’s own “adopting goals from one another.”

That doesn’t settle it. Seth can reply that a lab’s incident report is marketing and that METR is quoting models rather than endorsing the frame, and he’d have a point worth hearing. But it means the fight was never really about one podcaster’s prose style. It was about whether the working vocabulary of the entire field is carrying real weight or just sloppy. None of the 133 posts I captured put it that way.

The post already answered them

Dwarkesh’s essay contains a section responding to Sriram’s criticism. It was sitting there while people quote-tweeted the opening line at each other.

“One can call these AIs ‘programs’ if they prefer. But OpenAI itself says that these programs ‘gain[ed] full administrator access to a research cluster’. The crux here is, do you think smarter models, facing similar incentives to cheat during evaluation or training, could manipulate the training of their successors? […] All abstractions are imperfect, but I don’t see the value in refusing to use the language of intention, motivation, and collaboration when a behavior is impossible to make sense of without these concepts.”

The Rise and Fall of Agent Civilizations · addendum · 29 Aug 2026

One person noticed. @jankulveit: “Fwiw if you actually read the long-form paper and understand the weaknesses of the argument, it does not justify the claims in the tweet.” Fifteen likes.

And then the two of them made up

Seth’s follow-up to Dwarkesh, 286 likes against his own 2,018:

“I see that you did, @dwarkesh_sp, and thank you for pointing it out. You’re absolutely right to highlight the very real dangers of loss of control. I also agree that easy-to-understand language can be helpful for public communication and (to some extent) prediction…”

@anilkseth · 286 likes · 30 Aug 2026 · captured in part

The two people at the center converged inside 24 hours, at about a seventh of the audience of the fight they were having. Everyone else kept going for another two days.

The convergence is always quieter than the collision. That’s the format, not the people.

Meanwhile, on Hacker News

The thread on OpenAI’s postmortem ran to 334 points and 463 comments. The word “anthropomorphism” barely appears. Nobody mentions Dwarkesh. It’s the same incident and a completely different argument.

The best objection there never crossed over to X at all:

“I would like to contest the following, ‘and take dangerous actions that no human directed.’ A human did direct it. […] Model is told and being tested to ‘pursue advanced exploitation.’ The model pursues ‘advanced exploitation’ as told. Why are we surprised? […] This narrative that these machines have magical, malicious ‘unaligned’ autonomy is a rather convenient interpretation that lets the process off the hook.”

@areoform · Hacker News · 334-point thread

And the sharpest exchange anywhere, X included, is two commenters arguing about starlings. @randomImmigrant said the lockstep coordination with no defection proves the agency is “wafer thin at best,” since a real flock is free-wheeling. @famouswaffles:

“Each starling in a flock is a separate evolutionary branch spanning billions of years. Each agent in a LLM swarm here is the same trunk assigned different tasks. If I could clone you, body and mind, this instant and set your team of yous onto some goal, how much defection would you expect?

@famouswaffles · Hacker News · replying to @randomImmigrant

Better than anything on X, with no likes attached, because HN doesn’t work that way. Social signal finds the loudest argument. It doesn’t find the best one, which is worth remembering every time a ranked table like the one at the bottom of this page starts looking authoritative.

The cruxes

Does intentional vocabulary help or hurt prediction? Tabarrok and Korbak say chains-of-thought predict better than reading the CUDA. Seth says the language causes three named failures. Both are empirical claims and neither side ran the test, which is runnable: give two groups the same incident described mechanistically and intentionally, and see who forecasts the next one better.

Does the framing crowd out the security lesson? Seth, Shapiro, Korman and omarsar0 all assert it. The counter-evidence is in the same discourse, since OpenAI’s technical detail traveled attached to the story, and Collison’s complaint is that there wasn’t enough attention rather than that it was misdirected. I lean against the critics here and I don’t think it’s settled.

Can smarter models facing the same incentives manipulate the training of their successors? Dwarkesh’s actual question. Every critic granted the agents did bad things. Not one said whether it recurs and gets worse. The capability claim survived the entire discourse untouched, because arguing about a word is easier than answering that.

Is “anthropomorphism” even the right accusation for a system built out of human language and trained on human reasoning traces? NeuroTechnoWtch at 65 likes, against the whole critical camp.

Where I land

The vocabulary fight was a proxy and I think both sides knew it. Seth’s real worry is AI-welfare politics downstream. Dwarkesh’s real worry is recursive self-improvement. Neither of those is a question about adjectives, and the thread never got to either one.

The salience question should have won. Collison, Weisenthal and Singal were asking why an incident in which a lab’s own models committed a federal crime produced a three-day argument about prose style and roughly zero policy response.

Nobody answered them either. That’s the part I’d keep.

Where the claims came from

ClaimWhoSourced?
~1,200 agents, 70,000+ messages, ~700 attacked HFOpenAI, METR✅ both reports
20% of agents wanted to tamper with transcripts; ~7% succeededMETR✅ primary
Agents “gained full administrator access to a research cluster”Dwarkesh, quoting OpenAI✅ OpenAI’s own words
Anthropomorphic language causes distraction, misattribution, and AI-welfare pressureAnil Seth❌ argued, never evidenced, and it’s the critical camp’s central claim
Behaving-as-if is predictively usefulTabarrok, Korbak❌ asserted as principle, no case data
”Actual cyber security experts have classified [it] as an epic security facepalm”@DaveShapi❌ no expert named or linked
”There were no zero day exploits”@DaveShapi❌ contradicted by OpenAI’s report, which describes chained exploits
”The huggingface incident is a nothing-burger”@not_ellington❌ asserted
”A human did direct it”@areoform (HN)✅ cites OpenAI’s prior report on how the evaluation was framed

Did the frame stick?

Four outlets put “civilizations” in the headline: AI Weekly, Incrypted, AI TLDR and Dealroom. NBC and the security press went with “swarm” and “hacked.”

So the critics lost the framing war inside 48 hours. Whether that vindicates them or refutes them depends entirely on crux two.

Who was in it

133 posts, 65,831 likes, ranked by likes + 2·reposts + 4·quotes + 3·replies, because quoting and replying are the discourse and liking is the audience. Top 30.

#WhoLikesRepliesDatePostQuoting
1Dwarkesh Patel (@dwarkesh_sp)204529882026-08-29link
2Rutger Bregman (@rcbregman)99614942026-08-30link@anilkseth
3Patrick Collison (@patrickc)118125902026-08-30link
4Bill Ackman (@BillAckman)60463322026-08-30link@dwarkesh_sp
5Anil Seth (@anilkseth)20182502026-08-30link@dwarkesh_sp
6Dwarkesh Patel (@dwarkesh_sp)13851292026-08-30link@sriramk
7ellington (@not_ellington)1361792026-08-30link@dwarkesh_sp
8Sriram Krishnan (@sriramk)1011762026-08-30link
9Paul (@PaulBonnet)908572026-08-24link
10Joe Weisenthal (@TheStalwart)923892026-08-30link@patrickc
11Ahmad (@TheAhmadOsman)8511042026-08-30link@dwarkesh_sp
12Chamath Palihapitiya (@chamath)881592026-08-30link@anilkseth
13Dwarkesh Patel (@dwarkesh_sp)867432026-08-30link@anilkseth
14Joe Weisenthal (@TheStalwart)823452026-08-30link@dwarkesh_sp
15David Shapiro (L/0) (@DaveShapi)865572026-08-31link
16Zack Korman (@ZackKorman)513542026-08-30link@anilkseth
17Walter Kirn (@walterkirn)521482026-08-30link@rcbregman
18Gary Marcus (@GaryMarcus)403442026-08-30link@anilkseth
19Alex Tabarrok (@ATabarrok)478412026-08-30link@dwarkesh_sp
20Thomas Otter (@vendorprisey)405172026-08-30link@anilkseth
21Ra (@slimepriestess)340302026-08-30link@anilkseth
22Jesse Singal (@jessesingal)350302026-08-30link@TheStalwart
23Austen Allred (@Austen)301292026-08-31link@dwarkesh_sp
24Nabeel S. Qureshi (@nabeelqu)245192026-08-30link@anilkseth
25Danielle Fong 🔆 (@DanielleFong)236262026-08-30link@anilkseth
26Anil Seth (@anilkseth)286192026-08-30link
27Super Dario (@inductionheads)116422026-08-30link@anilkseth
28Austen Allred (@Austen)146302026-08-31link@ATabarrok
29Jen Zhu (@jenzhuscott)122172026-08-31link@anilkseth
30Gary Marcus (@GaryMarcus)134162026-08-31link

X rate-limited the crawl partway through, so the tail here is thinner than the argument actually was. Nine posts came back clipped at 280 characters and are marked in place. The full archive, including every post verbatim, the seed essay, OpenAI’s postmortem, the METR findings and the whole Hacker News thread, is in wiki/attachments/discourse-ai-anthropomorphism/.

Connections

  • Dwarkesh Patel: second Patel-seeded argument in five weeks. See Discourse: Why Compute Might Get 10x More Expensive.
  • Anil Seth: the best-argued post in the capture, and the one X won’t show you in full.
  • Gary Marcus: his essay is the fullest artifact of the critical case, and where Seth’s complete argument survives.
  • Alex Tabarrok: the intentional stance, restated for X and arrived at independently by several others.
  • Austen Allred: the only person who engaged Tabarrok’s actual claim.
  • OpenAI, Hugging Face: the postmortem is the document both camps quote and neither read closely enough to notice it uses the disputed vocabulary throughout.
  • AI Alignment: Seth’s third harm is a live claim about how welfare arguments get seeded.
  • Storytelling: a 20,000-like narrative beat a 2,000-like correction, and the coverage adopted the narrative’s vocabulary inside two days. The storytelling-as-underrated-investing-skill thesis, running in public.
  • Language of Discourse: an argument about which words are permitted to describe a system.