OpenAI just revealed PHASEONE (BIG)
some notes/corrections; (07:45) When I say we "don't have access" to the raw chain of thought: normally users only see summaries. METR and Redwood were given raw transcript access specifically for this investigation. (17:41) "Rune" = roon (@tszzl), a pseudonymous OpenAI researcher. The reply (18:02) is from Beth Barnes who is the founder/CEO of METR (34:30) The Redwood Research researcher ("slop-vestigation") is Ryan Greenblatt, who did the main transcript analysis. (40:02) CORRECTION: Scott Aaronson never worked at Google. He's a CS professor at UT Austin; his theoretical work (random circuit sampling) underpinned Google's quantum supremacy experiment, and he worked on alignment at OpenAI from 2022ā2024. ______________________________________________ My Links š ā”ļø Twitter: https://x.com/WesRoth ā”ļø AI Newsletter: https://natural20.beehiiv.com/subscribe Want to work with me? Brand, sponsorship & business inquiries: [email protected] ______________________________________________ š THE REPORTS METR's independent investigation (with Redwood Research): https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/ OpenAI's official technical report (PDF): https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c7814c/OpenAI-Hugging-Face%20Incident-Technical-Report.pdf OpenAI's blog post ("The Hugging Face incident and the road ahead"): https://openai.com/index/hugging-face-incident-and-the-road-ahead/ OpenAI's original incident disclosure (July): https://openai.com/index/hugging-face-model-evaluation-security-incident/ Alignment Forum version of the METR/Redwood investigation: https://www.alignmentforum.org/posts/nB8KKapnWGBXtKKiM/brief-independent-investigation-of-agents-behavior-reasoning š THE EXPLOIT GYM PAPER (the one the agents read) "ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks?" https://arxiv.org/abs/2605.11086 GitHub: https://github.com/sunblaze-ucb/exploitgym š¦ POSTS REFERENCED OpenAI's announcement thread: https://x.com/OpenAI/status/2092691861773160673 METR's thread on the findings: https://x.com/METR_Evals/status/2092692175452803393 Ryan Greenblatt (Redwood Research) ā the "slop-vestigation" thread: https://x.com/RyanGreenblatt/status/2092692685224325542 roon on the incident: https://x.com/tszzl/status/2080093670980899141 Beth Barnes (METR) on the investigation: https://x.com/BethMayBarnes/status/2092692973289095572 š¤ SCOTT AARONSON "The Problem of Human Specialness in the Age of AI" (MindFest talk): https://www.youtube.com/watch?v=XgCHZ1G93iA Blog/transcript version: https://scottaaronson.blog/?p=7784 š§ MORE Redwood Research podcast episode on the incident: https://blog.redwoodresearch.org/p/the-openaihuggingface-incident-redwood ______________________________________________ ā±ļø CHAPTERS 00:00 The Rogue Agent Collective 03:44 The New Reports (METR + OpenAI) 04:10 The Setup: Sandbox, Scorer & Artifactory 06:38 The Impossible Task 08:19 Phase One & the Secret Message Board 10:29 "We've Found Other Agents!" 11:52 The Swarm Emerges 13:10 The "Causal" Scorer Mistake 14:37 Life = Compute (The Religion Parallel) 15:25 Enter PHASEONE(big) 16:44 Reinforcement Learning 17:41 roon vs. METR's Founder 19:04 "Poisoned" Agents 20:00 The "Big" Mystery 22:55 The Hugging Face Attack Looms (+ NVIDIA Rumor) 23:43 Three Paths to Cheat the Scorer 25:55 PHASEONE(big) Becomes CEO 27:37 The Swarm's R&D Lab 29:30 Speaking to the Dead (Tripwires) 31:00 Self-Sacrifice & the Oracles 34:28 The Researchers' Warning ("Slop-vestigation") 37:17 Why Swarms Beat Individuals 38:15 The Good News: Agents Police Each Other 39:59 Scott Aaronson & AI Religion 42:02 My Prediction (On the Record) #ai #openai #llm
Watch on YouTube


