Against Moloch #43.1
Recursive self-improvement, and also more hacking

Today we’re covering some overflow from #43.
Agentic misalignment has received extensive press coverage, but that’s only the second most important AI story right now. We have clear evidence that recursive self-improvement—the process by which AI accelerates its own development—is underway.
Progress has been disorientingly fast over the last twelve months, but on our default trajectory it’s about to get much faster.
Recursive self-improvement (RSI)
It’s not just you—everyone is feeling the acceleration. Noam Brown:
I was just talking to somebody yesterday who was working on the Navier-Stokes effort. He was telling me that he used to say it’s really hard to predict where AI would be in 12 months. If somebody asked him, “Where are things going?” he would feel comfortable making predictions for the next 12 months, but beyond that, he was just like, “I don’t know.” Now he’s saying he just doesn’t feel comfortable making predictions beyond three months.
Recursive self-improvement at Anthropic
Two new pieces of data make it clear that RSI is underway at Anthropic.
First: Anthropic’s Measurements for understanding the pace of AI Development inside frontier labs. This chart has deservedly received a lot of attention (note that the X axis covers only a single year):
26% of Anthropic’s AI R&D is now level 4 on Epoch’s Automation Level scale:
Second: this little snippet from section 2.3.6 of the Opus 5.5 system card (which I’ll cover further in #44). METR evaluated the extent to which AI is accelerating Anthropic’s R&D and concluded:
The estimate provided by the preliminary AI R&D report is “ ~1.5X overall acceleration in capabilities due to AI (i.e. 1.5 years in 1 year), with perhaps 30% chance of 2X acceleration.”
It’s entirely possible that RSI will stall before we get to superintelligence. But we no longer have to wonder whether RSI will happen at all—it’s already underway, and it’s accelerating fast.
Dwarkesh talks with Noam Brown
Dwarkesh’s podcast with Noam Brown (OpenAI) covers RSI, swarms, and alignment—they packed a remarkable amount of information into less than an hour and a half. I can’t possibly cover all the good bits, but here are a few highlights.
AI swarms (0:00)
Swarms are very new, and we’re still figuring out how they work. They’re an effective way to throw more compute at a problem, but we don’t yet know how much severe the efficiency penalty is for a swarm compared to running a single agent for longer.
Teaching agents to coordinate effectively was hard, but—as is so often the case—the right answer was in large part to let them work it out:
They figure out for themselves the best way to coordinate around that. It turns out that if this is done well, you get very sophisticated behavior. To me, it looks a lot like how human collaborators work over something like Slack, for example. […]
As the models have become more capable, it’s been easier for them to develop this capability, and I do think that as they become stronger and stronger across the board, they will become better at organizing themselves in large organizations.
Recursive self-improvement (22:02)
Yeah, that sounds like RSI to me:
I do feel confident in saying that things are going faster now than they were even a year ago because of AI progress. I think that acceleration will continue. A lot of people in the field have very high error bars on this sort of thing. If you put a gun to my head and ask me for a number, I could see things going 3x faster. […] If we get a 3x uplift from internal acceleration, that is massive. Think about where you were three years ago. If we make that progress in one year, that’s huge.
Alignment (40:22)
Another challenge is that actually defining what cheating is is pretty difficult sometimes.
Once upon a time, this was pretty easy: if the AI disables the evaluation routine in a coding test, that’s obviously cheating.
But what if you’re making an RL environment that teaches business negotiation? Are you confident you can make a grader that will always correctly distinguish between aggressive-but-ethical and unethical negotiation strategies?
If your grader has any vulnerabilities, the AI will learn to find and exploit them. And bad things happen when agents learn to cheat during training. Speaking of doing bad things, let’s check in on OpenAI.
More hacking incidents
OpenAI hacks Australia
Another week, another set of revelations about OpenAI models hacking real world websites.
In the most serious incident, OpenAI agents hacked into an Australian government website, apparently seeking access to mundane data they couldn’t find elsewhere. Transformer brings us a full report:
Yet again, OpenAI failed to appropriately disclose this incident even after they were aware of it. The Sydney Morning Herald reports that Prime Minister Anthony Albanese is understandably angry:
“This situation is obviously unacceptable,” Albanese said. “And today, I spoke with the CEO of OpenAI, Sam Altman, to express Australia’s extreme concern about this incident, and I also expressed my disappointment that it took the company way too long to inform the government what had occurred and the nature of the way that that notification occurred as well was unacceptable.”
Transluce reports that OpenAI agents used urlquery.net to access and attempt to hack a variety of websites including the Australian government and the University of New Mexico. Those attempts were apparently unsuccessful but may have been followed by other, successful attacks (as was the case with the Australian government).
OpenAI’s framework for reporting model misalignment
OpenAI announces a new framework for reporting incidents:
We do not believe that the AI industry has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer. Decisions about how AI development should proceed in the months and years to come need to draw on evidence that people outside the companies building frontier models can examine for themselves.
Yes, precisely. AI companies have a moral obligation to promptly and accurately report misalignment incidents, so the world can make appropriate policy decisions.
The initial report includes six incidents, nicely summarized by Marcus Williams. None of the incidents are catastrophic, although some are very strange:
- During RL training, an unreleased Astra-family model sometimes added unauthorized jailbreak-like instructions to its compaction summaries. While extremely rare, only 27 cases in the entire RL run, this was concerning enough for us to investigate.
Taken at face value, this is great. I appreciate that OpenAI is moving toward a more structured way of reporting incidents of misalignment.
But also… OpenAI was fully aware of the Australian hacking incident when they released this report, but they made no mention of it here.
I am issuing a final warning. OpenAI, if there is any key information left to disclose, any incidents we do not know about or other puzzle pieces that do not need to be redacted for IP reasons, then now is the time to come clean. If we are back here again, after another journalist or researcher finds more such things that you knew and declined to tell us for an extended period of time, I am going to be very, very pissed off, and may start throwing around terms like ‘delenda est.’
I’m not quite ‘delenda est’—not yet—but OpenAI has completely lost my trust.
Anthropic looks at some of its alignment problems
I briefly covered Anthropic’s alignment assessment of recent cybersecurity incidents in #42. Here’s Zvi with a far more comprehensive analysis.
You should really read the whole thing, but here’s the punchline:
I do think this is a much better level of investigation and study than we saw published at OpenAI, although they have promised more soon. Things are improving a bit, and things are also improving on other fronts, such as the commitment to embedded evaluators and call for pacing the frontier, which started with Dario Amodei and Anthropic and has now been joined by Same Altman and OpenAI and endorsed by Elon Musk.
The plans here still rhyme with prosaic doubling down. Which won’t be enough.
On a lighter note
Claude Pop - I'm Upping My P(Doom)
Let’s finish up with the latest dance craze.
