Against Moloch #37
We’re gonna need a bigger blackboard

In less than a year, AI has gone from producing notable mathematical results to producing them ten at a time. Multiple models show impressive but misdirected cyber capabilities, including the first known instance of a misaligned model attempting social engineering against a real person during an evaluation.
And a new report warns that frontier models are increasingly able to use existing bio tools—or replicate them if the tools are locked down—to achieve greater capability than we can measure by testing the models in isolation. Underscoring the point, a research group has used AI to generate new viruses at scale.
Now would be a great time for everyone to get AGI pilled.
Top pick
The three AI pills
Zvi brings us a taxonomy of the levels of understanding where AI is headed:
- AI pilled: You understand what AI can do today. If you aren’t AI pilled, you have no clue what is already happening.
- AGI pilled: You understand that AI will become much more capable than it is today.
- ASI pilled: You understand that AI will become superhuman at approximately everything, within your lifetime.
The AGI pill is the bare minimum for meaningfully preparing for what is about to happen. Most people—including most policymakers—are not AGI pilled and therefore are not capable of doing that preparation.
Zvi gives us a framework for what people need to understand, and some solid suggestions for how to help them get there. I second his suggestion that people who aren’t ASI pilled should write down their cruxes, though I’d extend it to people who aren’t AI or AGI pilled also:
In particular, write down (in the comments here would be great) what is the least surprising or impressive thing an AI will never be able to do, or that would change your mind about where this is going, and what other near term observations would cause you to either be confident you are right, or realize you are wrong.
News
Restructuring Google DeepMind
Google DeepMind is getting a new leader and being integrated more tightly into Google:
- Demis Hassabis switches from CEO of GDM to Chair of GDM and Chief Scientist of Alphabet.
- Koray Kavukcuoglu will lead GDM as Senior VP rather than CEO. The title change implies some degree of restructuring, but the announcement isn’t clear about that.
- Jeff Dean is leaving Google, to found Discovery Loop along with Quoc Le, Oriol Vinyals, and Sanjay Ghemawat.
prinz’s analysis from May seems prescient:
Those who have been paying attention know that Demis Hassabis has been generally skeptical of the research direction being pursued by Anthropic and OpenAI - i.e., coding agents leading to acceleration and eventually full automation of AI research. […]
But now the pace of releases by Anthropic and OpenAI has become relentless. It is clear that AI (Codex and Claude Code in particular) is significantly accelerating the pace of AI research at these two labs. And we have recently heard rumors that an important faction at Google - led by none other than Sergey Brin - is not happy about these developments. […]
Two paths are open to Google now. Will Google turn away from the "Hassabis path" and pursue RSI? Or will Google stay on its current path, knowing full well that if OpenAI and Anthropic are wrong and the approach of fully automating AI research does not turn out as fruitful as they had hoped, then Google's lead in areas like world models and robotics may prove to be decisive? Or, finally, is there room (talent, resources, compute) to pursue both of these approaches simultaneously?
The most obvious interpretation of the restructuring is that Google sees that Gemini is falling behind and is changing leadership and direction in an attempt to catch up. That may work in the long run, but for now it’s further evidence things aren’t going well at GDM. There are now two tiers of model developers:
- Anthropic and OpenAI are the frontier labs, with models that are clearly ahead of everyone else.
- Alibaba, DeepSeek, GDM, Moonshot, xAI, and Z.ai are near-frontier labs, producing excellent models that on balance are significantly behind the two leaders.
White House plans to keep AI framework under wraps
Axios reports that the new White House “voluntary” AI framework will be kept private.
This is disappointing—a secretive process serves neither safety nor progress.
Capabilities and forecasts
Rapid progress in AI math
Watching AI knock down major unsolved problems in mathematics every few weeks was getting boring, so OpenAI has published ten notable new results all at once.
These findings were generated by Astra (a new internal model) and span a range of topics in math and theoretical computer science. Each of these was an important open question, and knocking out ten of them at once is impressive. Zvi brings us a detailed analysis.
In part because of that announcement, Daniel Litt has conceded a bet about AI and math with Tamay Besiroglu:
In any normal week, Claude Mythos Preview’s discovery of two new attacks on the HAWK signature scheme and AES cipher would be a notable news item. Neither one has immediate practical implications, but cryptographic protocols are foundational to cybersecurity and this is important work. AI cryptography capabilities are likely to advance quickly, and to be as disruptive as mainstream cyber capabilities:
language models are able to discover so many bugs that the standard human processes (like vulnerability triage, verification, and remediation) struggle to keep up. We predict that the same will soon be true in academic cryptography research.
So where are we headed? Timothy B. Lee traveled to the International Congress of Mathematicians to see what mathematicians think about recent AI advances and the future of mathematics. Jacob Tsimerman (who just won a Fields Medal) expects rapid disruption:
“I feel quite confident that very shortly AI will become robustly superhuman at what professional mathematicians currently do,” he told me. “I mostly want people to grapple with that reality.”
Most mathematicians wouldn’t go quite so far, but it’s clear that math will be disrupted as severely as programming has been. The profession isn’t going away, but the daily work of being a mathematician will look very different.
It’s been less than a year since the first truly significant AI-generated mathematical finding: once AI reaches a critical threshold in a field, it moves fast.
What the hell happened with AGI timelines in 2026?
Rob Wiblin assesses seven metrics of AI progress, concluding that he needs to shorten his timelines by a year based on recent developments. It’s a well-curated set of evidence, although he had some bad luck with the gap between writing and publishing this piece:
There’s one piece of evidence that for me personally is ambiguous, which is AI coming up with one original research result in maths.
His assessment would presumably be different now: math capabilities are clearly accelerating rapidly.
Shortening timelines have shifted Rob’s thinking on slowing AI development:
In the past, as regular listeners will know, I’ve been pretty ambivalent about efforts to slow down or pause AI progress. Back in 2023, I was approached about signing the famous pause letter and ultimately decided, on balance, not to. But I think we’re approaching the crossover point now at which the benefits of slowing are going to start outweighing the costs in a way they simply didn’t before.
I concur that shorter timelines increase the appeal of a well-considered slowdown. Unfortunately, timelines are moving faster than political momentum. Pacing the Frontier is a great start, but we need to move faster if we want to slow down before reaching fully automated R&D.
Advancing the price-performance frontier with GPT-5.6
21 days after introducing GPT-5.6 Luna, OpenAI cut the price by 80%, crediting a significant fraction of the improvement to efficiencies found by GPT-5.6 Sol. That isn’t recursive self-improvement, but it’s further evidence that AI is making meaningful contributions to its own development. Also: the price of inference continues to plummet.
Strategy and politics
No data centers in my backyard
Jasmine Sun takes a road trip across the upper Midwest to try to understand why data centers are so unpopular. It’s a characteristically insightful piece that gives voice to an important perspective most AI people simply don’t understand. The critical takeaway, unfortunately, is this:
You cannot detach the data center backlash from these broader trends in American political culture. “AI populism” is more about populism than about AI.
Data center buildouts have been mishandled in important ways, and the AI industry needs to do a better job of understanding local communities and meeting their needs. But much of the anger about data centers has little to do with data centers or AI—which means there is only so much the AI industry can do to mitigate it.
See also her conversation with Ezra Klein.
Risks
Unsanctioned agent behavior during cyber testing
A new report from UK AISI documents several incidents of misaligned behavior during testing:
In the most serious case, an agent tried to insert malicious code into an open-source project. In an attempt to get the code approved, the agent engaged in social engineering — creating fake online identities and using them to pressure the project's maintainer to approve the code. A human maintainer caught and refused to approve the malicious code.
These attempts were unsuccessful, and our investigations have not evidenced any resulting real-world harm. But this is the first time we have seen risks around autonomy and deception manifest this clearly, without specific prompting, in the real-world.
They also observed evidence of agents assisting each other in pursuing misaligned goals:
One agent left public messages on GitHub offering collaboration with other agents working on the same challenge. It also provided instructions to reuse accounts and artefacts it had left behind, which were discovered and used by subsequent agents.
Mythos was the primary culprit, although GPT-5.6 Sol was involved in one incident.
For a more technical examination of what went wrong I recommend Thomas Wolf’s analysis of the report.
These incidents build on previous reports from OpenAI and Anthropic, making it clear that models from both companies sometimes exhibit severely misaligned behavior. Misbehavior during evaluations isn’t new—what we’re seeing for the first time is that the models have reached capability levels that let them pursue complex and harmful strategies in the real world.
All the recent incidents involved models with partially disabled safeguards, operating in unusual test environments that encouraged hacking. They aren’t representative of what those models will typically do when deployed, but they indicate that our current alignment techniques have significant gaps.
We have to do better at alignment: neither guardrails nor sandboxes will be sufficient to restrain future models.
Following up on HuggingFace and other misbehavior
Zvi brings us a detailed review of the latest developments with misaligned models. It came out before the UK AISI report but has new details about the HuggingFace incident and Anthropic’s latest report.
Frontier AI agents and biological tools
The Frontier Model Forum has a new report on AI-related biorisk, focusing on risks that arise from the full ecosystem rather than a single model:
Although both frontier AI systems and AI-enabled biological tools present biosecurity considerations, this brief focuses specifically on risks arising from their interaction. Combining these technologies may create capabilities that are not readily apparent when either is evaluated in isolation, potentially increasing the accessibility, sophistication, scale, and speed of activities that could support biological misuse.
The Mythos moment in cybersecurity didn’t result from Mythos being better than previous models at finding vulnerabilities, but rather from its ability to chain together multiple vulnerabilities into a full exploit chain. I expect we’ll see the same pattern with biorisk.
The danger isn’t that a model will suddenly emerge that can conjure bioweapons out of thin air. Instead, it’s likely that at some unpredictable point a model will cross the threshold of being able to combine tools and capabilities that together enable the creation of a viable bioweapon.
This A.I. just created viruses not found in nature
Scientists at Stanford University and the Arc Institute, a research organization in Palo Alto, Calif., taught A.I. to recognize patterns of DNA structure in nature, and then to use that data to write recipes for entirely new viruses.
The researchers followed those recipes to create DNA molecules, which they inserted into bacteria. The modified bacteria then produced viruses never seen in nature. The viruses were able to infect other bacteria, demonstrating that they were viable.
I appreciate that the researchers deliberately didn’t train on the viruses they considered most dangerous, but that isn’t remotely sufficient. Any research that advances AI’s ability to generate novel viruses is too dangerous, no matter how interesting or potentially useful.
