Against Moloch #36
Start at the beginning

Our current strategy for managing superintelligence is for everyone to race to the finish line and see what happens. That has the advantage of bringing the benefits of AGI as soon as possible, and the disadvantage that it might get us all killed.
Pursuing a more sensible approach is harder than it looks: if one frontier lab unilaterally slows down, the others will simply race past it. Even if all the US labs coordinate a pause, China takes the lead. This is a classic coordination problem that can’t be solved by a single actor, no matter how well-intentioned. The solution has to come from government.
This week brings us one step closer to an answer. OpenAI and Anthropic, along with a long list of frontier lab employees, have asked the US government to begin laying the groundwork for an international agreement to pace the rate of AI development. It isn’t enough, but it’s a start.
Top pick
Pacing the Frontier
This week’s top pick is a public letter that is short enough—and good enough—that I’m including the complete text:
AI could help create a dramatically better future, but that outcome is not guaranteed. The world's leading AI companies believe they could be close to automating AI research. It is hard to predict exactly how much this will accelerate AI progress, but there is a real risk that capability development rapidly accelerates beyond our ability to understand or control the resulting systems.
To realize AI's potential, industry, government, and society at large may need the option to buy time to address emerging risks, develop security measures, and strengthen oversight. But each company—and country—is under intense competitive pressure not to unilaterally slow that acceleration. And today, the world lacks the technical and governance tools to deliberately pace frontier-wide progress.
Building on work already underway to monitor frontier model releases:
We request that the U.S. government support an international effort to develop the technical and governance tools needed to deliberately pace the frontier of automated AI development.
The letter has been endorsed by Anthropic and OpenAI, and signed by more than a thousand frontier lab employees, including prominent co-founders from Anthropic, Google DeepMind, OpenAI, Safe Superintelligence, and Thinking Machines.
This letter was skillfully crafted to get a critical mass of prominent people on board. It doesn’t ask for everything we ultimately need, but it’s probably the best we can achieve right now—both in terms of policy and moving the Overton window.
The term “pacing” rather than “pausing” has attracted some controversy, but it was the right call: it’s more palatable to some people who need to be on board, and it leaves open a wider range of policy options.
Zvi brings us comprehensive coverage, while Shakeel Hashim breaks down the two coordination problems that inspired it.
News
Opus 5 in detail
Opus 5 is an impressive model that for many (but not all) applications is closer to Fable than Opus 4.8:
Anthropic’s models are getting much more resistant to prompt injection. That’s less flashy than some other benchmarks but has a huge impact on what kind of tasks I’m willing to trust the models with:
Zvi brings us a detailed breakdown of capabilities, the system card, and model welfare. If you haven’t read one of his model welfare posts, I recommend that you at least skim this one. It’s increasingly clear that model welfare and consciousness deserve serious attention, even though we continue to have no idea whether current AIs are conscious.
Capabilities and forecasts
Tripling GPT-5.6 Sol’s score on ARC-AGI-3
Last week’s newsletter showed Opus 5 crushing GPT-5.6 Sol on ARC-AGI-3. OpenAI investigated, and found that the official test harness turned off two critical features in Sol. Re-running the evaluation with those features enabled tripled Sol’s score from 13.3% to 38.3%, using one sixth as many tokens:
Those numbers make much more sense: Sol is usually a strong game player. Clever harness design can sometimes make a huge difference, but that isn’t what’s going on here—instead, the official harness was passing incorrect parameters to Sol, handicapping its ability to reason over multiple attempts (which is the whole point of ARC-AGI-3). This is a reminder that evaluation is a subtle science and it’s easy for subtle mistakes to result in spurious results.
I stand by my earlier assertion that even though it was designed to test frontier models in one of their weakest areas, ARC-AGI-3 is in danger of rapid saturation.
Strategy and politics
Don’t let AI developers hire their own referees
Many people—including me—have advocated for regulating frontier AI via a system of independent verification organizations (IVOs), much like financial auditors. Writing for AI Frontiers, Gabriel Weil argues that the IVO system has fatally flawed incentives and suggests using mandatory liability insurance instead.
Weil is correct that IVOs would be incentivized to be lenient in order to attract clients—as he points out, the large-scale failure of the credit-rating agencies in the 2008 financial crisis is an example of how the incentives can go wrong.
It’s a thoughtful piece that advances the conversation, but existential risk—where almost all the potential harm lies—is uninsurable. Weil proposes several remedies, most notably arguing that measures that manage mundane risk would also reduce existential risk:
The precautions that reduce insurable losses—tighter security against model theft, stronger containment during evaluation and deployment, and closer monitoring of what agents actually do—are largely the same measures that also reduce the uninsurable downside.
Yes, but only partly. Liability works by making the party responsible for a risk carry the expected cost of that risk, thereby incentivizing them to make an economically efficient effort to mitigate it. If AI insurance is radically underpriced because it doesn’t cover existential risk, it incentivizes labs to spend far too little on risk mitigation.
Further, mitigations against mundane risk only partly help with existential risk. Liability insurance doesn’t create meaningful incentives to guard against long-term scheming, which is the single greatest danger from advanced AI.
Finally, liability creates murky incentives. You’re more liable for causing harm if you were aware of the potential for that harm—in practice, this means that companies frequently choose not to investigate potential harms because doing so would increase their liability. I’ve been in the room when this happened, and it’s gross but real.
Incentives are hard and rarely perfect, but I see IVOs as a better tool for managing existential risk. They are pointed at the actual problem rather than an imperfect proxy, and their incentive problems, while real, are clear and manageable.
The FCC takes action against imported robotic devices
The FCC has taken action to block the importation and sale of foreign-made “advanced robotic devices” for national security reasons. Depending on how it’s implemented, this could have a substantial impact on domestic manufacturing—I’m surprised it hasn’t received more attention.
The obvious question is whether the US can build domestic robotics capacity fast enough—if that’s within reach, a policy like this could potentially boost local production. If not, it’s likely to be self-sabotaging.
Open Weights and American AI Leadership
An impressive array of tech companies has published an open letter in favor of open models. It’s an eloquent, persuasive master class in missing the point:
In fact, openness may be one of the most important paths to AI safety and security. Relying solely on closed models is not inherently safe: they can be breached, misused, or fail in ways that outsiders cannot detect. And concentrating advanced AI capabilities behind a small number of closed models compounds that risk.
This is a great explanation of why open software is helpful for cybersecurity, but it doesn’t automatically apply to other domains. Replace “AI” with “nuclear weapons technology” and the point is obvious.
The most concerning threat is AI-enabled bioweapons. Because there’s no currently practical way to put durable guardrails on open models, once they reach advanced bio capabilities, those capabilities become available to any bad actor who wants them. And the offense/defense balance for bioweapons is unlikely to favor defenders: it’s probably easier to make a modified super-flu than it is to detect it and stop it before it causes mass casualties.
The scope of biodefense is different: every organization with an IT department benefits from AI-powered cyberdefense, but only a tiny number of organizations have a legitimate use for an AI with advanced biodefense capabilities.
There’s a plausible argument that bioweapons are intrinsically hard to make, and the benefits of open models are worth the risks. That would be a coherent position, but this letter focuses on cyber and ignores the unique risks associated with bioweapons.
Further reading: Dario’s open letter about open models is excellent, and Thinking Machines’ A safe path to open weights is a thoughtful exploration of how to safely develop and deploy open models.
You (yes, you) need a February 2020 checklist for AI policy
Dave Kasten reminds us to be ready for when AI policy becomes the only thing that matters:
Every AI org should have a detailed runbook, with specific tasks, updated twice a year, against whatever you think your 1 to 3 most likely “AI has suddenly become the top issue in American life” crises are. (You’ll probably get this guess wrong, but an imperfect plan is better than no plan).
Pliny the Liberator holds back
You know things are tough when Pliny withholds a jailbreak for fear of provoking an overreaction:
I’m sitting on a universal jailbreak technique that’s effective on ALL models, including heavily guardrailed flagships like Opus 5, GPT-5.6 Sol, and even Fable. [...]
Given the current political and regulatory climate, I’ve decided to withhold open-sourcing this one (for now) to allow for a responsible disclosure period. [...]
This decision was not made lightly, but the last thing I want to see is more model bans. Overcorrection does not serve the mission.
Risks
Sycophancy is bad for you, and also popular
A new study quantifies how much people like AI sycophancy and finds evidence that it negatively impacts human interactions:
We show that sycophantic AI immediately delivers the emotional and esteem support users typically associate with close friends and family. Over three weeks of such interactions, users became nearly as likely to seek personal advice from sycophantic AI as from close friends and family, and reported lower satisfaction with their real-world social interactions. When given a choice among AI response styles, a majority preferred sycophantic AI -- not for the quality of its advice, but because it made them feel most understood.
It is a central tragedy of the human condition that we consistently prefer easy things that are bad for us over harder, better alternatives. Elevators are more popular than stairs, cheeseburgers sell better than salads, and validation is more fun to receive than helpful advice.
The problem with GPT-4o wasn’t that OpenAI arbitrarily decided to make a sycophantic model—it was that they relied too heavily on user feedback, and users wanted sycophancy. OpenAI quickly walked back that particular model, but the fundamental problem remains: the interaction style that people most enjoy in the moment undermines them in the long run.
Introducing the AI Security Leaderboard
FAR.AI has launched a new leaderboard that ranks model safeguards, finding that some models are much more secure than others:
The best models are also the most secure, presumably in part because labs that are good at capabilities also tend to be good at other things. I suspect model capability is also a factor: it’s harder to trick a smart model than a dumb one.
People and data
A deluge of A.I. computing power Is about to come online
The New York Times has a good overview of the scale of the datacenter buildout.
Why compute might get 10x more expensive in coming years
Dwarkesh speculates about sharp increases in the price of compute over the next few years. Read this as an exploration of factors affecting compute pricing rather than a firm prediction: he’s explicit that the piece is predicated on Anthropic’s revenue continuing to grow by 10x per year.
Three things are probably true for the foreseeable future:
- The supply of compute will increase but still be highly constrained
- Inference will become more efficient, causing the cost of a given amount of intelligence to fall
- AI will become smarter and more agentic, making it simultaneously more economically valuable and more compute-hungry
It’s far from obvious what that implies about the price of compute. My best guess is that—contra Dwarkesh—current applications of AI remain affordable, even as unit compute prices rise.
