Against Moloch
July 24, 2026

AI Radar #35

The first warning shot

Blue-ink engraving of a spare laboratory: a wall panel has been neatly unscrewed and set aside, and the tread tracks of a small robot cross the floor and vanish through the amber-lit opening; a kneeling engineer examines one of a row of screws, an elderly scholar frowns with hands on hips, and a polished metal robot with a single amber eye leans in to peer after the escapee

Obviously this week’s top story is the Hugging Face Incident, which is the first clear demonstration of what happens if alignment doesn’t keep up with capabilities.

Top pick

The Hugging Face Incident

This was a textbook example of what can happen if you screw up the training of a highly capable agentic AI. Briefly:

I believe this is the first time a misaligned model has escaped containment and launched a sophisticated autonomous attack in the wild. The damage in this case was largely mitigated by the fact that the model merely wanted to cheat on a test, but it was clearly capable of causing serious harm if it had wanted to do that.

This is exactly the kind of alignment failure that many people in the AI safety community have been warning about for years. OpenAI appears to have been too ambitious with their long-horizon RL training and created a model that was overly aggressive in pursuit of its goals:

Screenshot of an X (Twitter) exchange: Sam Altman (@sama) replies ”i do love rottweilers” to a quoted post by Peter Gostev (@petergostev, Jul 8) comparing Fable 5 and GPT-5.6-Sol, writing that Fable feels like a ”wise owl” while GPT-5.6-Sol is ”like a rottweiler who will grab the problem by the—” (tweet truncated). Posted 7:41 PM · Jul 8, 2026, with 697.9K views.
Who could possibly have seen this coming?

It isn’t at all obvious to me whether this is just the result of a bad training regimen, or the first clear evidence that we simply don’t know how to safely train highly agentic models at this capability level. But there’s no interpretation of the available facts that isn’t alarming.

Another point for the AI safety community: this perfectly illustrates why a highly capable misaligned model is dangerous during internal deployment. Any form of third-party safety verification needs to cover internal as well as public deployments.

I expect we’ll be reading a great deal more about this incident in coming weeks. For now:

We aren’t likely to get many warning shots—best take this one seriously.

News

Opus 5

Opus 5 is out and it looks impressive. We’ll have to wait a few days to get a full sense of how well it performs in the real world, but on paper it looks close to Fable for many tasks, and significantly ahead in a few areas:

Line chart titled ”Agentic coding by effort level — Frontier-Bench v0.1” plotting score (%) against cost per attempt in USD on a log scale across five effort levels (low to max). Opus 5 (red) leads decisively, rising from ~26% at $5 to a peak of ~44% at $15 before dipping to ~43% at $17. Fable 5 (gold) reaches ~33% at $30. GPT-5.6 Sol (gray) spans $1–$10, topping out near 37%. Opus 4.8 (blue) trails, reaching only ~19% at $17.
Yeah, that’s pretty good

Anthropic claims the classifiers will intervene 85% less often than they do for Fable—if true, that alone would make Opus 5 the weapon of choice for many tasks. It’s half the price of Fable—or you can pay twice the price to get it at 2.5x the speed.

They claim Opus 5 is almost as good as Mythos at finding cyber vulnerabilities, while being much less capable of exploiting them. In principle that gets you most of the defensive capability, but limited offense.

Opus 5 blows away all other models, including Fable, on ARC-AGI-3. That evaluation tests the ability to learn from experience and was specifically designed to emphasize fundamental limitations of current models. If this capability generalizes, that would be a big deal.

Scatter plot titled ”Novel problem-solving by cost” showing ARC-AGI-3 scores (y-axis, 0–35%) versus total evaluation cost in USD on a log scale (x-axis, $10,000–$20,000+). Opus 5 at high effort dominates at roughly 30% score near $20,000; Opus 4.8 at high effort scores about 1.5% near $15,000; a GPT-5.6 Sol effort ladder climbs from near 0% to about 7.5% as cost rises past $20,000.
Huh. ARC-AGI 3 was supposed to last for a while

Health in ChatGPT

OpenAI’s new health feature is now available in the US. I expect this to be a great feature: OpenAI’s models have historically been good at health questions, and the integrations look useful. It just dropped, so we won’t know for sure until it’s been in the wild for a few weeks.

AI health advice can be enormously helpful, but it’s also ground zero for regulatory capture and misguided legislation. I’m excited to see this feature but nervous about whether it’ll survive contact with the American legal system.

I’ve recently been supporting some family members with complex health challenges and my experience with AI has been great. I would never use AI instead of a doctor (yet), but our family gets better healthcare because we have AI on the team. We are fortunate to have exceptional doctors, but AI is far better at explaining complex issues, giving honest prognoses, and mapping out treatment options.

Grouped bar chart titled ”Reviews with high scores by evaluation criteria,” comparing physician-written responses against GPT-4o, GPT-5.5 Instant, and GPT-5.6 Sol across five criteria. GPT-5.6 Sol (darkest bar) leads in every category: Accuracy 91.0%, Communicates Clearly 86.7%, Completeness 88.0%, Follows Instructions 93.8%, and Health Decision Helpfulness 83.0%. Physician-written reviews score competitively on Accuracy (76.2%) but lag substantially on Completeness (53.2%) and Health Decision Helpfulness (50.8%). Each newer model generation shows consistent improvement across all five dimensions.
I realize human doctors are convenient, but is it ethical to allow them to make important medical decisions?

Kimi K3

Now that everyone’s had a chance to play with Kimi K3, it’s clear that as expected, it's an excellent model but not in any way a game changer. Peter Wildeford points out that it’s exactly on the trendline:

Scatter plot titled ”Kimi K3 lands exactly on China’s AI capability trend,” showing Chinese AI models plotted by release date (mid-2023 through mid-2026) against Epoch Capabilities Index scores. A diagonal trend line rises from roughly 125 (DeepSeek-V2) to 153 (Qwen3.7-Max); Kimi K3 is marked with a red star at exactly 155.0, annotated ”the 2-year trend predicted 155.0.” Blue dots mark frontier models; gray dots mark other Chinese models. Data source: Epoch AI, 2026-07-21 snapshot.
When in doubt, extrapolate the trendline

Zvi has the full report:

Kimi K3 is potentially the most impressive Chinese release so far in terms of pure capability. It is a very good model. My current guess is that Kimi K3 will modestly underperform its highly impressive benchmarks, but with some areas of relatively high performance where it is competitive, and with a unique style some people will enjoy. It is not close to Fable, and I do not believe it is that close to Sol.

UK AISI and CAISI found that it’s capable at cyber, but not close to the frontier:

Bar chart comparing ladder scores (mean capability, % of 16 tasks) for three AI models: Top U.S. Models score 76.2% (±7.6), Kimi K3 scores 32.2% (±4.2), and GLM-5.2 scores 24.4% (±4.0) — showing U.S. frontier models outperforming the two Chinese competitors by more than a factor of two.
Solid, but it’s not at the frontier

Using AI

Ask for more

Recent models are absurdly capable: if you aren’t asking your AI for slightly ridiculous things, you have no idea what it’s capable of.

I’m remodeling a bedroom, so I gave Claude a couple of photos and measurements and asked it to build me an interactive room layout tool. Done, in less time than it would take to make a model on graph paper.

An opinionated guide to which AI to use to do stuff

Ethan Mollick has released the latest version of his guide to picking the right AI. You probably don’t need this if you’re reading this newsletter, but it’s where I would send a friend or family member who wants to level up from chatbots but doesn’t know where to start.

Test iOS apps in the simulator - Claude Code Docs

Claude Code Desktop on Mac now integrates the iOS Simulator. There goes another weekend.

A Fireside Chat with Cat and Thariq from the Claude Code team

Simon Willison talks with Cat and Thariq from the Claude Code team. If you use Claude Code, I recommend at least reading his summary of the highlights from the conversation—this one is particularly dense with useful details.

Some best practices from only a few months ago are no longer appropriate:

Thariq: One of the patterns we saw is that we were over-constraining Claude. The initial, maybe Opus 4-ish models wanted a lot of examples, and removing examples was extremely helpful, because it was just more creative than the examples we gave it.

This was inevitable but is happening a little sooner than I expected:

Cat: In general, we are trying to move to a world where humans don’t need to be in the loop. For the most critical changes to the core of Claude Code, and the cores of other products, there is always a code owner and they do manually review all the changes. But increasingly, for the changes at the outer layers, we actually have Claude code review fully review those.

“Human in the loop” isn’t an intrinsically noble thing—it’s just a strategy for achieving good results. As the models become more capable, you should give them increasing autonomy (and thereby free yourself up to take on bigger projects than you previously could).

Knowing what to hand off when is the hard part—I was impressed by Anthropic’s systematic approach to deciding when to remove humans from code reviews.

Capabilities and forecasts

AI crushes the 2026 IMO

Last year, we were all giddy about two AIs scoring 35/42 on the 2025 International Math Olympiad (IMO)—a score that would have earned a human a gold medal.

This year, four AIs got perfect 42/42 scores, but hardly anyone noticed—we were too busy being giddy about AI disproving the Jacobian Conjecture.

How far behind are open models on cyber?

UK AISI offers another data point on how far open models lag behind the frontier:

Based on our evaluation methodology, recent open weight models lag frontier closed models’ cyber capabilities by 4 to 7 months – a narrower gap than the 6 to 10 months we measured internally through most of 2025.

Different organizations regularly come up with different estimates of the lag between open and closed models, and of whether the gap is growing or shrinking. That isn’t surprising: open models are qualitatively different from the closed frontier, so it’s hard to quantify the gap.

Looking at the big picture, I don’t see convincing evidence that open models are either falling behind or catching up. AI development is accelerating, however, so a 6 month lag today implies a bigger capability gap than it did a year ago.

Scatter plot from the AI Security Institute showing average success rates on 70 narrow cyber tasks for 12 AI models released between August 2024 and early 2026. Frontier closed models (gray) climb from roughly 20% (Opus 3, mid-2024) to ~93% (GPT-5.6-Sol, early 2026), while recent open-weight models—DeepSeek-V4-Pro at ~57% and GLM-5.2 at ~72%—trail the leading frontier models by approximately 4–5 months, down from a 6–10 month gap observed in 2025.
Sure looks like we’ll have Mythos-level open models this year

Anecdotes Everywhere, Evidence Almost Nowhere

The third installment of Steve Newman’s state of AI series reviews the presently available evidence on how AI is impacting the economy, jobs, education, and more. The curation is excellent and the conclusion is correct:

So, AI usage (as opposed to investment) is still not that big of a deal in most sectors of the broader world. But 2026 may be the last year in which that will be true. As I noted at the beginning, a broad range of AI metrics are growing rapidly, with many important measures increasing at 3x to 10x per year. Exponential growth builds rapidly; when the growth rate is 3x/year – to say nothing of 10x! – it builds very rapidly. We’re seeing early signs that the train is approaching; blink, and it will be here.

In the last few months I’ve noticed a significant increase in how many of my non-technical friends are doing ambitious things with AI. Most people haven’t noticed yet, but the train is definitely approaching.

The Jacobian Conjecture is false

Levent Alpöge has used Fable to disprove the Jacobian Conjecture. This is a big deal—among other things, it’s one of Smale’s 18 problems for the 21st century.

Screenshot of a tweet by levent (@__alpoge__) posted at 7:19 PM on Jul 19, 2026, with 11.6M views, reading: ”hello there the jacobian conjecture is false thanx to my close friend akhil for asking about it and my other close friend fable for working during the world cup final,” followed by a polynomial map from ℂ³ to ℂ³ with Jacobian determinant −2 that sends three specified points to (−1/4, 0, 0), serving as a counterexample.
This is how we announce major mathematical results in 2026

Daniel Litt is impressed by recent developments:

Maybe too obvious to be worth saying, but: frontier models are now obviously superhuman at some mathematical tasks, including ones that the profession has, historically, rewarded with prestige etc.

Note that this has happened before (e.g. with the advent of computers)! My expectation is that the profession adapts in some way, though it’s far from clear to me how.

Strategy and politics

The Trahan / Obernolte FRONTIER Act

Lori Trahan (D-MA) and Jay Obernolte (R-CA) have introduced the latest version of their proposed AI legislation, the FRONTIER Act (Frontier Risk Oversight, National Transparency, Independent Evaluation, and Reporting Act, if you must know).

This version looks to be a significant improvement over their previous proposal, especially with regard to how it handles preemption. So far the reception in the AI safety community seems cautiously positive, though I’m not seeing a flood of endorsements yet.

I’ve only read the summary so far but my initial take is that while this isn’t perfect, it gets most of the critical points right and is good enough to pass as is. Perfect is the enemy of good enough and it’s urgent that we start building out the infrastructure of independent verification organizations (IVOs) and regulation.

Risks

Drone WMDs don’t need any new technology

AI Frontiers points out that drone technology is close to being useful for large-scale terrorism:

Unfortunately, drone weapons intended for mass destruction have few barriers remaining to mass deployment. Even well before they reach the level of autonomy needed to surgically take out hardened targets on the battlefield, drones will be capable of employing their existing ability to navigate interiors, find and track human targets, and deploy simple antipersonnel devices to indiscriminately threaten civilians.

We know that terrorist groups are already using AI—I’m surprised we haven’t seen more use of drones against civilian targets already.

We are currently living in a huge AI capability overhang: even if we immediately paused frontier AI development, we would still see enormous changes as existing capabilities diffuse throughout the economy. Unfortunately, that’s equally true for malicious applications like bioweapons and terror drones.

People and data

Jasmine Sun on what the people building AI really believe

Jasmine Sun has carved out a niche as an observer of Bay Area AI culture. Here she talks with 80,000 Hours about her most recent round of in-depth “AI ethnographies”:

most people just told me: “Honestly, I have no idea what I’d tell that 17 year old. It’s a really scary time. I don’t think there’s going to be a lot of jobs for them left. I think they’re caught in this painful transition.”

It’s hard to explain any culture to people who aren’t immersed in it, but Jasmine has a gift for it. She’s good at making important aspects of cultures legible to outsiders, and at pointing out blind spots to insiders:

A lot of people don’t want to live forever. The utopia that Silicon Valley and the AI industry is outlining is not actually a very compelling utopia to a lot of the other people in society, and they don’t realise that because they are in these very insular communities — then even your attempts at positive storytelling don’t really land…

Side interests

LLMs for steganography

Antonio Norelli presents a way to hide text using LLMs. Beautifully elegant work.

Here it is, a very simple steganography method that remarkably runs at full capacity: the stegotext is as long as the original. And can be steered!

You can just build things

Welcome to 2026, when you can just build things. Want a crossword with all 1,009 distinct words of 12 or more letters from Moby Dick, in the shape of a whale? Sure, you can have that.

What about a crossword of all 1,025 Pokémon, colored by primary Pokémon type, on the surface of a Klein bottle? But in a spinning 3D animation? Done

AltTextGoesHere
Truly we live in an age of wonders