An outside inspector in a grey suit holding a folder, standing in a cold steel-and-glass laboratory corridor lit by overhead strip lights
EP 23September 15, 20266 min read

AI is going to kill us all

PodcastAI SafetyGovernanceRegulation
MW
Matthew J. Wozniak
September 15, 2026 · 6 min read

Dario Amodei says frontier AI is moving too fast for safeguards, and Anthropic is putting permanent outside evaluators inside the building. OpenAI says it will follow. My read on whether that's oversight, capture, or a speed limit only the big labs can afford. Companion notes to Human in the Loop Episode 23.

Anthropic CEO Dario Amodei says frontier AI is improving too quickly for safeguards to keep up. Anthropic's answer is to put permanent third-party evaluators inside the company, with employee-like access to systems, training processes, and incidents. Sam Altman says OpenAI will adopt the same access model.

The title of this episode is a joke, mostly. The reason Oscar and I spent an hour on this is that the story got flattened into "the AI labs agreed to slow down," and that is not what happened. Anthropic made a detailed commitment. OpenAI promised to follow but hasn't published implementation details. Elon Musk endorsed Amodei's argument without committing xAI to the evaluator program. Three very different levels of commitment, one headline.

So the question we kept circling: are embedded evaluators genuine oversight, regulatory capture, or a way for frontier labs to coordinate a speed limit that smaller competitors can't afford?

Signal or Noise

Oscar and I run every story through the same filter: is this signal you should act on, or noise dressed up as news? Here's how I called them this week.

Anthropic and OpenAI back embedded outside evaluators

This is the real thing, and I don't want to be cynical about it just because cynicism is easy. Employee-like access is a big deal. Evaluators who can see training runs and incidents are a different animal from evaluators who get a demo and a deck. My read: Signal, with a follow-through test. The test is whether OpenAI publishes the implementation details, and whether anyone outside the two companies can verify what access the evaluators actually got. Until then it's one commitment and one promise.

OpenAI's Navier–Stokes claim and the credit dispute

OpenAI published a Navier–Stokes solution claim. Then a statement from Buckmaster and WIRED's reporting on the academic response turned it into a dispute over credit. I'm not qualified to referee the math. I am qualified to notice the pattern: the attribution fight is going to show up every time a lab announces a scientific result, and the labs are not ready for it. Signal, and the credit dispute is the part to watch.

Meta Muse and its separate Sentinel approval agent

Meta launched Muse, its personal AI agent, alongside a separate Sentinel agent that handles approvals. The architecture is the story for me. Splitting the agent that does the work from the agent that approves the work is the pattern I'd copy, whatever you think of Meta. Signal.

DeepMind's predictions for 9 billion DNA variants

AlphaGenome Atlas publishes predictions for every possible single-letter change in the human genome. Predictions are not validations, and I'd want a biologist in the room before I got excited. But a public map at that scale is real. Signal.

California's new framework for AI auditors

Governor Newsom signed first-in-the-nation AI safeguards, including a framework for AI auditors. I'm glad it exists. My worry is the same one that drives the No Jargon segment below: an audit label only means something when the scope and the evidence are public. Signal, with limits.

No Jargon Required

Two terms your team is going to hear a lot this quarter.

  1. Independent AI audits. The word "audit" is doing a lot of work right now, and it usually arrives without the two things that make an audit worth anything: what was in scope, and what evidence was examined. When someone says a system was independently audited, ask for both. If they can't produce them, the label is marketing.
  2. Recursive self-improvement. The idea that a better model helps build its successor. Anthropic and OpenAI both published on it this week, and the reason it matters for the evaluator story is simple: if the loop speeds up, the safeguards have to sit inside the loop, not review it afterward. That's the strongest argument for embedded evaluators I heard all episode.

Where I landed

Oscar and I didn't end up in the same place on the motive question, and I don't think we need to. Here's the frame I walked away with:

  • Oversight is what this becomes if the access is real, the findings are reported somewhere outside the company, and OpenAI ships the details.
  • Capture is what it becomes if the evaluators are chosen by the labs, paid by the labs, and never contradict the labs in public.
  • A speed limit is what it becomes for everyone else if regulators start treating "we have embedded evaluators" as the bar, because only a handful of companies can afford to clear it.

All three can be true at once. That's the uncomfortable part.

Watch or listen to the full conversation below, and tell me which call you'd flip.


Human in the LoopWatch on YouTube · Listen on Spotify · Listen on Apple Podcasts

Every episode — full show notes, sources, and the archive — lives at podcast.vallyseed.com. The show is produced by VallySeed, where Oscar and I help organizations design, build, and deploy AI systems that create measurable competitive advantage.