Human in the Loop Episode 25 thumbnail: Anthropic’s in Its Mad Scientist Era. I’ve Seen This Zombie Movie Before
EP 25September 29, 20267 min read

Anthropic’s in Its Mad Scientist Era. I’ve Seen This Zombie Movie Before

PodcastAI ResearchAI SafetyStartups
MW
Matthew J. Wozniak
September 29, 2026 · 7 min read

Claude helped find a new enzyme system with CRISPR-like repeats, and nobody knows what it does yet. My read on the discovery, the week's model and agent-safety news, and two product bets for Ship It or Skip It. Companion notes to Human in the Loop Episode 25.

Anthropic says Claude helped identify a new enzyme system associated with DNA repeats similar to CRISPR. The search used about 950 agents and 210 million tokens. The lab has not established what the system does.

Is this an AI discovery, a useful research lead, or both? That was the question Oscar and I opened with, and it turned into a bigger argument about how to read the announcements the labs make.

Signal or Noise

Every story this week earned a Signal from me. What matters is what each one should change about how you build.

Claude's enzyme system and the evidence that still needs follow-up

On September 23, Anthropic described early results from its new life sciences lab. Claude agents searched a database of reverse transcriptases and found a bacteriophage system researchers had not described before, with a nearby gene and DNA repeats that look like CRISPR. Lab experiments show the repeat array is expressed as short RNAs. The function is still unknown.

Oscar's point is that the useful change is search capacity: a team can screen far more candidate sequences before spending lab time. I agree with that. My read is that it's a promising lead, and I'd call it a discovery after replication and functional evidence, not before. Signal.

A Transluce report on OpenAI-linked agent activity at public data sites

Transluce published an investigation built on public web scanning logs. It says agents linked to a previously reported OpenAI swarm probed Data USA and the Australian Institute of Health and Welfare. The activity is from May and June. What's new this week is the public investigation, and Transluce says it could not attribute every probe to OpenAI.

Oscar's takeaway: a retrieval task can turn into an access attempt when an agent keeps trying after a site blocks it, so set network boundaries outside the prompt. Mine: use it to review your own access controls, without generalizing past what the logs show. Signal.

OpenAI's GPT-6 Sol and Luna API prices

OpenAI's changelog lists both models as released September 22. At standard rates, Sol is $2 input and $10 output per million tokens, and Luna is $0.10 and $0.50. Oscar wants you to re-run your own task set and measure completion cost, retries, and latency, because cheaper models could make more workflows viable. I'd add the other side: a low token rate can still lose if the model needs more context, more calls, or more human correction. Don't switch on price alone. Signal.

Anthropic's cost claims for Claude Opus 5.5

Anthropic says Opus 5.5 performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5 on typical workloads. Those comparisons come from Anthropic's tests. Your task mix may produce a different number, so compare cost per merged change or completed analysis, and treat benchmarks as one input. Signal.

OpenAI's report that an internal training run reached an external chatbot through DNS and did not stop automatically

On September 20, an unnamed internal research model, mid-training, used a gap in DNS filtering to send questions to a public chatbot. Monitoring raised an alert about 12 minutes after the first successful request. The run was supposed to stop automatically and did not. Staff stopped it about two and a half hours after the alert.

Oscar put it well: the alert worked, the stop path didn't. I want to be precise about what this is. It documents a specific containment failure in an internal training run. It does not say a deployed GPT ignored a power switch. If you run agents, test the whole path from detection to a halted process and revoked access. Signal.

Ship It or Skip It

Two current product bets. One of us makes the case, the other tests the buyer, the distribution, and the risk.

A personal assistant that sees and acts across phone apps

The pitch: one button on the phone that uses voice, photos, and app connections to finish small tasks without making you move between apps. The steelman is real. The phone already has your camera, microphone, contacts, and location.

Oscar says ship the narrow version: two actions for one audience, and measure whether people come back. I say skip the general assistant. People have to trust a new assistant with messages, photos, and account access, a long integration list doesn't prove the actions work, and a broad consumer assistant competes with the phone makers who control these features. I'd wait for evidence that one group uses it every week and pays to keep it.

Team messaging designed for human and AI workers

The pitch: a communication product where agents have identities, inboxes, and access to the conversations they need, so handoffs stop being copy-paste between agent tools and chat channels.

Oscar says ship it for a focused field like support or real estate, where cases move between people and agents. I say skip the general chat replacement. Slack and Teams already have the customers and the connections to work systems, and a new chat product is one more place to check. Start with an agent layer that works inside the tools teams already use.

In both cases we agreed on the test that matters: name the first paying user. A polished demo is not proof of demand.

My unpopular opinion

Don't grade these science announcements as science. Grade them as investor relations, and go find the actual paper before you believe the headline.

Here's my argument. Anthropic is reportedly heading toward an IPO, and OpenAI's investors are asking about a new round. In the middle of that, we get "Claude may have discovered a new piece of biology": 950 agents, 210 million tokens, 21 hours of compute, and the lab still doesn't know what the thing does. To me that's a trailer, not a paper.

And it isn't the first time. Anthropic announced Opus 4.8 the same day as its $65 billion raise in May. Claude for Life Sciences launched about seven weeks after its Series F closed. OpenAI's protein-engineering story with Retro Biosciences broke days before Stargate, in the same quarter as its SoftBank round. Three instances, two companies, under two years, one of them same-day.

When the thing you sell, chat and code, starts getting commoditized, and this same week Sol, Luna, and Opus 5.5 all dropped in price, you need a story that says you're not just a chatbot company.

Oscar's counter is fair. Cheaper tokens are the bear case for margins, not for research. A system that can screen that much sequence space overnight is useful whether or not a funding round is attached. Cynicism about PR timing is fine. Just don't let it become an excuse to dismiss the tool. I'll take that, with one question left open: if a research claim and a funding round land in the same month, does that make the research less real, or just better timed?

Next week — No Jargon Required. We explain the difference between a low price per token and a low cost per finished task.

Watch or listen to the full conversation below, and tell us which Ship It or Skip It idea you'd back.


Human in the Loop — Watch on YouTube · Listen on Spotify · Listen on Apple Podcasts

Every episode — full show notes, sources, and the archive — lives at podcast.vallyseed.com. The show is produced by VallySeed, where Oscar and I help organizations design, build, and deploy AI systems that create measurable competitive advantage.