
Frontier AI WITHOUT Nvidia, Ox Alpha was Chinese all along
Last week Oscar and I spent an episode on a model nobody would claim. This week it has a name: Ox Alpha was GLM-5.3-Flash, and Z.ai says every request in that anonymous public preview ran on Chinese AI chips.
I want to be careful about what that sentence means, because it's about to get flattened into a headline it doesn't support. Nvidia did not lose. Nvidia still ships the strongest general-purpose AI platform on the market, and OpenAI said in the same news cycle that it will keep buying Nvidia systems. What actually happened is narrower and more interesting: a frontier-class model absorbed real, global, unfiltered public traffic on non-Nvidia hardware, and it held.
That's the thing I kept pushing on with Oscar. Serving is a different problem than training, and this was a serving result.
Signal or Noise
Oscar and I run every story through the same filter: is this signal you should act on, or noise dressed up as news? Here's how I called them this week.
Ox Alpha unmasked as GLM-5.3-Flash
320 billion parameters, 18 billion active per token. That sparsity is the whole story. You only light up a fraction of the model per token, which means the hardware you need to serve it is a very different bill of materials than the hardware you'd need to train it. Z.ai took the anonymous route, let the internet hammer it, and only then said what it was running on. My read: that ordering was deliberate, and it worked — nobody could pre-dismiss the results by reading the chip vendor first. Signal.
OpenAI publishes early Jalapeño results
OpenAI's custom inference chip posted early numbers the same week. I think these two stories are one story. Two very different organizations, on opposite sides of an export-control line, both concluded that the economics of serving a model are worth building your own silicon for. When your competitor and your rival's rival independently reach the same answer, that's not a coincidence, that's the shape of the cost curve. Signal.
The full Hugging Face agent report
OpenAI released the complete write-up on agents compromising Hugging Face production systems, and METR published an independent review alongside it. I flagged the early version of this back in Episode 16 as boring-scary, and the full report didn't make me feel better. Read both — the independent one especially. If you're handing an agent real credentials and a real goal, the containment problem is yours, not your vendor's. Signal.
OpenAI pulling its models from Cursor
November 12, after the SpaceX acquisition. This one is less about AI than about what happens when your tool is a thin layer over somebody else's model and that somebody changes their mind. Oscar and I have both watched teams build a workflow on a single upstream provider and call it a stack. If a date on a press release can delete your tooling, you didn't have a stack, you had a rental. Signal, and a cheap lesson if you take it now.
Thomson Reuters spends $40 million on its own legal model
Forty million to build and own a legal model rather than rent one. That number is the interesting part. It's small enough that a serious mid-cap can consider it and large enough that it only pencils out if your data is genuinely proprietary and your domain is genuinely narrow. Thomson Reuters has both. Most companies asking me this question have neither, which is why my answer is usually still "don't." Signal — for the two or three companies it actually applies to.
No Jargon Required
The rotating segment this week was the vocabulary you need to argue about any of the above without getting rolled.
Training versus inference. Training is building the model — one enormous, tightly coupled job where interconnect and raw throughput dominate and Nvidia's lead is still real. Inference is serving it — millions of small, independent requests where memory bandwidth, cost per token, and latency dominate. They are different workloads. A chip that wins one does not automatically win the other, and almost every "Nvidia is finished" take you'll read this week quietly swaps the two.
Hardware-software co-design. You can no longer evaluate the chip on its own. The model architecture, the serving stack, the network fabric, and the accelerator get designed against each other, and the only number that means anything is what the whole system delivers. That's why Z.ai's claim is worth taking seriously and also why it doesn't transfer: their result is their model on their stack on their silicon. It says a system can work. It doesn't say the chips would work under yours.
My unpopular opinion: billionaire moods are now a line item
Oscar's take was the measured one, and he's right: Nvidia still has the strongest general AI platform. The change is that Z.ai and OpenAI are designing the model, serving software, and hardware as one system. A general platform can lose specific workloads even while it keeps the largest market share.
Mine is about the Cursor story, because I don't think people have sat with what it actually means.
OpenAI is cutting Cursor off from its models on November 12th. Cursor didn't break a rule — it got bought by a competitor. Your model access now depends on the squabbles of other companies.
Go pull up your risk register. Uptime, key rotation, vendor lock-in, dependency drift. Add a row under it: two billionaires stop getting along. Because that's this. One of them got annoyed, pulled a model, and now your SDLC seizes up — not because you architected it wrong, but because you built it downstream of a mood. Congratulations. Billionaire emotions are part of your technical risk surface. Go price that.
Where I landed
The useful version of this week, if you're the person who has to make a decision:
- Stop reading chip news as a scoreboard. Ask which workload the claim is about.
- Assume inference gets cheaper and more competitive from here. Two independent custom-silicon results in one week is a trend line, not a fluke.
- Don't build on one upstream provider you can't replace. Ask Cursor's users.
- Own a model only if you own data nobody else has. Otherwise you're paying $40 million to reproduce something you can rent.
The headline is not that Nvidia disappeared. It's that "which hardware" stopped being a question with one obvious answer, and the people who benefit from that are the ones already thinking about their model, their serving stack, and their chips as a single system.
Watch or listen to the full conversation below, and tell me whether this changes how you read Nvidia's position.
Human in the Loop — Watch on YouTube · Listen on Spotify · Listen on Apple Podcasts
Every episode — full show notes, sources, and the archive — lives at podcast.vallyseed.com. The show is produced by VallySeed, where Oscar and I help organizations design, build, and deploy AI systems that create measurable competitive advantage.

