A small stack of glowing minimalist mini-computers on a desk radiating power, on a dark navy background
EP 5May 12, 20265 min read

Self-Hosted AI: Everything You Need to Know

PodcastSelf-HostingEconomics
MW
Matt Wozniak
May 12, 2026 · 5 min read

My notes from our guest deep-dive on running production AI on Mac Minis — killing the token bill, and when self-hosting actually beats the API. Companion notes to Human in the Loop Episode 5.

Anthropic and OpenAI are burning billions a year, and your token bill is on track to climb 4–8x over the next two years. Meanwhile, a profitable AI company is running production workloads for small businesses on a rack of Mac Minis — no token costs, 600–800K API calls a month on a single machine. So we brought the founder on to explain how.

Every fifth episode we drop the segments and go one topic, one expert. This one is a guest deep-dive with Jackson Oaks, founder of Recursion AI. Here's what stuck with me.

The token bill is the real threat

Most teams treat inference cost as a line item. Jackson treats it as an existential one — because for a business serving lots of small, repetitive requests, usage-based pricing scales with your success in exactly the wrong direction. The more customers you win, the more the meter runs. Self-hosting flips that: a fixed cost up front, then near-zero marginal cost per call.

Mac Minis are the punchline, but the point is architecture

The eye-catching part is the hardware — a stack of consumer machines doing what people assume needs a data center. But the real lesson is that a huge share of production AI work doesn't need the frontier. It needs a good-enough model, run efficiently, on hardware you own, for workloads you understand well enough to size.

When self-hosting actually wins

My honest read after the conversation: self-hosting is a fit when your workload is high-volume, repetitive, and latency-tolerant, and when you have the engineering discipline to run it. It is not a fit when you need the absolute frontier, unpredictable spiky load, or you don't have someone who wants to own the ops. The answer is boring and correct: it depends on your workload — but far more teams should be running the numbers than currently are.

What I took away

The frontier labs want you to believe the API is the only path. It isn't. For the right workload, owning your inference is one of the highest-leverage cost decisions a small AI company can make — and almost nobody talks about it because it isn't the exciting part.

Watch or listen to the full conversation with Jackson below.