The proxy problem
Since Opus itself is a black box, I used the closest public stand-in I could find. Kimi K3 is an open-weight model with 2.8 trillion parameters and a 1-million-token context window. Open weights means you can actually download the model files and run them yourself, which makes it a fair proxy for frontier-scale hardware math.
With a model that size on the table, I walked through the realistic options for local LLM hardware cost. A Mac Studio. Four H100s. A single DGX B200. A DGX B300. And finally a proper GPU cluster. Each step up buys you something specific, and each one changes the answer to "can I run Claude locally, or something like it."
The part nobody prices in
The single-user math is only half the story. The interesting question is what happens when 100 employees need concurrent access. Serving one careful engineer and serving a whole company are completely different hardware problems. That is where the shape of the answer changes, and it is the part I walk through in detail in the video.
Where I landed
Useful local AI is real. Smaller models on hardware you already own can do genuine work today. But frontier-scale private AI looks a lot more like a data-center project than a laptop feature. If your board is asking about self-hosted AI for privacy or compliance reasons, the honest framing is capital planning, not a weekend install.
Running this math for real?
I specialize in cloud and AI cost analysis for teams making exactly this build-versus-buy decision. If you want a second set of eyes on your numbers before you commit budget, book a call.
Book a call