Last week, Anthropic broke my firm's operations.
Not with a bad model. Not with an outage. With an upgrade.
For over a month, we'd been running a core business process on Cowork - their desktop agent product. Locally, daily, maxing out our subscription. It worked. It compounded. It became part of how we operate.
Then they shipped an update to introduce cloud features, and everything broke. No migration path. No deprecation notice. No "pin to previous version" button. Our processes were dead for 24 hours, and then came the real bill: the hours we burned diagnosing, patching, and rebuilding what their release broke. Downtime you can measure. The cleanup tax is worse.
This isn't a complaint about Anthropic. Their model is excellent. Their team ships fast. This is about something structural that every founder and every LP needs to internalize - because what happened to us is coming for everyone building on the model layer's shiny new apps.
The model layer wants to be your app layer
Every frontier lab has reached the same conclusion: models are commoditizing, and the margin lives upstack. So OpenAI, Anthropic, and Google are racing to be the software - chat apps, coding agents, desktop agents. Rational. Probably inevitable.
But here's what nobody's pricing in. The entire promise of these agent products is long-running work: processes that hold state, run on schedules, accumulate context across weeks. A chat session is disposable - ask, answer, done. A business process is infrastructure. And the moment your process runs longer than their release cycle, you're in trouble, because their release cycle is measured in weeks and it ships without asking you. Every hour you invest tuning a workflow is an asset sitting on land that gets re-zoned quarterly.
These companies are research labs wearing product company costumes. Business software has a sacred, boring contract - it works the same way tomorrow as it did today - and frontier labs can't honor it, because at the frontier everything is provisional. The label admits it out loud: "research preview." Here's the tell: the same labs behave like adults on the API side, with pinned model strings, versioned deprecations, and SLAs. They know how to sell stability. They just don't sell it where they want you living.
And breakage is only half the bill. Satya Nadella warned this month that enterprises using proprietary AI pay twice: once in subscription fees, and again by handing over the business knowledge embedded in their prompts and corrections. Every workflow you tune inside their app is expertise flowing out of your company and into their next model. You're not just renting their software. You're training their next product.
They don't version. They don't warn. They don't migrate.
You're not a customer with a contract. You're a beta tester with a credit card.
The open-source alternative
Here's what makes this moment interesting: the same week Anthropic's update was breaking my processes, the alternative announced itself. Twice.
On Wednesday, Thinking Machines dropped Inkling, its first model. Open weights. A 975B mixture-of-experts you can download, fine-tune, and run yourself. And the detail that matters most: their own materials call it "not the strongest model available today, closed or open." A $12B lab led by OpenAI's former CTO shipped its first model and led with modesty - because the pitch isn't "smartest model." The pitch is that AI organizations shape and own beats one-size-fits-all AI that someone else keeps changing.
Two days later, Moonshot dropped Kimi K3. 2.8 trillion parameters. A million-token context window. Open weights. It trails Claude Fable 5 and GPT-5.6 Sol on overall benchmarks - but in Arena's blind testing, developers preferred K3 over every leading US model for front-end coding. From not "good enough" to preferred. At Sonnet-tier pricing.
The open ecosystem isn't trailing anymore. It's converging. And it hands you the one thing no closed model-layer app will ever sell you: you control the update button.
Pin the weights. Pin the orchestration. Pin the entire stack, run it for three years, and nothing changes unless you change it. When a better open model drops, you upgrade on your schedule - run it in staging, regression-test the outputs, roll it out when it passes or roll it back when it doesn't. You still ride the capability curve. You just ride it with a test window and an undo button. That's the real difference between open and closed at the application layer. Not frontier versus laggard. Forced updates versus chosen ones.
And for most business processes, the frontier premium was never the point. Your CRM doesn't need frontier reasoning. Your reporting pipeline doesn't need deep scientific level research. It needs the same good-enough answer, every time, forever. At the level of routine operations, models are already commodities - and you don't pay a premium, or accept a moving target, for a commodity.
Remember where your expertise goes inside a closed product? With open weights, it flows the other way. Fine-tune a model on your process and your knowledge compounds into an asset that sits on your balance sheet - weights you own, on hardware you choose, that no vendor can deprecate, reprice, or absorb. Bridgewater just proved the math: an open model fine-tuned on the fund's own expertise scored 84.7% on financial reasoning, beating top proprietary models at roughly a fourteenth of the cost. Closed AI turns your operational knowledge into their training data. Open AI turns it into your equity.
Is it free? No. Self-hosting converts vendor risk into operational risk - the weights are pinned, but the ecosystem around them churns, and somebody on your team owns that now. But the tooling gap that made self-hosting a 2023 horror story has mostly closed: mature inference servers, one-command deployment, fine-tuning platforms built for teams without a research lab. Hugging Face's Clem Delangue predicts frontier models get reserved for experimentation while production work shifts to open and private alternatives - and we've watched that exact split play out at every other layer of the stack. Operating systems ended in Linux. Databases ended in Postgres. Container orchestration ended in Kubernetes. Proprietary was better first. Open was good enough, then standard, then invisible.
Debian is boring. Boring is the feature.
What to do with this
For founders, here's the hierarchy of platform risk, safest to most reckless: open weights you self-host - if you have the team to carry the ops burden. Frontier models through versioned APIs, behind an orchestration layer that treats the model as a swappable part. Closed models glued directly into your product with no abstraction. And at the bottom: business processes inside someone else's consumer app. That last one isn't a strategy. It's a countdown.
The moat is the workflow, not the model. If you own the orchestration - the logic, the data flows, the process - the model underneath is a swappable part. If someone else owns the orchestration, you're the swappable part.
The bottom line
The labs will keep pushing into the app layer. It's where the money is. And their apps will keep breaking things, because that's what research velocity does. Both things are true. Your move is to stop pretending otherwise.
Use the frontier to explore. Own what you operate. Pin what you depend on. The gap between "most capable" and "most dependable" is where the next generation of great AI companies gets built - because someone has to sell certainty in an industry that only sells change.
We got everything running again. It cost a day of dead processes and far more in cleanup. Lesson bought and paid for.
Don't buy it twice.