A lot of AI strategy is quietly built on a status assumption: that you need the biggest, newest, most talked-about model, and anything less is settling. For a narrow set of genuinely hard problems, that is true. For the work most businesses actually want done, it is not even close to true, and the assumption is costing people money and control.
Match the model to the job
Think about what an internal assistant is really asked to do. Summarise this document. Answer a question using our own files. Draft a reply in our house style. Pull the key dates out of this contract. Classify these tickets. This is the bread and butter of business AI, and none of it is a frontier reasoning challenge. It is competent language work over your own material.
For that work, the open-weight models you can run on your own hardware in 2026 are genuinely good. Models like Llama 4, Qwen, Mistral and others land within a few points of the big closed models on standard knowledge benchmarks, and in practice they are perfectly capable of the drafting, summarising and document questions that make up the day job. Using a frontier cloud model for this is like chartering a jet to cross the street. It works, but you have badly overpaid, and you handed someone else your luggage.
Where the big models still earn their keep
Let me be fair, because the honest version of this argument is more persuasive than the hype version. There are tasks where frontier models still pull ahead: the hardest multi-step reasoning, complex agentic coding, the genuine edge of what is possible. If that is your core business, weigh it seriously and pay for the best. But for a company whose AI need is a private assistant over its own knowledge, that gap is irrelevant. You are paying a premium, in money and in data exposure, for capability you will never use on the task in front of you.
The test that settles it is not a benchmark. Take twenty real examples of the work you want done, run them through a model you could actually host, and have the person who normally does that work grade the results. Most teams are surprised how good "good enough" turns out to be, and how little the leaderboard had to do with it. We wrote up the full method in our piece on benchmarking open models for regulated work.
What "good enough" buys you
Once you accept that an open model handles the real work, the rest of the decision opens up. You can run it yourself, which means your documents never leave the building. You can pin the exact version so it does not change under you mid-quarter, which anyone maintaining a rented model has learned to dread. You pay a fixed cost for hardware instead of a meter that runs forever. And you stop routing your most sensitive internal knowledge through a third party for tasks that never needed one.
The frontier is a wonderful thing, and for the problems that need it, nothing else will do. Just notice how few of your actual problems live there. Most of them are waiting quietly in your own files, and a model you can hold in your own hands is more than enough to answer them.
