A CPG leadership team I sat with spent the better part of an hour arguing about which model to standardize on. GPT or Claude or Gemini. Someone had a slide with benchmark scores. Someone else had a procurement view. It was a good, serious debate, and it was the wrong one.
Here is the uncomfortable truth about that debate. By the time the slide was built, the benchmark had already moved. The leaderboard reshuffles almost weekly now. ChatGPT's own market share slipped below half for the first time this year. Whatever you standardize on today is a decision with a very short half-life.
The model is the most replaceable part of your AI stack. And most companies are pouring their strategy into the one component that commoditizes fastest.
Look at what the work actually shows. Bridgewater's AI lab ran a subjective judgment task past the frontier models. With basic prompts, the best models scored about 50 percent. A coin flip. Expert-crafted prompts pushed the same models into the mid-70s. Then they paid for the newer, pricier model. That upgrade bought them one to two points. Separately, they took a cheap open-weight model, fine-tuned it on expert-labeled data, and hit 85 percent at roughly a fourteenth of the cost.
Read that back slowly. The encoded expertise moved the needle 25 points. The shiny model upgrade moved it two. Gartner said the quiet part in April: AI project success comes down to integration, governance, and operational alignment, not model sophistication.
So if the model is not the moat, what is.
It is the stuff the vendor cannot ship in the box. Your first-party shopper data, the kind no frontier lab has and cannot buy. Your category judgment, the pattern recognition earned over decades of knowing why a launch works in one channel and dies in another. And the unglamorous plumbing, the integration and the governance and the workflow redesign that turns a clever demo into a number on the P&L. None of that comes pre-loaded. All of it is yours, and none of it commoditizes next Tuesday.
There is a second edge to this, and Satya Nadella named it well. He calls it the Reverse Information Paradox. The flow of learning runs one way. Every prompt you send teaches the provider a little more about how your company actually operates, and that accumulated knowledge becomes their advantage over time. Convenience-first adoption feels like speed. It can quietly be you renting your own edge back to the people who will sell it to your competitor next quarter.
For a data-rich CPG, that reframes the whole conversation. The question stops being "which model is best this month" and becomes "what are we building that the model cannot take with it when we switch." Your proprietary category and shopper knowledge is the asset. The model is a rental car. You would not build your business plan around which rental you booked this week.
So here is the one thing to do this week. Take your current AI roadmap and run a single test on it. If you swapped out the underlying model tomorrow, what would survive. The prompts tuned to your category. The proprietary data feeding the system. The workflows your team rebuilt around it. The governance that keeps your edge in-house.
If the honest answer is "not much would survive," you do not have an AI strategy. You have a subscription.
The companies that win the next few years will not be the ones who picked the right model. They will be the ones who spent this window building the thing the model can never be.
-- Imteaz
KEEP READING
If this hit a nerve, two earlier pieces on the same thread:
FORWARD IT
This one is for the operator standing in front of a model-selection slide, about to solve the wrong problem.
Or just hit reply and tell me one thing: what in your AI plan would actually survive swapping the model tomorrow? I read every one.
Talk to your AI tools the way you'd talk to a colleague.
You don't send a colleague a three-word brief. You explain the context, the constraints, what you've already tried. But typing all that into ChatGPT takes forever — so you don't.
Wispr Flow lets you speak your prompts instead. Talk through your thinking naturally and get clean, paste-ready text. No filler words. No cleanup. Just detailed prompts that actually get you useful answers on the first try.
Millions of users worldwide. Works system-wide on Mac, Windows, and iPhone.



