In my last piece, AI Tokenomics, I made the case that intelligence has become a metered utility you have to govern. Here’s the sequel – because the price of that utility is now collapsing, and the reason should reorganize your strategy.
Executive Summary
- The shock: Frontier models now ship in weeks, not years – and the price of a given level of intelligence is falling by roughly an order of magnitude a year.
- The engine: This isn’t just smart labs working harder. AI is now an input to its own production – a flywheel where each generation helps build and optimize the next. Faster releases and falling prices are the symptoms; self-improvement is the cause, and it compounds.
- The catch: Cheaper per token isn’t cheaper per task, and not every lab is racing to the bottom (Kimi K3 just launched at frontier pricing). Judge on cost per successful task, not the sticker price.
- The inversion: When the model improves itself and its price is in freefall, the model stops being your moat. Advantage moves to what it can’t copy – your data, your governance, and how fast you can swap in the next one.
- The mandate: Architect for model mobility, not model loyalty. Make model choice a dial; put your durable capital into the system around it.

Intelligence vs. cost per task (log scale): at a similar intelligence score, cost per task spans a ~50x range, and near-frontier models increasingly fill the low-cost “most attractive quadrant.” (Source: Artificial Analysis, Intelligence vs. Cost per Intelligence Index Task.)
The engine: AI improving AI
Something broke the release calendar. A major model used to be a once-a-year keynote event; lately the industry ships new frontier and open-weight models close enough together that a strategy slide goes stale in a week. That’s the first curve. The second is quieter but decisive: the price of a unit of intelligence is falling off a cliff – by public analysis, roughly an order of magnitude a year for a fixed capability level, faster than the hardware curves most leaders built their intuitions on.
The executive question is why, because it decides whether to wait or to build. The answer: AI has started improving AI. It shows up in this month’s releases, not science fiction: OpenAI says it trained GPT-5.6 “to get more useful work from every token,” and reports its smaller models rival the prior frontier at a fraction of the cost; agentic coding loops now build software – and the tooling, data, and evaluations behind the next model – under human supervision; and open labs increasingly reuse each other’s methods and data, so every breakthrough becomes everyone’s starting line. Cheaper, better models make better tools; better tools make better models. That flywheel explains both curves at once – and it won’t stop on a convenient schedule.

The token-efficiency axis: at a similar intelligence score, models range from ~7,000 to over 70,000 output tokens to finish the same task – a tenfold spread. Cheap per token is not the same as cheap per task. (Source: Artificial Analysis, Intelligence vs. Output Tokens per Task.)
Cheaper isn’t simple
Two complications separate leaders who profit from the flywheel from those who just get a cheaper invoice.
Price per token is not price per task. A reasoning model can burn many times more tokens to finish the same job – Artificial Analysis shows a tenfold spread (roughly 7,000 to 70,000+ output tokens) at a similar intelligence score. A model that looks cheap per token can cost more per finished task than a pricier, terser one. Meter cost per successful task.
Not every lab is racing to the bottom. Moonshot AI’s Kimi K3 (July 16) – long the value pick – launched at roughly the price of a leading U.S. mid-tier frontier model, about triple its predecessor, while ranking among the top models on the index. The “cheap challenger” trade is no longer automatic. The deciding factors sit beyond the leaderboard: token efficiency on your workloads, plus data residency, compliance, and reliability for regulated enterprises.
The model is no longer the moat
When the best model is a moving target, its price is in freefall, and the technology improves itself, standardizing on one model is a long lease signed at the top of the market. The scarce resource has moved to what the model touches but doesn’t contain: your data and its governance, the system around the model (routing, evaluation, guardrails, and the cost control I’ve called AI Tokenomics), and your switching speed.
Falling prices only reach your P&L if your architecture can capture them – and most can’t, because one model is welded into the plumbing. The teams pulling ahead put a model router and a governed catalog between their apps and any single vendor – on the Microsoft stack, in Microsoft Foundry, fronted by the Azure API Management AI gateway – route each task to the cheapest model that clears a quality bar, and promote a better one behind an evaluation gate without touching the app. Model mobility turns every price cut and capability jump into a config change, not a migration.
What to do Monday
- Decouple. Put a gateway and model router between your apps and any single model – choice becomes config, not code.
- Re-evaluate on a cadence. The market re-prices monthly; your selection should too, on evidence.
- Default to the attractive quadrant. Near-frontier, low-cost model as the default; escalate to premium only where a task measurably needs it.
- Fund the durable layer. Data, governance, evaluation, cost control – the assets that don’t depreciate when a new model ships.
- Measure cost per successful task, not raw tokens or credits.
One caution: mobility isn’t free. A public benchmark leader can still fail your domain, so trust evals on your own data; switching needs a test harness, not just a config flag; and cheaper-per-unit intelligence often raises total spend by inviting more usage – which is exactly why governance matters more, not less.
The bottom line
We’re in one of the fastest price collapses of a valuable input we’ve seen – driven by the input improving itself – and it won’t wait for the dust to settle. Build for a world where intelligence is cheap, abundant, and better every quarter. The winners won’t be the organizations that picked the best model; they’ll be the ones who made the model the easy part and put their advantage everywhere else. When intelligence is manufactured by intelligence, the last scarce resource is human judgment about what is worth building.
FAQ
Is “AI improving AI” just singularity hype? No – it’s the mundane, observable version: models making tooling, data, evaluations, and other models incrementally better and cheaper. No science fiction required to reprice your budget.
How does this connect to AI Tokenomics? Two sides of one coin. The flywheel lowers the price per unit of intelligence; governance decides whether you capture the savings or lose them to runaway usage.
Go deeper
- Artificial Analysis – Intelligence vs. Cost (intelligence index and price-per-task map). https://artificialanalysis.ai/evaluations/artificial-analysis-intelligence-index
- Artificial Analysis – Intelligence vs. Output Tokens per Task (the token-efficiency axis). https://artificialanalysis.ai/evaluations/artificial-analysis-intelligence-index?eval-token-usage=score-vs-output-tokens-per-task
- Epoch AI – LLM inference price trends (the ~order-of-magnitude-per-year cost decline). https://epoch.ai/data-insights/llm-inference-price-trends
- Model router in Microsoft Foundry – route each task to the most cost-effective model that clears a quality bar. https://learn.microsoft.com/en-us/azure/ai-foundry/openai/concepts/model-router

Leave a Reply