Menu di accessibilità (premi Invio per aprire)

September 30, 2026

AI token costs: why being able to switch models matters

Token prices drop with every release, yet enterprise AI spending keeps climbing. Companies tied to a single LLM miss out on price cuts and pay a high price for every migration.

The cost of AI tokens can vary up to a hundredfold from one model to another for the exact same task, and pricing shifts with every release: on September 22, 2026, OpenAI and Anthropic unveiled three new models on the very same day, all cheaper than their predecessors. Yet only organizations capable of shifting their workflows to the most cost-effective model without rewriting software can actually reap the benefits. Everyone else keeps paying yesterday’s rates.

Meanwhile, enterprise AI spending is growing faster than prices are falling. Far more than the model chosen today, what truly matters is how much it will cost to switch tomorrow.


What changed in AI model pricing?

According to reporting by Agenda Digitale, OpenAI introduced two lower-cost models alongside its flagship GPT-6 Astra. GPT-6 Sol, engineered for complex coding and agentic workflows, is priced at $2 per million input tokens and $10 per million output tokens; GPT-6 Luna, targeted at high-volume classification, extraction, and summarization, drops to $0.10 and $0.50. Anthropic launched Claude Opus 5.5 at $4 for input and $20 for output, a 20% cut compared to Opus 5.

The advertised savings require a closer look. OpenAI calculates its 50% reduction against the promotional pricing tier of the GPT-5.6 family. Anthropic cites an estimated 40% cost reduction on typical workloads, but this figure is a projection: it relies on the assumption that the model requires fewer tokens to complete tasks and on reduced prompt cache read rates, down from $0.50 to $0.20 per million tokens. Full pricing terms must also be considered. Beyond 272000 input tokens, Sol and Luna shift to long-context pricing tiers, regional processing adds a 10% premium for eligible models, and European Union data residency is only offered under standard processing. In practice, a regulated enterprise rarely pays the baseline headline rate.

The more meaningful comparison lies elsewhere. On a theoretical baseline workload of one million input tokens and 200000 output tokens (excluding caching, tool calls, or retries), Luna costs $0.20, Sol $4, Opus 5.5 $8, and Astra $20. True cost-effectiveness ultimately hinges on how many responses pass quality validation thresholds and how much manual rework is required to correct failures.


Why is AI spending growing while unit prices fall?

Over the past three years, the price per million tokens has dropped roughly a hundredfold, yet enterprise expenditure continues to surge: between early 2025 and early 2026, prices fell by approximately 80%, while enterprise consumption multiplied by 13x. This dynamic stems from a foundational nineteenth-century economic principle: Jevons’ paradox. When a resource becomes more efficient and less expensive to consume, aggregate demand rises rather than declines.

In the AI sector, this consumption surge is propelled by agents. A task that required 500 tokens in 2023 can easily consume 50000 tokens today because an agentic pipeline does not stop at a single generation: it reasons step by step, verifies intermediate outputs, queries databases, and triggers recursive model invocations. In Vercel AI Gateway data from May, requests ending in a tool call represented just 22.2% of total traffic but accounted for 58.9% of all consumed tokens. A negligible unit cost per call, multiplied across a multi-step execution chain, quickly compounds into a substantial operating expense.

Corporate financial disclosures already reflect this reality. Uber depleted its entire 2026 AI budget within the first four months of the year, driven primarily by agentic coding tools: prior to usage controls, individual engineers were racking up monthly bills between $500 and $2000, prompting the company to cap monthly spend at $1500 per employee per tool. Royal Bank of Canada saw its token consumption surge by 500% in six months. According to KPMG, only 26% of companies have clear visibility into their actual AI spending, while 22% discover it only when invoices arrive. These case studies are examined in our analysis of the hidden cost of tokens in corporate balance sheets.


Is the newest model always the best choice?

No. In Artificial Analysis’s Coding Agent Index, GPT-6 Sol climbs from 55 to 57 points while cutting cost per task roughly in half, whereas Luna drops from 43 to 41 despite becoming significantly cheaper. On office knowledge work benchmarks assessing tasks across 44 professions, both models show performance regressions compared to their predecessors.

A price list alone never reveals whether a model fits a specific operational workflow. The definitive operational metric is the cost per accepted result: the cost per support ticket resolved without reopening, per document field accurately extracted and validated, or per code change that passes automated tests and peer review. Measured this way, a premium model can prove far more economical if it minimizes errors and manual oversight, whereas an ultra-low-cost model wins decisively on deterministic, easily verifiable routines.

Selecting a model is therefore never a one-time decision. It must be reassessed with every market release, workflow by workflow.


Why is switching models so difficult for many enterprises?

Because software architectures were hardcoded around a single vendor. Even when underlying APIs appear standardized, proprietary tool calling implementations, memory management patterns, safety filters, and caching mechanisms couple the application directly to a specific provider. Every model migration turns into an engineering overhaul: rewriting code, redesigning evaluation pipelines, and securing fresh budget allocations. By the time that migration project wraps up, the next market price drop has already arrived.

Meanwhile, the range of alternative models companies miss out on continues to expand. Open-weight models, many developed in China, offer downloadable weights that can run on any chosen cloud provider or on-premises infrastructure. Over the three months leading up to September 22, open-weight models represented 77.1% of all token traffic routed through Vercel’s gateway, with DeepSeek V4.1 Flash alone capturing 58%, operating at API rates of $0.30 input and $1.20 output during peak hours. While these numbers reflect a single platform that recently broadened its open-weight taxonomy, they clearly demonstrate where high-volume workloads migrate once cost and performance thresholds are met. Running open-weight models carries its own cost structure: hardware provisioning, specialized talent, cybersecurity, service continuity, and, for Chinese models, geopolitical governance assessments.

Conversely, organizations face the opposite trap: uncontrolled model switching. When engineering teams upgrade from a lightweight model to a frontier model to chase marginal quality improvements without finance oversight, token costs can multiply tenfold to a hundredfold overnight. Enterprises face two symmetric operational risks: being trapped when switching makes economic sense, and switching haphazardly without governance.


How to build a governed model portfolio

Relying on a single model across an entire enterprise is no longer sustainable. Agenda Digitale recommends a tiered architecture: a lightweight or open-weight model for classification, standard summarization, and batch extraction; an intermediate model for coding and moderately complex agentic flows; and a frontier model reserved for ambiguous edge cases with high economic stakes. Traffic routing to the appropriate model can be orchestrated dynamically based on task type, data classification tier, latency SLAs, and runtime model confidence scores.

To operate effectively, this architecture demands key prerequisites established well before they are urgently needed:

  • provider-agnostic abstractions, version-controlled prompts, and standardized output schemas;
  • an evaluated fallback model ready for automated failover;
  • evaluation suites built on internal enterprise data, specific language edge cases, proprietary documentation, and known error modes, backed by strict acceptance thresholds and rollback playbooks;
  • deep process telemetry: cost per task, token distribution profiles, error rates, latency percentiles, and escalation ratios to higher-tier models;
  • granular cost allocation per individual request, mirroring established cloud chargeback frameworks, alongside automated spending quotas and approval workflows.

Model efficiency also depends heavily on consumption discipline. Across all three newly released models, output tokens cost five times more than input tokens: concise prompting, structured response schemas, verbosity limits, and prompt cache optimization impact total billing just as profoundly as model selection. Without granular telemetry, lower catalog pricing merely fuels higher call volumes and untracked expenditures.


How AVA eliminates model lock-in

AVA is Aidia’s proprietary enterprise AI platform that connects business applications to language models without binding operations to any single vendor. The underlying model becomes an interchangeable utility component: this is the foundation of AVA’s core architectural principle, no vendor lock-in.

AVA integrates natively with any LLM, whether open-source or commercial. When a more cost-effective model launches for high-volume tasks, or when an open-weight model is provisioned on dedicated enterprise infrastructure, companies can adopt it instantly without modifying application code. Dynamic routing ensures every single call reaches the optimal engine: routine classification never wastes frontier tokens, while mission-critical reasoning is never offloaded to models prone to failure. It is the tiered portfolio architecture described by Agenda Digitale, operationalized at request level.

New models deploy without engineering overhauls: existing ERP and CRM integrations, custom tools, and automated pipelines remain untouched while the underlying model shifts seamlessly. When the next price drop hits the market, capturing margin improvements requires no migration overhead.

AVA deploys on-premises with zero data leakage and fixed compute costs: enterprise data stays entirely within company borders, decoupling operations from volatile usage-based billing. For software engineering, the exact area where Uber exhausted its budget, AVA interfaces directly with enterprise codebases, providing centralized governance over developer permissions, authorized models, and token consumption limits.

Contact us to evaluate your enterprise workflows and discover how dependent your AI processes are on a single provider today.


Frequently asked questions

How much do tokens cost for the latest AI models?

Under September 2026 pricing, GPT-6 Luna costs $0.10 per million input tokens and $0.50 per million output tokens; GPT-6 Sol costs $2 and $10; Claude Opus 5.5 costs $4 and $20; and GPT-6 Astra costs $10 and $50. On a workload of one million input tokens and 200000 output tokens, the price spread between the most economical and most expensive model is a hundredfold.

Why does AI spending rise when token prices are falling?

Due to Jevons’ paradox: as a resource becomes cheaper and more efficient, total consumption increases substantially. Agentic workflows reason across multiple sequential steps and trigger repeated model invocations, meaning a task that consumed 500 tokens in 2023 can easily require 50000 tokens today. Between 2025 and 2026, unit prices fell by 80% while enterprise token volumes expanded thirteenfold.

What is vendor lock-in in AI models?

Vendor lock-in is technical dependency on a single model provider. Proprietary tool integrations, memory mechanisms, safety moderation layers, and caching APIs bind software applications to a specific vendor platform even when standard endpoints seem interchangeable. Switching models becomes a complex migration project, preventing the organization from capturing competitors’ price reductions or adopting superior specialized architectures.

Is it advisable to standardize on a single AI model across an enterprise?

Rarely. A lightweight model is sufficient for high-volume classification and extraction; a balanced intermediate model handles software development and agentic logic; and a frontier model should be reserved for high-stakes, ambiguous problem spaces. Model selection must be driven by cost per accepted result on specific workflows, measured against proprietary enterprise data rather than generic marketing benchmarks.

How can companies keep AI token costs under control?

Effective cost governance requires three operational pillars: attributing expenses to specific business units and use cases on a per-request basis; setting automated spending limits with workflow approval gates; and aligning IT and finance on shared FinOps chargeback frameworks. These controls should be paired with intelligent routing to lightweight models for repetitive tasks and end-to-end telemetry tracking error rates, latency, and escalation frequency.


Sources

Marta Magnini

Marta Magnini

Digital Marketing & Communication Assistant at Aidia, graduated in Communication Sciences and passionate about performing arts.

Aidia

At Aidia, we develop AI-based software solutions, NLP solutions, Big Data Analytics, and Data Science. Innovative solutions to optimize processes and streamline workflows. To learn more, contact us or send an email to info@aidia.it.