All articles
Research/September 29, 2026/7 min read

Your AI Cost Model Is Already Wrong: Tokenizers, Context Cliffs and Session Hours

Claude 4.7 and later use a new tokenizer that produces about 30 percent more tokens for the same text. The price per million went down. The number of millions went up.

TL;DR

  • Gemini 2.5 Flash Image shuts down on 2 October 2026. That is days away. If any .NET code still calls it, move now.
  • Claude 4.7 and later use a new tokenizer. It produces about 30 percent more tokens for the same text. Every cost estimate and every token counting check you wrote before that is now wrong.
  • OpenAI is on GPT-6. GPT-5.5 is capped at 272K context, and the GPT-6 models charge roughly double above 272K input tokens.
  • Token price is no longer the whole bill. Claude Managed Agents add $0.08 per session hour on top of tokens. Web search is billed per search.
  • Microsoft Agent Framework 1.0 shipped on 3 April 2026 and replaced Semantic Kernel agents and AutoGen. If you are still on the old packages, plan that migration.

What Changed

OpenAI

The lineup moved a full generation since July.

Model Input Output Note
gpt-6-astra $10 / MTok $50 / MTok $20 / $75 above 272K input
gpt-6.1-sol $2 / MTok $10 / MTok $4 / $15 above 272K
gpt-6-luna $0.10 / MTok $0.50 / MTok the cheap one
gpt-5.6-sol $4 / MTok $20 / MTok previous generation
gpt-5.5 $5 / MTok $30 / MTok limited to 272K context

Two things matter here more than the numbers.

First, the 272K line is now a price cliff, not just a limit. A request that goes one token over costs about twice as much per token. If you build prompts from a RAG pipeline where chunk count varies with the query, your cost per call is not stable. You need a hard cap in code, not a hope.

Second, processing modes multiply the same rates. Batch is half price. Fast is 2x. Ultrafast is 6x. A team that turns on Fast mode to fix a latency complaint can double the bill without anyone changing a model string.

The regional data residency uplift from earlier this year is still there: 10 percent extra on models released on or after 5 March 2026 when you call a regional endpoint. That hits GDPR scoped ASP.NET Core deployments directly.

Anthropic

The Claude 5 generation is out, and the pricing shape is different from OpenAI in one important way.

Model Input Output
Claude Fable 5.1 $10 / MTok $50 / MTok
Claude Opus 5.5 $4 / MTok $20 / MTok
Claude Opus 5 $5 / MTok $25 / MTok
Claude Sonnet 5.5 $2 / MTok $10 / MTok
Claude Sonnet 5 $2 / MTok $10 / MTok
Claude Haiku 4.5 $1 / MTok $5 / MTok

Claude 4.6 and later include the full 1M token context window at standard pricing. A 900K request costs the same per token as a 9K request. There is no cliff. If your workload is long context by nature, document analysis, large diffs, whole repository reads, that single line is worth more than a headline benchmark score.

The Sonnet 5 introductory rate of $2 and $10 was made permanent. The increase to $3 and $15 that was scheduled for 1 September 2026 did not happen.

Now the part that will surprise people.

Claude 4.7 and later models use a newer tokenizer that produces about 30 percent more tokens for the same text. Sonnet 4.6 and earlier use the old one. So if you upgraded from Sonnet 4.6 to a 5.x model and your bill went up more than the per token price suggested, this is why. The price per million went down. The number of millions went up.

This breaks three things in a typical .NET codebase:

  1. Any cost projection built from a token count you measured on an older model.
  2. Any middleware that counts tokens locally to enforce a budget before the call.
  3. Any chunking logic that packs a prompt to a fixed token target.

Re-measure. Do not assume.

Retirements on the first party API: Opus 4.1, Opus 4, Sonnet 4 and Haiku 3.5 are retired. Some are still reachable through Bedrock or Google Cloud, which is exactly the kind of difference that makes a model string work in one environment and fail in another.

Google

Model Input Output
gemini-3.8-flash $0.75 / MTok $3.75 / MTok
gemini-3.1-pro-preview $2.00 / MTok $12.00 / MTok
gemini-2.5-flash $0.30 / MTok $2.50 / MTok

The 3.8 Flash rate is introductory and runs through 31 December 2026. After that it doubles to $1.50 and $7.50. Put that date in the calendar rather than finding out in January.

Gemini 3.1 Pro tiers at 200K, not 272K. Above that it is $4 and $18. So the three vendors now have three different long context rules. Claude has none, OpenAI cuts at 272K, Google cuts at 200K. If you route between providers, the cheapest choice changes with prompt length.

Gemini 2.5 Flash Image is shutting down on 2 October 2026. This is the only item in this digest with a deadline inside the week.

.NET Tooling

Microsoft Agent Framework reached 1.0 on 3 April 2026. It merges what used to be Semantic Kernel agents and AutoGen into one supported SDK, with multi agent orchestration, MCP and A2A support, checkpointing and human in the loop built in.

Under it sits Microsoft.Extensions.AI. That is the layer worth caring about. Microsoft.Extensions.AI.Abstractions gives you IChatClient and IEmbeddingGenerator<TInput, TEmbedding>, and the Microsoft.Extensions.AI package adds the middleware: automatic function tool invocation, caching, and OpenTelemetry through UseOpenTelemetry on the chat client builder. There is also an IImageGenerator interface, still experimental.

Agent Framework 1.0 can build an agent on any inference service that provides an IChatClient, and ships connectors for Azure OpenAI, OpenAI, Anthropic, Bedrock, Gemini and Ollama.


Why It Matters

Write against IChatClient, not against a vendor SDK. This is the practical lesson of the last three months. Between July and today, OpenAI moved from GPT-5.5 to GPT-6, Anthropic released a new generation with a different tokenizer, and Google changed its context tiering. A codebase that depends on IChatClient handles that with a config change and a re-measure. A codebase full of OpenAIClient calls handles it with a refactor.

Go find your hardcoded model strings today. Not the interesting work, but the cheapest hour you will spend. A retired model does not degrade politely. It returns an error, usually in production, usually on a path nobody covered with a test. Grep the whole solution for model ids, including appsettings files, environment variable defaults and test fixtures.

Your cost model needs two new columns. Token price alone no longer predicts the bill. Claude Managed Agents bill $0.08 per session hour on top of tokens, and runtime only counts while the session is actually running, not while it waits for you. Web search is $10 per 1,000 searches on Anthropic. Code execution is free when used with web search or web fetch, and otherwise gives 1,550 free container hours per month before it starts charging. If you are running agents rather than single calls, model these as line items.

Prompt caching is where the real savings are now. A cache hit costs 10 percent of the input price on most Claude models, 5 percent on Opus 5.5, and 2.5 percent on Fable 5.1. On OpenAI most models cache input at 10 percent as well. For an agent that carries a large stable system prompt across many turns, this is the difference between a workable bill and a surprising one. It pays off after one read on the 5 minute cache and after two on the 1 hour cache.

Treat context length as a cost decision, not only a capability. With a cliff at 272K on OpenAI and 200K on Gemini, a RAG pipeline that grows its prompt with recall is a pipeline with an unpredictable bill. Cap the chunk budget in code and log the input token count per call so you can see the distribution instead of guessing at it.


Sources

Sources

  1. Pricing | OpenAI API
  2. Pricing | Claude Platform Docs
  3. Model deprecations | Claude Platform Docs
  4. Gemini Developer API pricing | Google AI for Developers
  5. Microsoft.Extensions.AI libraries | Microsoft Learn
  6. Agent based on any IChatClient | Microsoft Learn
  7. Releases, microsoft/agent-framework | GitHub
Share