
Your AI Cost Model Is Already Wrong: Tokenizers, Context Cliffs and Session Hours
Claude 4.7 and later use a new tokenizer that produces about 30 percent more tokens for the same text. The price per million went down. The number of millions went up.
TL;DR
- Gemini 2.5 Flash Image shuts down on 2 October 2026. That is days away. If any .NET code still calls it, move now.
- Claude 4.7 and later use a new tokenizer. It produces about 30 percent more tokens for the same text. Every cost estimate and every token counting check you wrote before that is now wrong.
- OpenAI is on GPT-6. GPT-5.5 is capped at 272K context, and the GPT-6 models charge roughly double above 272K input tokens.
- Token price is no longer the whole bill. Claude Managed Agents add $0.08 per session hour on top of tokens. Web search is billed per search.
- Microsoft Agent Framework 1.0 shipped on 3 April 2026 and replaced Semantic Kernel agents and AutoGen. If you are still on the old packages, plan that migration.
What Changed
OpenAI
The lineup moved a full generation since July.
| Model | Input | Output | Note |
|---|---|---|---|
| gpt-6-astra | $10 / MTok | $50 / MTok | $20 / $75 above 272K input |
| gpt-6.1-sol | $2 / MTok | $10 / MTok | $4 / $15 above 272K |
| gpt-6-luna | $0.10 / MTok | $0.50 / MTok | the cheap one |
| gpt-5.6-sol | $4 / MTok | $20 / MTok | previous generation |
| gpt-5.5 | $5 / MTok | $30 / MTok | limited to 272K context |
Two things matter here more than the numbers.
First, the 272K line is now a price cliff, not just a limit. A request that goes one token over costs about twice as much per token. If you build prompts from a RAG pipeline where chunk count varies with the query, your cost per call is not stable. You need a hard cap in code, not a hope.
Second, processing modes multiply the same rates. Batch is half price. Fast is 2x. Ultrafast is 6x. A team that turns on Fast mode to fix a latency complaint can double the bill without anyone changing a model string.
The regional data residency uplift from earlier this year is still there: 10 percent extra on models released on or after 5 March 2026 when you call a regional endpoint. That hits GDPR scoped ASP.NET Core deployments directly.
Anthropic
The Claude 5 generation is out, and the pricing shape is different from OpenAI in one important way.
| Model | Input | Output |
|---|---|---|
| Claude Fable 5.1 | $10 / MTok | $50 / MTok |
| Claude Opus 5.5 | $4 / MTok | $20 / MTok |
| Claude Opus 5 | $5 / MTok | $25 / MTok |
| Claude Sonnet 5.5 | $2 / MTok | $10 / MTok |
| Claude Sonnet 5 | $2 / MTok | $10 / MTok |
| Claude Haiku 4.5 | $1 / MTok | $5 / MTok |
Claude 4.6 and later include the full 1M token context window at standard pricing. A 900K request costs the same per token as a 9K request. There is no cliff. If your workload is long context by nature, document analysis, large diffs, whole repository reads, that single line is worth more than a headline benchmark score.
The Sonnet 5 introductory rate of $2 and $10 was made permanent. The increase to $3 and $15 that was scheduled for 1 September 2026 did not happen.
Now the part that will surprise people.
Claude 4.7 and later models use a newer tokenizer that produces about 30 percent more tokens for the same text. Sonnet 4.6 and earlier use the old one. So if you upgraded from Sonnet 4.6 to a 5.x model and your bill went up more than the per token price suggested, this is why. The price per million went down. The number of millions went up.
This breaks three things in a typical .NET codebase:
- Any cost projection built from a token count you measured on an older model.
- Any middleware that counts tokens locally to enforce a budget before the call.
- Any chunking logic that packs a prompt to a fixed token target.
Re-measure. Do not assume.
Retirements on the first party API: Opus 4.1, Opus 4, Sonnet 4 and Haiku 3.5 are retired. Some are still reachable through Bedrock or Google Cloud, which is exactly the kind of difference that makes a model string work in one environment and fail in another.
| Model | Input | Output |
|---|---|---|
| gemini-3.8-flash | $0.75 / MTok | $3.75 / MTok |
| gemini-3.1-pro-preview | $2.00 / MTok | $12.00 / MTok |
| gemini-2.5-flash | $0.30 / MTok | $2.50 / MTok |
The 3.8 Flash rate is introductory and runs through 31 December 2026. After that it doubles to $1.50 and $7.50. Put that date in the calendar rather than finding out in January.
Gemini 3.1 Pro tiers at 200K, not 272K. Above that it is $4 and $18. So the three vendors now have three different long context rules. Claude has none, OpenAI cuts at 272K, Google cuts at 200K. If you route between providers, the cheapest choice changes with prompt length.
Gemini 2.5 Flash Image is shutting down on 2 October 2026. This is the only item in this digest with a deadline inside the week.
.NET Tooling
Microsoft Agent Framework reached 1.0 on 3 April 2026. It merges what used to be Semantic Kernel agents and AutoGen into one supported SDK, with multi agent orchestration, MCP and A2A support, checkpointing and human in the loop built in.
Under it sits Microsoft.Extensions.AI. That is the layer worth caring about. Microsoft.Extensions.AI.Abstractions gives you IChatClient and IEmbeddingGenerator<TInput, TEmbedding>, and the Microsoft.Extensions.AI package adds the middleware: automatic function tool invocation, caching, and OpenTelemetry through UseOpenTelemetry on the chat client builder. There is also an IImageGenerator interface, still experimental.
Agent Framework 1.0 can build an agent on any inference service that provides an IChatClient, and ships connectors for Azure OpenAI, OpenAI, Anthropic, Bedrock, Gemini and Ollama.
Why It Matters
Write against IChatClient, not against a vendor SDK. This is the practical lesson of the last three months. Between July and today, OpenAI moved from GPT-5.5 to GPT-6, Anthropic released a new generation with a different tokenizer, and Google changed its context tiering. A codebase that depends on IChatClient handles that with a config change and a re-measure. A codebase full of OpenAIClient calls handles it with a refactor.
Go find your hardcoded model strings today. Not the interesting work, but the cheapest hour you will spend. A retired model does not degrade politely. It returns an error, usually in production, usually on a path nobody covered with a test. Grep the whole solution for model ids, including appsettings files, environment variable defaults and test fixtures.
Your cost model needs two new columns. Token price alone no longer predicts the bill. Claude Managed Agents bill $0.08 per session hour on top of tokens, and runtime only counts while the session is actually running, not while it waits for you. Web search is $10 per 1,000 searches on Anthropic. Code execution is free when used with web search or web fetch, and otherwise gives 1,550 free container hours per month before it starts charging. If you are running agents rather than single calls, model these as line items.
Prompt caching is where the real savings are now. A cache hit costs 10 percent of the input price on most Claude models, 5 percent on Opus 5.5, and 2.5 percent on Fable 5.1. On OpenAI most models cache input at 10 percent as well. For an agent that carries a large stable system prompt across many turns, this is the difference between a workable bill and a surprising one. It pays off after one read on the 5 minute cache and after two on the 1 hour cache.
Treat context length as a cost decision, not only a capability. With a cliff at 272K on OpenAI and 200K on Gemini, a RAG pipeline that grows its prompt with recall is a pipeline with an unpredictable bill. Cap the chunk budget in code and log the input token count per call so you can see the distribution instead of guessing at it.
Sources
- Pricing (OpenAI API Docs)
- Pricing (Claude Platform Docs)
- Model deprecations (Claude Platform Docs)
- Gemini Developer API pricing (Google AI for Developers)
- Microsoft.Extensions.AI libraries (Microsoft Learn)
- Agent based on any IChatClient (Microsoft Learn)
- Releases, microsoft/agent-framework (GitHub)
Sources
Keep reading

September 30, 2026 · 10 min
WebMCP: Letting Agents Call Your Web App Instead of Clicking It
Agents interact with your site by reading the DOM and guessing which element to click. WebMCP lets the page declare named tools instead. Here is the real API, the security model, and how it fits next to an MCP server in C#.
Read
September 28, 2026 · 7 min
Graph-Native Data Structures in C#, Part 6: DAGs and Dependency Graphs
Model a build system dependency graph in NebulaGraph, keep it acyclic from C# at insert time, get the execution order with Kahn's algorithm, and answer what is blocked right now. With real nGQL and C# code.
Read
September 28, 2026 · 7 min
Building an Agentic System in .NET, Part 4: Keeping Memory True
Stale agent memory is not just wasteful, it causes real failures that are hard to trace. This part builds the whole defence: timestamps, decay scoring, deduplication, contradiction detection, verification on read, and scheduled pruning as .NET hosted services.
Read