Anthropic’s Claude Opus 4.7 features a new tokenizer that changes how the model processes inputs, according to analysis by OpenRouter published April 27, 2026. While Anthropic maintained the official price at $5 per million input tokens and $25 per million output tokens, the Opus 4.7 tokenizer inflation means users see higher costs for equivalent workloads. Anthropic disclosed a 1.0 to 1.35 times inflation range depending on content type.

OpenRouter analyzed over one million requests from users who switched from Opus 4.6 to Opus 4.7. The research compared token usage patterns across both models to measure real-world cost impacts. For prompts above 2,000 tokens, costs increased between 12 and 27 percent when prompt caching discounts are factored in. Short prompts under 2,000 tokens actually became more cost-efficient, with completions becoming 62 percent shorter, offsetting tokenizer overhead entirely.

How Tokenizer Inflation Works

OpenRouter uses two independent token counting methods for every request. The company’s own tokenizer, called QuadChars, groups every 4 printable ASCII characters as one token and counts each non-ASCII character separately. This consistent baseline isolates the tokenizer change from differences in prompt content.

Opus 4.7 produces 32 to 45 percent more native tokens than Opus 4.6 for equivalent text. Smaller prompts under 2,000 tokens see the highest inflation at 42 to 45 percent. Production-scale prompts between 10,000 and 128,000 tokens experience 32 to 34 percent inflation. The same tokenizer inflation applies to completion tokens, not just prompts.

Cache Absorption Reduces Token Cost Impact

Prompt caching absorbs a significant portion of the token inflation. Cached tokens are billed at a 90 percent discount, so extra tokens that land in cache create minimal cost impact. For the longest prompts exceeding 128,000 tokens, 93 percent of the extra tokens generated by the new tokenizer are captured by the cache.

Shorter prompts between 2,000 and 10,000 tokens see 56 percent of the token increase absorbed by caching. Prompts between 50,000 and 128,000 tokens achieve 77 percent cache absorption. Only very short prompts under 2,000 tokens show minimal cache benefits, with less than 10 percent of requests hitting the cache at all.

Completion Length Changes by Prompt Size

Opus 4.7 exhibits significantly different completion behavior based on prompt length. Short prompts under 2,000 tokens receive drastically shorter responses, with median completion lengths dropping 62 percent compared to Opus 4.6. For simple queries, the model generates fewer tokens in response.

Longer context prompts produce moderately longer responses. Prompts between 10,000 and 128,000 tokens see response lengths increase 13 to 30 percent at the median. Prompts exceeding 128,000 tokens see responses grow 26 percent longer. These changes combine with tokenizer inflation and caching to determine final costs.

Real-World Cost Breakdown

OpenRouter calculated average cost per million tokens across the user cohort that switched models. The 2K to 10K token range saw the highest cost increase at 27.2 percent, rising from $6.65 to $8.46 per million tokens. Prompts between 10,000 and 25,000 tokens increased 25.2 percent in cost.

Longer prompts saw more modest increases due to stronger cache absorption. Prompts between 50,000 and 128,000 tokens increased 11.9 percent in cost, while those exceeding 128,000 tokens rose 15.3 percent. The short prompt bucket under 2,000 tokens actually decreased slightly at negative 1.6 percent, because shorter response lengths more than offset the tokenizer inflation.

The analysis examined OpenRouter request logs excluding media files, cancelled requests, and zero-token requests. The methodology normalized costs by dividing by OpenRouter token counts rather than native token counts, controlling for prompt length differences across model versions and isolating the tokenizer change itself.

Source: OpenRouter