Anthropic reduced Claude Sonnet 5.5 prompt-cache reads from $0.20 to $0.10 per million tokens in its 7 October update. The official Claude Developers account repeated the change on 8 October. Current model documentation lists ordinary input at $2 and output at $10 per million tokens. The reduction applies to reading context already in the prompt cache.
Prompt caching reuses an existing prompt prefix. That prefix can include tools, system instructions, and messages. A coding agent repeatedly consulting the same instructions or background material is the kind of workload where this matters. Newly added material still has to be processed, and generated output still carries its normal price.
A simple example makes the change easier to see. Reading a cached prefix containing 100,000 tokens previously cost two cents; at the new rate it costs one cent. That calculation covers the cache-read portion alone. It excludes the initial cache write, fresh input, and the response. A long answer or frequent changes to the prompt can still dominate the total.
In its Haiku 5.5 announcement, Anthropic estimates that the reduction makes Sonnet 5.5 around twenty percent cheaper on most agentic tasks. That is the company's workload-level estimate. An individual application's saving depends on its cache-hit rate and the balance between repeated input and output.
Look at reuse before changing the system
The practical builder question is whether stable context is being reused. Instructions, tool definitions, and recurring reference material are candidates for a consistent prefix. Changing those elements unnecessarily can weaken the benefit.
A useful first check is to compare fresh-input, cache-write, cache-read, and output token usage across a real sequence of tasks. That reveals whether the application benefits from the new price before a larger redesign. Cheap repeated context is valuable when the work genuinely repeats it.