Skip to main content

Claude Opus 4.8 New Features: Fast Mode, System Messages, effort

A practical look at what is new in Opus 4.8 — fast mode, mid-conversation system messages, the new effort default, prompt-caching improvements, and adaptive thinking.

Claude Opus 4.8 keeps the same tools and platform features as the previous-generation Opus 4.7, while adding a handful of new capabilities and behaviors. This guide focuses on the changes that are genuinely noticeable or immediately useful in development. For spec comparisons and what changed from 4.7, see the separate "Opus 4.8 complete guide."

🟢 Claude Opus 4.8 · model notice · Fable subscription
🟡 This article reflects Claude Opus 4.8. As of September 2026, the current top model is Claude Opus 5, and the full lineup is Claude Opus 5.5 / Claude Sonnet 5 / Claude Haiku 4.5 (higher tier: Claude Fable 5.1).

Fable 5 and 5.1 subscription (updated September 7, 2026): Claude Fable 5.1, released September 1, 2026, is the current Fable model and Fable 5 is now legacy. Plan terms are the same for both — Max and Team Premium plans include Fable at up to 50% of the weekly usage limit; Pro and Team Standard use usage credits (

🟡 This article reflects Claude Opus 4.8. As of September 2026, the current top model is Claude Opus 5, and the full lineup is Claude Opus 5.5 / Claude Sonnet 5 / Claude Haiku 4.5 (higher tier: Claude Fable 5.1).
0/M input, $50/M output tokens). The one-time
🟡 This article reflects Claude Opus 4.8. As of September 2026, the current top model is Claude Opus 5, and the full lineup is Claude Opus 5.5 / Claude Sonnet 5 / Claude Haiku 4.5 (higher tier: Claude Fable 5.1).
00 credit applied only to the Fable 5 transition and is not offered for 5.1. Some coding and debugging requests may be answered by an Opus model due to a security classifier (both models). See the Fable 5.1 guide and the Fable 5 availability guide for details.

What is new in Opus 4.8Fast modeOutput speedup to 2.5xAPI previewMid-convosystem messageinsertablecache kepteffortdefault highdeep reasoningadjustablePrompt cachemin 1,024tokensshort cachedAdaptivethinkingonly as neededoff by default

1. Fast mode — same model, faster output

Fast mode is offered as a research preview in the Claude API (not a general release). Set speed: "fast" on a request and the same Opus 4.8 model raises output tokens per second by up to 2.5x. In exchange, pricing is premium (higher). It helps when you need a long response quickly, but since cost rises, use it only on calls that truly need it.

2. Mid-conversation system messages — change instructions without breaking the cache

Until now, the system message (base instructions) could only sit at the very start of a conversation. From Opus 4.8 you can also place a role: "system" message after a user turn. As a result, adding new instructions midway through a long conversation preserves the earlier prompt cache, so you continue quickly without paying to recompute it. No beta header is needed. This is especially beneficial for long-running agentic work.

3. effort now defaults to 'high'

effort controls how deeply the model thinks before answering. In Opus 4.8 the default is high across all environments, including the Claude API and Claude Code. That means if you set nothing, it runs in the most thorough reasoning mode. If you want fast, lightweight responses, set effort lower explicitly.

4. Prompt cache minimum length is now 1,024 tokens

Prompt caching saves and reuses repeated leading content to cut cost and time. In Opus 4.7 the minimum cacheable length was longer, so short prompts were not cached. In Opus 4.8 that minimum drops to 1,024 tokens, bringing shorter prompts into caching range with no code changes.

5. Adaptive thinking — think only when needed

Opus 4.8 supports a single 'thinking' mode: adaptive thinking. When enabled (thinking: {type: "adaptive"}), the model decides each turn on its own — answering simple lookups directly and reasoning first on complex, multi-step problems. So it does not spend thinking tokens on easy tasks. Note that unless you explicitly enable it, thinking is off by default. The previous generation's 'extended thinking budget (budget_tokens)' approach is not supported.

Other improvements

According to official documentation, Opus 4.8 improves on 4.7 in long-horizon agentic coding (less compaction and better compaction recovery), reasoning-effort calibration, and fewer skipped tool calls. The stop_details carried in refusal responses is now officially documented, making it easier for applications to handle refusal types distinctly.

Opus 4.8 new features at a glancefast modeResearch preview, up to2.5x output speed(premium pricing)Mid-chat system msgAdd instructions mid-conversation, cache kepteffort defaults highDeepest reasoning modeunless you set itCache min 1,024 tokensShort prompts can nowbe cached tooAdaptive thinkingDecides per turn —reasons only when neededOther improvementsLong-horizon agent coding,fewer missed tool callsfast mode, system messages, effort and cache mainly matter for API work

Good to know

Among these, fast mode, system messages, effort, and prompt caching mainly apply when you build with the Claude API. If you use Opus 4.8 in regular chat (claude.ai), you will mostly feel improvements in response quality and speed and in the stability of long tasks. Also, like 4.7, Opus 4.8 does not support changing sampling values such as temperature, top_p, or top_k (defaults only).

Related: Claude Model Selection Guide.

Was this helpful?

Keep reading