Anthropic has released Claude Opus 4.6, a major upgrade to its most capable AI model with a 1 million token context window and much better coding ability.
What's new
Opus 4.6 can take in and produce much more text in a single session. The 1M token context window is in beta on the Claude Developer Platform, and output now goes up to 128k tokens, so the model can generate far longer responses. Anthropic has also made it more reliable on sustained multi-step agentic work and better at code review and debugging in large codebases.
Benchmark performance
The model sets new state-of-the-art results on several benchmarks. It has the highest score on Terminal-Bench 2.0, which measures agentic coding, and leads on Humanity's Last Exam for multidisciplinary reasoning. On GDPval-AA it outperforms GPT-5.2 by approximately 144 Elo points. In long-context retrieval it scores 76% on MRCR v2, compared with 18.5% for Sonnet 4.5.
Developer features
With adaptive thinking, the model decides for itself when extended reasoning will help, so developers no longer have to tune prompts by hand to get it. Effort controls give four levels (low, medium, high and max) for balancing intelligence against speed and cost.
During long tasks, context compaction automatically summarises older context to keep the model effective over extended sessions. Claude Code now supports agent teams, where multiple agents work together.
Safety
Anthropic reports low rates of misaligned behaviour in Opus 4.6 and the lowest over-refusal rate among recent Claude versions. In practice it is less likely to turn down legitimate requests while keeping its safety guardrails in place.
Pricing and availability
Claude Opus 4.6 is available now on claude.ai, the Anthropic API (model ID claude-opus-4-6) and major cloud platforms. Standard pricing stays at $5 per million input tokens and $25 per million output tokens. When input exceeds 200k tokens, the premium rate is $10 per million input tokens and $37.50 per million output tokens.