Brings a 1M-token context window in beta, 128k output tokens, and the lead on Terminal-Bench 2.0 for agentic coding. Pricing stays at $5/$25 per million tokens, with a premium tier for prompts over 200k, relevant if you feed it whole repos.
Opus 4.6 and GPT-5.3-Codex ship the same day, Sonnet 4.6 becomes the new workhorse, Qwen3.5 adds open weights, and the Pentagon labels Anthropic a supply chain risk.
In February, coding agents became the frontier battleground for good: two flagships on the same day, tooling moving into managed VMs and remote sessions, and geopolitics starting to dictate who builds with what. Every link goes to the primary source.
Brings a 1M-token context window in beta, 128k output tokens, and the lead on Terminal-Bench 2.0 for agentic coding. Pricing stays at $5/$25 per million tokens, with a premium tier for prompts over 200k, relevant if you feed it whole repos.
The first model unifying the Codex and GPT-5 training stacks, around 25 percent faster than GPT-5.2-Codex and steerable mid-task instead of fire-and-forget. Launch was ChatGPT Pro only; the smaller Codex-Spark followed on February 12 as a research preview.
Near-Opus capability at unchanged Sonnet pricing of $3/$15 per million tokens, with 1M context in beta and hardened prompt-injection resistance. It became the default on claude.ai, and it is the price-performance point most agent workloads will actually run on.
A natively multimodal mixture-of-experts model with 397B parameters and only 17B active, combining gated-delta linear attention with sparse MoE across 201 languages. Smaller sizes followed on February 24, the month's biggest open-weights drop for self-hosters.
A reasoning step over Gemini 3 Pro: 77.1 percent verified on ARC-AGI-2, more than double its predecessor, and 94.3 on GPQA Diamond. Immediately available in the Gemini API, Vertex AI, and the CLI, and in public preview in GitHub Copilot the same day.
Each cloud agent runs in its own VM with a full dev environment and can run the software it builds, attaching videos, screenshots, and logs to the finished PR. Bugbot Autofix followed two days later, proposing tested fixes directly on PRs.
Remote Control drives a live local Claude Code session from web, iOS, or desktop while execution and the filesystem stay on your machine. Scheduled tasks run recurring agent jobs, initially only while the desktop app is open.
Jointly learns a hierarchical skill library and the policy that uses it, with the library co-evolving during RL; over 15 percent better than strong baselines on ALFWorld and WebShop. Directly relevant for agents that should accumulate reusable skills instead of relearning every time.
13 frontier labs signed the Frontier AI Impact Commitments. The concrete part for builders is infrastructure: large-scale data-center capacity from OpenAI and Tata, and a strategic agreement between Anthropic and Infosys.
Around 24,000 fraudulent accounts and over 16 million exchanges are attributed to DeepSeek, Moonshot, and MiniMax, targeting agentic reasoning and coding. The practical fallout for builders is tightened API access controls and stricter account verification.
After Anthropic refused to drop contract terms barring mass surveillance and fully autonomous weapons, the administration directed federal agencies to stop using it. If you sell into government-adjacent supply chains, model choice is now a compliance question.