Model News H2 2025
knowledgeAdopt
The most relevant foundation-model releases of the second half of 2025.
Trends - what moved
- Agentic capabilities became the main battleground. Every flagship was optimized for long autonomous runs, tool use and computer use: Claude Sonnet 4.5 sustained up to 30h agent runs, Google shipped Gemini 3 with an agent-first strategy, xAI added agent-tool APIs.
- Chinese open-weight models closed the gap to months. Kimi K2 Thinking beat proprietary models on some agentic benchmarks, DeepSeek halved API prices via sparse attention, Qwen and GLM shipped continuously - and OpenAI answered with its first open weights since GPT-2 (gpt-oss).
- Price war alongside the frontier race. Claude Opus 4.5 cut prices by two thirds while being the first model above 80% on SWE-bench Verified; Haiku 4.5 and Gemini 3 Flash brought near-frontier quality to the budget tier. At the same time OpenAI raised GPT-5.2 API prices by 40% - the price spread within one vendor lineup is widening.
- The December escalation: Gemini 3 Pro's benchmark jump triggered OpenAI's internal "Code Red" and an accelerated GPT-5.2 release - frontier leadership changed hands within weeks, twice.
- Meta went silent. No frontier release in H2 2025; Llama 4 Behemoth was shelved during the reorganization into Meta Superintelligence Labs.
Releases by player
- OpenAI: GPT-5 (Aug, unified router system as ChatGPT default), GPT-5.1 (Nov, adaptive reasoning), GPT-5.2 (Dec, expert-level knowledge work) - openai.com/news
- Anthropic: Claude Opus 4.1 (Aug), Sonnet 4.5 (Sep, coding/agent leader at release), Haiku 4.5 (Oct, Sonnet-4 level at a third of the price), Opus 4.5 (Nov, >80% SWE-bench Verified, price cut to $5/$25 per MTok) - anthropic.com/news
- Google: Gemini 2.5 Flash Image "Nano Banana" (Aug, viral image editing), Gemini 3 Pro (Nov, major benchmark jump with Deep Think), Gemini 3 Flash (Dec, Pro-level reasoning at Flash pricing) - blog.google
- xAI: Grok 4 (Jul), Grok 4.1 (Nov)
- Amazon: Nova 2 family (Dec, price/performance positioning plus Nova Forge for custom models)
- Alibaba: Qwen3-Max (Sep, first 1T+ model, closed weights)
Notable open-weight releases
- gpt-oss-120b / 20b (OpenAI, USA) - first open OpenAI weights since GPT-2, Apache 2.0, runs locally
- Kimi K2 / K2 Thinking (Moonshot AI, China) - 1T-parameter agentic MoE; the Thinking variant competed with proprietary models on agentic benchmarks
- DeepSeek V3.1 / V3.2-Exp (DeepSeek, China) - hybrid thinking/non-thinking; V3.2 halved API prices via sparse attention
- Mistral 3 / Large 3 (Mistral, France) - Europe's largest open-weight MoE, Apache 2.0