AI Technology RadarAI Technology Radar

Model News H2 2025

knowledge
Adopt

The most relevant foundation-model releases of the second half of 2025.

Trends - what moved

  • Agentic capabilities became the main battleground. Every flagship was optimized for long autonomous runs, tool use and computer use: Claude Sonnet 4.5 sustained up to 30h agent runs, Google shipped Gemini 3 with an agent-first strategy, xAI added agent-tool APIs.
  • Chinese open-weight models closed the gap to months. Kimi K2 Thinking beat proprietary models on some agentic benchmarks, DeepSeek halved API prices via sparse attention, Qwen and GLM shipped continuously - and OpenAI answered with its first open weights since GPT-2 (gpt-oss).
  • Price war alongside the frontier race. Claude Opus 4.5 cut prices by two thirds while being the first model above 80% on SWE-bench Verified; Haiku 4.5 and Gemini 3 Flash brought near-frontier quality to the budget tier. At the same time OpenAI raised GPT-5.2 API prices by 40% - the price spread within one vendor lineup is widening.
  • The December escalation: Gemini 3 Pro's benchmark jump triggered OpenAI's internal "Code Red" and an accelerated GPT-5.2 release - frontier leadership changed hands within weeks, twice.
  • Meta went silent. No frontier release in H2 2025; Llama 4 Behemoth was shelved during the reorganization into Meta Superintelligence Labs.

Releases by player

  • OpenAI: GPT-5 (Aug, unified router system as ChatGPT default), GPT-5.1 (Nov, adaptive reasoning), GPT-5.2 (Dec, expert-level knowledge work) - openai.com/news
  • Anthropic: Claude Opus 4.1 (Aug), Sonnet 4.5 (Sep, coding/agent leader at release), Haiku 4.5 (Oct, Sonnet-4 level at a third of the price), Opus 4.5 (Nov, >80% SWE-bench Verified, price cut to $5/$25 per MTok) - anthropic.com/news
  • Google: Gemini 2.5 Flash Image "Nano Banana" (Aug, viral image editing), Gemini 3 Pro (Nov, major benchmark jump with Deep Think), Gemini 3 Flash (Dec, Pro-level reasoning at Flash pricing) - blog.google
  • xAI: Grok 4 (Jul), Grok 4.1 (Nov)
  • Amazon: Nova 2 family (Dec, price/performance positioning plus Nova Forge for custom models)
  • Alibaba: Qwen3-Max (Sep, first 1T+ model, closed weights)

Notable open-weight releases

  • gpt-oss-120b / 20b (OpenAI, USA) - first open OpenAI weights since GPT-2, Apache 2.0, runs locally
  • Kimi K2 / K2 Thinking (Moonshot AI, China) - 1T-parameter agentic MoE; the Thinking variant competed with proprietary models on agentic benchmarks
  • DeepSeek V3.1 / V3.2-Exp (DeepSeek, China) - hybrid thinking/non-thinking; V3.2 halved API prices via sparse attention
  • Mistral 3 / Large 3 (Mistral, France) - Europe's largest open-weight MoE, Apache 2.0