News
Releases, price changes, retirements and new capabilities, from each provider’s official pages and announcements. Summaries are ours; every item links to its source.
- Oct 3, 2026 xAI retires grok-voice-transcribe-1.0 grok-voice-transcribe-1.0 reached end of life on October 2, 2026. Requests to its slug are routed to grok-voice-transcribe-2.0 at the same price.
- Oct 2, 2026 Google retires Gemini 2.5 Flash Image Gemini 2.5 Flash Image is now retired.
- Oct 1, 2026 OpenAI deprecates gpt-4o-mini-tts (2025-12-15) and gpt-4o-mini-tts (2025-03-20) gpt-4o-mini-tts (2025-12-15) and gpt-4o-mini-tts (2025-03-20) are now deprecated and shut down on 2027-01-06.
- Oct 1, 2026 OpenAI deprecates tts-1, tts-1-hd, GPT-5.4 nano, GPT-5.3 Codex and GPT-5.1 tts-1, tts-1-hd, GPT-5.4 nano, GPT-5.3 Codex and GPT-5.1 are now deprecated and shut down on 2027-01-06.
- Oct 1, 2026 OpenAI retires gpt-5.4-cyber gpt-5.4-cyber is now retired.
- Sep 30, 2026 Alibaba Qwen retires deepseek-r1-distill-llama-8b deepseek-r1-distill-llama-8b is now retired.
- Sep 30, 2026 Anthropic schedules Claude Sonnet 4.5 retirement Anthropic plans to retire Claude Sonnet 4.5 (claude-sonnet-4-5-20250929) from the Claude API on November 30, 2026. It recommends migrating to Claude Sonnet 5.5.
- Sep 30, 2026 Black Forest Labs releases FLUX 3 Image FLUX 3 Image is available through one endpoint for image generation and editing, with up to ten reference images and output up to 4K. Per-image prices range from $0.041 at 768sq to $0.607 at 4K.
- Sep 30, 2026 Google announces Gemini 4 Argon Google announced Gemini 4 Argon, with an output token limit of 1 million. It is rolling out to trusted cyber defenders and is not yet available to developers; Google says developer access will come later.
- Sep 30, 2026 Google retires gemini-omni-flash-preview gemini-omni-flash-preview is now retired.
- Sep 30, 2026 Runway adds Eleven v4 to Runway Dev Eleven v4 is available on Runway Dev for speech generation, with scripts up to 2,500 characters. It costs 2.2 credits per 1,000 characters through October 12, 2026 PT, then 5 credits, with a 1-credit minimum.
- Sep 29, 2026 Inception makes Mercury Voice generally available Inception made Mercury Voice generally available to enterprise customers. It supports 128K-token contexts, up to 50K output tokens and three reasoning settings. It costs $0.40 per 1M input tokens and $1.50 per 1M output tokens, with 50% off at launch ($0.20 and $0.75).
- Sep 29, 2026 MiniMax releases minimax-m3.1-flash-preview 1M-token context window.
- Sep 29, 2026 Mistral AI retires Leanstral 1.5 and GLM 5.2 Leanstral 1.5 retires on September 30, 2026; Z.ai GLM 5.2 retires on October 31, 2026. Mistral AI recommends Z.ai GLM 5.3 as its replacement at the same price.
- Sep 29, 2026 OpenAI adds computer use to the Agents API OpenAI added computer use to the Agents API, allowing agents to complete tasks in an OpenAI-hosted browser. Applications handle website access approvals and sign-in.
- Sep 29, 2026 OpenAI adds Ultrafast mode for GPT-6 Astra OpenAI added Ultrafast mode for GPT-6 Astra in the Responses API. API customers can use the ultrafast service tier to reduce the time between generated output tokens, subject to rate limits.
- Sep 29, 2026 OpenAI releases GPT-6.1 Sol OpenAI released GPT-6.1 Sol (gpt-6.1-sol) for complex coding and professional work. Standard prices for prompts of up to 272K input tokens are $2 per 1M input tokens, $0.10 cached input and $10 output.
- Sep 28, 2026 Anthropic releases Claude Sonnet 5.5 Anthropic launched Claude Sonnet 5.5 (claude-sonnet-5-5) on the Claude API, Amazon Bedrock, Claude Platform on AWS, Google Cloud, and Microsoft Foundry. Context window, output limits, and prices are on its model page.
- Sep 28, 2026 ElevenLabs releases Eleven v4 and Eleven v4 Turbo Eleven v4 and Eleven v4 Turbo are now available. Eleven v4 supports voice cloning in more than 90 languages; Turbo is intended for real-time use and has approximately 100 ms median inference latency.
- Sep 25, 2026 Meituan releases LongCat-2.5-Preview 1M-token context window. On LongCat API: $0.30 input, $1.20 output per 1M tokens.
- Sep 25, 2026 Perplexity cuts fast preset search price The Agent API fast preset now uses Fast Search. web_search falls from $2.50 to $1.00 per 1,000 invocations and is about 800 ms faster.
- Sep 24, 2026 Perplexity adds Fast Search option Search API and Agent API web_search can set search_type to fast for a lower-latency path at $1.00 per 1,000 requests or invocations. Agent API model tokens are billed separately; search_type web is standard search.
- Sep 23, 2026 Black Forest Labs releases FLUX 3 Action Black Forest Labs released FLUX 3 Action, an open-weight 7B world-action model from its FLUX 3 backbone, with weights available. On RoboLab-120 it uses under half the parameters of the prior best open model and runs up to 3.95x faster.
- Sep 23, 2026 Meta introduces Muse Realtime Avatar On September 23, 2026, Meta introduced Muse Realtime Avatar. It uses Muse Realtime Voice speech tokens and reference media to generate live, synchronized facial, hand, and full-body video.
- Sep 22, 2026 Anthropic adds mid-conversation tool definitions On the Claude API, the inline-tools-2026-09-15 beta header lets a system message added mid-conversation carry a full tool definition or, with the MCP connector beta header, an MCP toolset, without editing the tools list.
- Sep 22, 2026 Anthropic previews fast mode for Claude Opus 5.5 Fast mode is available as a research preview for Claude Opus 5.5 on the Claude API.
- Sep 22, 2026 Anthropic releases Claude Opus 5.5 Anthropic released Claude Opus 5.5 (claude-opus-5-5): 1M context, 128k max output, always-on adaptive thinking, at $4/$20 per MTok (Claude Opus 5 is $5/$25). It is on the Claude API, Amazon Bedrock, AWS, Google Cloud, and Microsoft Foundry.
- Sep 22, 2026 Google releases Gemini 3.8 Flash TTS models Gemini 3.8 Flash TTS (gemini-3.8-flash-tts) and Gemini 3.8 Flash-Lite TTS (gemini-3.8-flash-lite-tts) are generally available, with a Gemini API Voices endpoint. Flash-Lite TTS is meant to replace gemini-3.1-flash-tts-preview.
- Sep 22, 2026 Xiaomi releases MiMo-V2.6-Flash 1M-token context window. On Xiaomi MiMo API: $0.14 input, $0.28 output per 1M tokens.
- Sep 22, 2026 Xiaomi releases MiMo-V2.6-Pro 1M-token context window. On Xiaomi MiMo API: $0.435 input, $0.87 output per 1M tokens.
- Sep 22, 2026 Xiaomi releases MiMo-V2.6-Pro-UltraSpeed 1M-token context window. On Xiaomi MiMo API: $4.35 input, $8.70 output per 1M tokens.
- Sep 22, 2026 OpenAI improves prompt caching for GPT-6 OpenAI says GPT-6 models hit the prompt cache more often by default, and discounts now apply to shared prefixes reused within 30 minutes. New tools include a caching dashboard, a diagnostics tool that explains cache misses, explicit cache breakpoints, and changing reasoning effort mid-conversation without losing the cache.
- Sep 22, 2026 OpenAI releases GPT-6 Sol and GPT-6 Luna OpenAI released GPT-6 Sol (gpt-6-sol) and GPT-6 Luna (gpt-6-luna), reasoning models that take text and images and return text through the Responses and Chat Completions APIs. Standard prices per 1M tokens for prompts of up to 272K input tokens: Sol $2 input, $0.20 cached input, $10 output; Luna $0.10, $0.01 and $0.50.
- Sep 21, 2026 xAI releases Grok 4.7 500K-token context window. On xAI API: $2.00 input, $6.00 output per 1M tokens.
- Sep 21, 2026 Alibaba Qwen releases qwen3.8-omni-flash-realtime 197K-token context window.
- Sep 20, 2026 Qwen releases Qwen-Image-2.1 Qwen open-sourced Qwen-Image-2.1 on September 20, 2026. The 7B-parameter image model combines text-to-image generation and image editing, with native support for transparent images.
- Sep 20, 2026 Alibaba Qwen releases qwen-audio-3.1-realtime-plus 262K-token context window.
- Sep 18, 2026 Qwen releases Qwen3.8-LiveTranslate Qwen announced Qwen3.8-LiveTranslate on September 18, 2026, describing an Interleave architecture for real-time simultaneous interpretation. Its average lagging metric drops from 2.8 seconds to 2.3 seconds.
- Sep 18, 2026 Google limits access to Gemini 2.5 models Google is limiting Gemini 2.5 API access to users who have actively used those models. They are not deprecated and remain available until further notice. New projects should use 3.5 Flash-Lite or 3.8 Flash.
- Sep 17, 2026 Google releases Antigravity Agent 09-2026 Google released antigravity-preview-09-2026, replacing antigravity-preview-05-2026. Local tool parameters and file edits changed; remote output-only use needs only the new agent string. The May preview shuts down on October 5, 2026.
- Sep 17, 2026 Alibaba Qwen releases qwen3.8-omni-flash 1M-token context window. On Alibaba Cloud Model Studio: $0.15 input, $0.47 output per 1M tokens.
- Sep 17, 2026 Runway adds Enhance Frame Rate on Dev Runway Dev can now convert video to 24, 25, 30, 48, 50, 60, 120, 23.98, 29.97, or 59.94 fps. Inputs are at most 300 seconds, billed at 1 credit per 2 seconds via the video upscale endpoint with model enhance_frame_rate.
- Sep 16, 2026 Alibaba Qwen releases qwen-mt-uni qwen-mt-uni is now available.
- Sep 15, 2026 Google releases Gemini 3.8 Live models Gemini 3.8 Live (gemini-3.8-live) and Gemini 3.8 Live Extended Thinking (gemini-3.8-live-extended-thinking) are generally available as audio-to-audio models on the Live API.
- Sep 15, 2026 TypeSafe releases Jev (preview) 64K-token context window. On TypeSafe API: $0.042 input, $0.00 output per 1M tokens.
- Sep 15, 2026 Kling AI releases Virtual Try-On 3.0 Virtual Try-On 3.0 improves face consistency and image quality. It accepts flat-lay, mannequin, and on-model clothing images, and can lock pose, keep the face, and choose whether to keep the original background. V1 and V1.5 parameters map to the 3.0 pipeline.
- Sep 14, 2026 Anthropic adds on-demand compaction to Messages API Anthropic's Messages API now supports on-demand conversation compaction in beta via the compact-2026-09-04 header. A compaction parameter returns a signed summary block you can send later instead of the original messages, while keeping recent turns verbatim.
- Sep 13, 2026 Perplexity adds custom MCP connectors A Project can register a remote MCP server once and Perplexity stores its credential. Agent API calls use the connector ID with type connector. API-key or no auth, and Streamable HTTP or SSE, are supported.
- Sep 11, 2026 ElevenLabs makes Scribe v2 Medical generally available Scribe v2 Medical is now generally available for batch medical and clinical speech recognition. It is billed at the same rate as Scribe v2.
- Sep 10, 2026 Black Forest Labs adds 2K and 4K FLUX 3 Video The FLUX 3 Video endpoint now returns qhd (2560×1440) and uhd (3840×2176) clips in one request, and draft_enhance can commit a draft at either size. Added per-second prices are $0.40 and $0.80 for text or image input, and $0.65 and $0.95 for continuation.
- Sep 10, 2026 DeepSeek releases DeepSeek-V4.1-Flash DeepSeek released DeepSeek-V4.1-Flash, its smallest new-architecture model with native multimodal vision, callable as deepseek-flash. The retired names deepseek-v4-flash and deepseek-v4-flash-vision-exp route to it for now. V4 Pro continues after 14 September 2026 with unchanged billing, and API prices were reduced with the release.
- Sep 10, 2026 OpenAI makes GPT-Live 1 generally available GPT-Live 1 is generally available for full-duplex voice sessions that continue while a backend model or agent handles reasoning and tools. Voice costs $0.05 per minute, billed per second; model and tool use is charged separately.
- Sep 10, 2026 OpenAI releases Agents API in public beta OpenAI released the Agents API in public beta, with a managed Codex harness for session orchestration, context compaction, and recovery. Sessions can stream progress, use custom tools and MCP servers, and run in OpenAI or customer sandboxes.
- Sep 8, 2026 Inception releases Mercury 2.5 Inception released Mercury 2.5, a quality step up from Mercury 2, with a 260K-token context and 1,107 tokens per second. List price is $0.20 per million input tokens and $0.75 per million output tokens, 80% off at launch.
- Sep 8, 2026 OpenAI makes GPT-Rosalind generally available OpenAI made GPT-Rosalind (gpt-rosalind-research) generally available for approved internal life sciences research under its trusted-access program. Prices are $5, $0.50 cached, and $25 per 1M input, cached input, and output tokens. Billing starts October 5, 2026.
- Sep 8, 2026 OpenAI releases GPT Image 2.5 Sunburst and Flare OpenAI released GPT Image 2.5 Sunburst and GPT Image 2.5 Flare for generation and editing via the Image API and Responses image tool. Both add xhigh and max quality and use GPT Image 2 token rates.
- Sep 7, 2026 Kling AI launches Video Commerce API Kling AI launched Video Commerce via API and Agent. A character image plus script makes a talking-head video; adding product images produces promotion shots. Options cover audio speed, aspect ratio, resolution, voice, speech rate, and background music.
- Sep 4, 2026 Ant Group releases Ling-3.0-flash-VL On September 4, 2026, Ant Group launched multimodal Ling-3.0-flash-VL. It can be tried in the chat UI and called through OpenAI-compatible and Anthropic-compatible APIs.
- Sep 3, 2026 Google releases Lyria 3.5 music model Lyria 3.5 (lyria-3.5) is generally available for full-length song generation. It accepts text and image inputs and outputs 44.1 kHz stereo audio, with duration and structure controls.
- Sep 3, 2026 OpenAI adds long-running controls for GPT-6 Astra The Responses API now offers async tool calling, mid-turn steering over WebSockets, and mid-conversation reasoning-effort changes for GPT-6 Astra, while preserving the cached prompt prefix.
- Sep 3, 2026 OpenAI releases GPT-6 Astra OpenAI released GPT-6 Astra on the Responses and Chat Completions APIs for reasoning, coding, computer use, research, and documents. It rejects none reasoning effort, custom temperature, top_p, and logprobs. Tool calling requires the Responses API.
- Sep 2, 2026 Google releases Gemini 3.8 Flash gemini-3.8-flash is generally available. Google describes it as its most intelligent Flash model, aimed at long-horizon software engineering, agents, and enterprise workflows.
- Sep 2, 2026 Meta releases Muse Spark 1.3 On September 2, 2026, Meta released Muse Spark 1.3, including a max-reasoning option, on Muse Code and Meta Model API. It targets stronger agentic and coding work and more reliable long instructions.
- Sep 2, 2026 Alibaba Qwen releases Qwen3.8 Max (2026-09-02) 1M-token context window. On Alibaba Cloud Model Studio: ¥12.00 input, ¥36.00 output per 1M tokens.
- Sep 1, 2026 Anthropic prices Claude Fable 5.1 cache reads Cache-read price for Claude Fable 5.1 and Claude Mythos 5.1 is $0.25 per million tokens, 0.025 times base input, versus 0.1 times on other models. Cache write prices are unchanged.
- Sep 1, 2026 Anthropic releases Claude Fable 5.1 Anthropic launched Claude Fable 5.1 (claude-fable-5-1), successor to Claude Fable 5. Context is 1M tokens, max output 128k, price $10/$50 per MTok, cache reads $0.25 per MTok, on Claude API, Bedrock, AWS, Google Cloud, and Microsoft Foundry.
- Sep 1, 2026 Anthropic releases Claude Mythos 5.1 1M-token context window.
- Sep 1, 2026 Google adds agentic video understanding to Gemini Agentic video understanding is available for Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite on the Interactions and GenerateContent APIs. Google says it can use up to 88% fewer tokens on long-form video than static processing.
- Sep 1, 2026 Meta releases Muse Voice Transcribe On September 1, 2026, Meta released Muse Voice Transcribe, a real-time audio model with streaming speech recognition, diarization for more than 20 speakers, endpointing, and multilingual code-switching.
- Sep 1, 2026 Alibaba Qwen releases qwen3.7-text-embedding-flash qwen3.7-text-embedding-flash is now available.
- Sep 1, 2026 Alibaba Qwen releases qwen3.7-text-rerank qwen3.7-text-rerank is now available.
- Sep 1, 2026 MongoDB releases rerank-3-lite rerank-3-lite is now available.
- Sep 1, 2026 MongoDB releases rerank-3 rerank-3 is now available.
- Aug 31, 2026 Mistral AI makes OCR 4.1 generally available On August 31, 2026, Mistral AI made OCR 4.1 (mistral-ocr-4-1) generally available.
- Aug 28, 2026 Tencent releases Hy4 Preview 1.02M-token context window. On Tencent Cloud TokenHub: ¥6.00 input, ¥18.00 output per 1M tokens.
- Aug 28, 2026 Alibaba Qwen releases qwen-mt-image-2.0 qwen-mt-image-2.0 is now available.
- Aug 27, 2026 Google releases Gemini Omni Flash Gemini Omni Flash is now available.
- Aug 26, 2026 Google releases Gemini 3.5 Transcribe Live Gemini 3.5 Transcribe Live is now available.
- Aug 26, 2026 Google releases Gemini 3.5 Transcribe Gemini 3.5 Transcribe is now available.
- Aug 26, 2026 Zhipu AI releases GLM-5.3-Flash 1M-token context window. On Z.ai API: $0.15 input, $0.50 output per 1M tokens.
- Aug 26, 2026 OpenAI deprecates whisper-1, gpt-4o-transcribe, gpt-4o-mini-transcribe and gpt-4o-transcribe-diarize whisper-1, gpt-4o-transcribe, gpt-4o-mini-transcribe and gpt-4o-transcribe-diarize are now deprecated and shut down on 2027-02-26.
- Aug 26, 2026 Alibaba Qwen releases Qwen3.8 Flash 1M-token context window. On Alibaba Cloud Model Studio: $0.15 input, $0.47 output per 1M tokens.
- Aug 20, 2026 Alibaba Qwen releases Wan3.0 Video Prime Wan3.0 Video Prime is now available.
- Aug 18, 2026 Zhipu AI releases GLM-5.3 1M-token context window. On Z.ai API: $1.40 input, $4.40 output per 1M tokens.
- Aug 17, 2026 Alibaba Qwen releases qwen3.8-27b 1M-token context window. On Alibaba Cloud Model Studio: $0.50 input, $3.00 output per 1M tokens.
- Aug 13, 2026 Google releases Gemini 3.7 Flash 1.05M-token context window. On Gemini API: $0.75 input, $3.75 output per 1M tokens.
- Aug 12, 2026 xAI releases Grok 4.6 500K-token context window. On xAI API: $2.00 input, $6.00 output per 1M tokens.
Written from verified changes in our dataset or from providers’ official announcements, in our words, and published automatically with source links.