The Current

OpenAI Previews 'Ultrafast' Mode Running GPT-5.6 Sol Up to 14x Faster

The API-first service tier, powered by Cerebras, is available in a limited preview to select customers.

useful models · for technical · August 14, 2026

OpenAI announced on August 13, 2026 an early look at Ultrafast, a new service tier that runs its GPT-5.6 Sol model up to 14 times faster than Standard processing, launching first in the OpenAI API. According to the company's blog post, Ultrafast generates up to 750 output tokens per second and is powered by a partnership with chipmaker Cerebras. Tokens represent the distinct pieces of text a model generates when interacting with a user. OpenAI said that until now, getting real-time speed 'typically meant choosing a smaller or more specialized model,' and framed Ultrafast as offering 'more useful work per second' without sacrificing intelligence. The company is testing the mode with an initial group of companies across coding, commerce, financial research, and support, and internally for uses such as incident response and research. OpenAI included statements from customers at Jane Street, Podium, Basis, and Rogo describing speed improvements in their workflows. TechCrunch reported that competitors including Anthropic have launched accelerated versions of their models, noting Claude's fast mode does not match the speed OpenAI describes. The preview is currently available only to a small group of customers, with OpenAI saying it will expand access 'as capacity grows.'

  • Ultrafast runs GPT-5.6 Sol up to 14x faster than Standard, generating up to 750 output tokens per second
  • The mode is powered by OpenAI's partnership with chipmaker Cerebras and launches first in the OpenAI API
  • Announced August 13, 2026, currently a limited preview open only to a select group of customers
  • Named test areas include coding, commerce, financial research, support, and incident response

What it means for you

OpenAI now offers a way to run its most capable model much faster, but only through its API and only to a handful of hand-picked companies right now. For most people, this changes nothing today — you can't sign up and use it, and it won't show up in ChatGPT. It matters mainly if you build software that needs an AI to respond in real time during a live interaction, like a phone call or an interactive tool.

Try this

If you run a product where AI response speed genuinely limits the experience — say a voice assistant or live support tool — sign up for OpenAI's Ultrafast notification list so you're in line when access widens.

Who should care

Developers and technical teams building real-time, interactive AI products where a delay of a few seconds hurts the user experience.

Skip this if

You use ChatGPT through the normal app or web interface, or your work doesn't depend on instant AI responses — this is a preview you can't access yet and don't need.

Sources: OpenAI, TechCrunch AIread the original

← All stories