OpenAI's Ultrafast mode pushes GPT-5.6 Sol to 750 tokens per second

OpenAI's Ultrafast mode pushes GPT-5.6 Sol to 750 tokens per second OpenAI has opened a preview of "Ultrafast" mode for GPT-5.6 Sol, its flagship reasoning model. The mode streams up to 750 output tokens per second, which OpenAI frames as roughly 14 times the speed of the standard API, by running inference on Cerebras hardware. Cerebras and OpenAI signed a ten billion dollar partnership earlier this year, and Ultrafast is the first major use of that capacity. Access is limited for now to

1 min
OpenAI's Ultrafast mode pushes GPT-5.6 Sol to 750 tokens per second

OpenAI's Ultrafast mode pushes GPT-5.6 Sol to 750 tokens per second

OpenAI has opened a preview of "Ultrafast" mode for GPT-5.6 Sol, its flagship reasoning model. The mode streams up to 750 output tokens per second, which OpenAI frames as roughly 14 times the speed of the standard API, by running inference on Cerebras hardware.

CloudSEK supply-chain breach visualization

Cerebras and OpenAI signed a ten billion dollar partnership earlier this year, and Ultrafast is the first major use of that capacity. Access is limited for now to the OpenAI API and to a set of selected customers. OpenAI says it will widen availability as capacity grows and is collecting sign-ups through a public form.

The pitch is speed without downsizing. OpenAI says Ultrafast keeps the full capabilities of a large reasoning model while matching the responsiveness of a smaller one, which it calls more useful work per second. The company points to live incident response, where logs, code changes, and postmortems could be analyzed while an outage is still unfolding, and to finance, support, and research workflows that currently run as overnight batch jobs.

Ultrafast is also a pricing move. OpenAI already sells a Fast Mode API tier that promises about 2.5x speed for GPT-5.6 Sol at roughly twice the price. Ultrafast adds a third, faster, and likely more expensive tier, turning inference latency into a metered product the way cloud providers meter compute performance.

Sources:

Written by

More to read

  • Activation Checkpointing in Large Language Models: How Selective Recomputation Eliminates Memory Bottlenecks

    Large language model pre-training and fine-tuning are fundamentally constrained by GPU memory (VRAM). While distributed techniques such as Fully Sharded Data Parallel (FSDP), ZeRO, and Tensor Parallelism successfully shard model parameters, optimizer states, and gradients across hundreds or thousands of GPUs, activation memory presents a distinct scaling bottleneck. During the forward pass of a transformer model, intermediate tensor outputs must be preserved in GPU memory so that backpropagatio

    1 min
  • Anthropic Demonstrates Autonomous De Novo Protein Design and Chemical Analysis with Claude

    Anthropic Demonstrates Autonomous De Novo Protein Design and Chemical Analysis with Claude Anthropic has published experimental results demonstrating Claude's ability to autonomously design de novo protein binders with physical wet-lab validation and automate complex analytical chemistry workflows. The findings show frontier LLMs acting as autonomous agents across computational biology and molecular characterization pipelines. In the primary experiment, Anthropic evaluated Claude Mythos Previe

    1 min
  • Cerebras Unveils CS-4 Rack-Scale System Powered by Three WSE-3 Turbo Chips and Nexus Architecture

    Cerebras Unveils CS-4 Rack-Scale System Powered by Three WSE-3 Turbo Chips and Nexus Architecture Cerebras Systems has announced the CS-4, a rack-scale AI accelerator system designed around three of its next-generation Wafer Scale Engine 3 Turbo (WSE-3 Turbo) chips and a modular hardware architecture dubbed Nexus. Cerebras confirmed that initial customer shipments for the CS-4 are scheduled to begin in the current quarter. The new system marks a structural shift from Cerebras's single-wafer CS

    1 min