SAN FRANCISCO — OpenAI introduced a preview of Ultrafast mode for its GPT-5.6 Sol model, an API tier built for enterprise customers requiring high-speed, high-volume processing. The service runs GPT-5.6 Sol up to 14 times faster than standard speeds, generating up to 750 output tokens per second.

Cerebras powers the accelerated processing, a specific infrastructure bet OpenAI is making to hit those performance numbers. The initial launch targets a select group of customers through the OpenAI API.

GPT-5.6 Sol is OpenAI's latest model, with expanded capabilities in coding, scientific research and cybersecurity, and incorporates the company's most advanced safety stack.

The model is rolling out globally, with full availability expected within 24 hours. ChatGPT users on Plus, Pro, Business and Enterprise plans can access GPT-5.6 Sol through medium and higher effort settings.

OpenAI plans to expand Ultrafast access as capacity grows. The tiered structure lets the company charge a premium for performance—a direct monetization of latency for enterprises running real-time customer support, rapid data analysis or automated content generation at scale.

Alongside the Ultrafast launch, OpenAI released o3-mini, a cost-efficient reasoning model optimized for coding, mathematics and scientific applications. On the API, o3-mini supports Structured Outputs, function calling, developer messages and streaming, with three adjustable reasoning effort levels—low, medium and high—letting users trade speed for depth.

The dual strategy—premium throughput at the high end, cost-optimized reasoning at the low end—gives OpenAI coverage across the enterprise budget spectrum as it competes against other frontier AI developers.