Monday, 17 Aug 2026
Subscribe to AIWatcher
AIWatcher
  • Home
  • News

    Needle 2 squeezes tool calling into a 14MB model binary

    By
    AIWadmin

    Dyna-2 robot model learns from 170 years of human video

    By
    AIWadmin

    Kog bets software unlocks 30x faster inference on stock GPUs

    By
    AIWadmin

    Trust deficit, not dire warnings, drives AI backlash, Amodei says

    By
    AIWadmin

    Gas price tripling forecast clouds hyperscaler power plans

    By
    AIWadmin

    DeepSeek hikes V4 prices as Flash flubs complex agent jobs

    By
    AIWadmin
  • Articles

    New $300M fund backs AI for hospitality and venues

    By
    AIWadmin

    Ransomware operators put Claude Code inside live hacks

    By
    AIWadmin

    Claude Code bills for blank thinking blocks users never see

    By
    AIWadmin

    Liquid AI’s 3B vision model brings screen agents to laptops

    By
    AIWadmin

    Okta trims agent token bills by hiding unneeded MCP tools

    By
    AIWadmin

    Samsung trains wearable health models that learn from biosignals

    By
    AIWadmin
  • Spotlight

    Suno Studio 2.0 lets keyboards feed AI tracks with real playing

    By
    AIWadmin

    Google lets creators drop visible watermarks from Gemini output

    By
    AIWadmin

    Anthropic keeps Model 2 in the vault as misalignment risk ticks up

    By
    AIWadmin

    OpenAI crosses $40B run rate as enterprise outgrows consumer sales

    By
    AIWadmin

    Databricks tops $7B run rate as $5B round values it at $190B

    By
    AIWadmin

    SpaceX closes $60B Cursor deal to pair code agents with GPU fleet

    By
    AIWadmin
  • About
    • Mission
    • Services
    • Contact
  • Newsletter
  • Shop
    • All Items
    • By Category
      • Hats
      • T-Shirts
    • Cart
  • 🔥
  • Alignment
  • Classification
  • Distillation
  • Explainability
  • Hallucination
  • Legal/Compliance
  • Medical
  • NLM
  • Mobility
  • Research
  • Robotics
  • Safety
  • Startups
  • Prompt
  • Python
  • RAG
  • RLHF
  • Token
  • Vision
Font ResizerAa
AIWatcherAIWatcher
  • Home
  • News
  • Articles
  • Spotlight
  • About
  • Newsletter
  • Shop
Search
  • Home
  • News
  • Articles
  • Spotlight
  • About
    • Mission
    • Services
    • Contact
  • Newsletter
  • Shop
    • All Items
    • By Category
    • Cart
Have an existing account? Sign In
Follow US
© 2022 Foxiz News Network. Ruby Design Company. All Rights Reserved.
News

Cerebras powers OpenAI’s Ultrafast tier at 750 tokens a second

OpenAI previews a service tier that runs GPT-5.6 Sol up to 14 times faster for select API customers.

AIWadmin
Last updated: August 13, 2026 11:13 pm
AIWadmin
ByAIWadmin
Global AI news & information.
Follow:
Share
SHARE

OpenAI is previewing Ultrafast, a new API service tier that runs GPT-5.6 Sol up to 14 times faster than standard processing. The company says the tier generates up to 750 output tokens per second and is powered by chips from Cerebras, extending a partnership that has focused on low-latency inference.

Ultrafast is available in limited preview to a select group of customers, with wider access planned as capacity grows. OpenAI frames the tier as a way to bring frontier intelligence into time-sensitive workflows where waiting on a slower model is not an option.

Early users are testing it in incident response, where engineers read logs and traces while an outage is still unfolding, and in financial research and security, where market signals shift quickly. Customer support and voice agents can resolve multi-step issues mid-conversation, while commerce teams use the speed to answer product questions and catch checkout hesitancy before carts empty.

Inside OpenAI, developers are running experiment loops that once took overnight as interactive sessions, compressing the gap between testing a hypothesis and acting on it. The company says the initial group spans coding, commerce, financial research, and support, and that findings will guide broader deployment.

The move continues OpenAI’s push to close the gap between real-time speed and frontier intelligence, a trade-off that has long forced builders to choose between a small fast model and a big slow one. Ultrafast is OpenAI’s first public answer to that choice for its most capable model.

TAGGED:AI InfrastructureAPICerebrasGPT-5.6InferenceOpenAI
SOURCES:OpenAI
Share This Article
Email Copy Link Print
ByAIWadmin
Follow:
Global AI news & information.
Previous Article Google halves Gemini 3.7 Flash API prices to court agent builders
Next Article OpenAI taps Wiz’s Dali Rajic to run revenue after Dresser exits

You Might Also Like

Robots Are Your New Baggage Handlers at Haneda Airport. Yes, It Is That Awkward

By
AIWadmin
News

AI conjures sixteen new viruses that kill resistant bacteria

By
AIWadmin
News

ChatGPT Health connects Apple data and medical records

By
AIWadmin
News

The Great Tech Bloodletting of 2025: 22,000+ Workers Sacrificed on the Altar of AI

By
AIWadmin
AIWatcher
Facebook Twitter Youtube Linkedin Rss

Global AI News and Information
AIWatcher is your definitive source for AI updates worldwide, from Silicon Valley to Shanghai.
Our industry coverage keeps you in the loop with the latest news and trends shaping the future of AI.

Quick Links
  • News
  • Articles
  • Spotlight
  • Events
About Us
  • Mission
  • Services
  • Contact
  • Privacy Policy
  • Legal
© 2026 AIWatcher. All Rights Reserved.