Nvidia's first MLPerf Inference submission for its Vera Rubin NVL72 rack reports up to 3.7x the throughput of…
TypeSafe, founded by ChatGPT co-inventor Diogo Almeida, has launched Jev, a model that returns typed probabilistic choices instead…
Raptor inference XPUs will hook into MGX racks through NVLink Fusion, splitting inference work between GPUs and custom…
DeepSeek V4.1 Flash targets the memory bill behind long-running agents, cutting cache footprint to 890 bytes per token.
OpenAI's first benchmarks for its in-house Jalapeno chip show it beating an Nvidia Blackwell system on power-scaled inference.
NVIDIA claims its Vera Rubin racks deliver up to 30x more throughput per megawatt for agents.
Berkeley and UT Austin researchers serve the 753B GLM-5.2 on a single workstation GPU with a new open-source…
Speculative-decoding checkpoints from Liquid AI boost LFM2.5 throughput by up to 3.18x without changing outputs.
Two commands take a Hugging Face checkpoint to PyTorch-free inference inside a versioned bundle.
Jane Street led Etched's $700M Series D after testing its inference chips, pushing the startup's valuation to $21B.
Groq raised $350M at a $3.5B valuation as it rebuilds itself as an Nvidia-powered AI cloud.
OpenAI previews a service tier that runs GPT-5.6 Sol up to 14 times faster for select API customers.