Tag: Inference

Nvidia’s Vera Rubin racks triple throughput in first MLPerf run

Nvidia's first MLPerf Inference submission for its Vera Rubin NVL72 rack reports up to 3.7x the throughput of…

ChatGPT co-inventor leaves stealth with Jev for typed decisions

TypeSafe, founded by ChatGPT co-inventor Diogo Almeida, has launched Jev, a model that returns typed probabilistic choices instead…

Chipmaker d-Matrix books a rack slot in NVIDIA’s AI factory plan

Raptor inference XPUs will hook into MGX racks through NVLink Fusion, splitting inference work between GPUs and custom…

DeepSeek shrinks its KV cache to fit million-token contexts

DeepSeek V4.1 Flash targets the memory bill behind long-running agents, cutting cache footprint to 890 bytes per token.

OpenAI custom silicon outruns Nvidia on first inference tests

OpenAI's first benchmarks for its in-house Jalapeno chip show it beating an Nvidia Blackwell system on power-scaled inference.

NVIDIA says Vera Rubin racks handle agents at 30x the efficiency

NVIDIA claims its Vera Rubin racks deliver up to 30x more throughput per megawatt for agents.

Berkeley researchers run a 753B model on a single workstation GPU

Berkeley and UT Austin researchers serve the 753B GLM-5.2 on a single workstation GPU with a new open-source…

A small drafter model makes Liquid AI’s LFM2.5 run three times faster

Speculative-decoding checkpoints from Liquid AI boost LFM2.5 throughput by up to 3.18x without changing outputs.

NVIDIA’s TensorRT Model Connect skips ONNX for native C++

Two commands take a Hugging Face checkpoint to PyTorch-free inference inside a versioned bundle.

Jane Street leads Etched’s $700M round after testing its silicon

Jane Street led Etched's $700M Series D after testing its inference chips, pushing the startup's valuation to $21B.

Groq banks $350M for a second act running Nvidia clouds

Groq raised $350M at a $3.5B valuation as it rebuilds itself as an Nvidia-powered AI cloud.

Cerebras powers OpenAI’s Ultrafast tier at 750 tokens a second

OpenAI previews a service tier that runs GPT-5.6 Sol up to 14 times faster for select API customers.