Tag: AI Inference

Fujitsu puts a 2nm Arm chip behind Japan-built AI servers

Fujitsu will start selling its MONAKA processor and Japan-built servers in November, pitching them at buyers who want…

Alibaba previews Qwen4 architecture in a 6B-active open model

An experimental open-weight release sketches the design direction for Alibaba's next flagship.

Ramp’s free model gateway steers AI queries to the cheapest option

Ramp launched Router, a model routing API that is free through 2026 and picks the cheapest LLM for…

Kog bets software unlocks 30x faster inference on stock GPUs

The French startup says its software can push 30x faster decoding out of the GPUs enterprises already own.

d-Matrix snaps up Wallaroo to run AI across mixed silicon

Chipmaker d-Matrix acquires Wallaroo.ai to orchestrate AI inference across GPUs and its Corsair accelerators.

DigitalOcean’s Stealth Edge AI Network Targets Real-Time Inference for Startups

DigitalOcean is building a distributed edge AI network with purpose built data centers and inference technology, targeting startups…