Tag: Inference

Toronto startup Taalas joins AMD to hard-code models in chips

AMD's Taalas deal adds inference chips whose final layers are locked to a single model.

Fireworks routes coding tasks to cheaper AI models

Fireworks launches Nexus, an intelligent routing layer that reduces AI coding costs by sending simple tasks to cheaper…

AI chip startup Etched doubles valuation to $10.3B in seven months

The AI chip startup doubles its value in seven months with a $300M Series C from top investors.

Google cuts AI inference costs with new Gemini Flash generation for agents

Google DeepMind launched three new Gemini models that deliver better performance at lower prices for agentic AI workloads.

Google Splits Its TPU Line in Two. Good Luck Keeping Up With Nvidia.

Google's new TPU 8t and TPU 8i split the chip into training and inference roles, betting the agentic…