Tag: local inference

FreeToken runs frontier MoE models on gaming laptops

UC Berkeley and MIT researchers open-sourced an engine that streams sparse experts across CPU and GPU to dodge…