The Chinese lab squeezed 600B parameters into a sparse model that activates 27B per token and holds context…