Tag: MoE models

StepFun’s Step 5 Preview stacks 92 layers to stretch agent runs

The Chinese lab squeezed 600B parameters into a sparse model that activates 27B per token and holds context…