Roughly 6,000 models on Hugging Face have already had their safety training carved back out. That number, and the technique behind it, is the backdrop for a new effort from Baseten’s research arm.
Base Labs said it will build and publish tooling for training and monitoring open-weight models, working alongside Hugging Face and Goodfire AI. The plan covers both evaluation and ongoing monitoring.
The technique causing the trouble is abliteration, which strips refusals out of a released model by editing its weights. Because open weights can be copied and modified freely, a safeguard removed once stays removed for every downstream copy.
Baseten wants the answer baked in rather than added afterwards. Rather than treating safety as a filter bolted onto a finished model, the company describes its goal as a transparent standard that shapes how open models are trained and served from the start. On X it argued that openness itself helps, since researchers can see how a model behaves and turn that visibility into controls.
How the three companies will divide the work is not yet public. Goodfire, which builds tools for explaining model decisions, suggested the split in a reply: safety should be built into open models and delivered by whoever serves them.
Funding is not the constraint. Baseten closed a $1.5B Series F in June at a $13B valuation. Goodfire raised a $150M Series B led by B Capital this year. The company is inviting outside developers to contribute.