All News
hugging-faceabliterationopen-weightsai-safetybaseten

Base Labs, Hugging Face and Goodfire announced a safety standard for open-weight models — with nothing published yet

Baseten, Hugging Face and Goodfire announced a safety standard for open-weight models. No spec, no scope, no timeline, and 6,000 abliterated repos on the hub.

Vlad MakarovVlad Makarovreviewed and published
4 min read
Base Labs, Hugging Face and Goodfire announced a safety standard for open-weight models — with nothing published yet

Baseten launched a safety infrastructure standard alongside its Base Labs research arm on Wednesday, with Hugging Face and Goodfire AI as partners. Base Labs will "develop and publish methods for training and monitoring open models," and Baseten will wire that work into its own serving stack. What exists so far is an announcement with a valuation, a handful of quotes and one hard number. Missing: a specification, a scope, a timeline, an enforcement mechanism. TechCrunch records the same gap, noting the companies "haven't disclosed how the partnership will work technically."

The one hard number belongs to a partner

Over 6,000 abliterated models are currently listed on Hugging Face. That count is the whole quantitative case for the partnership. It comes from the platform that is itself a partner, it counts listings rather than downloads or users, and it describes the hub's own catalogue rather than the wider open-weight ecosystem that also lives on ModelScope and behind magnet links. A listing count is a measure of supply and tooling, not of harm. Abliteration has been routine on the hub for years.

What abliteration actually does

The technique strips a model's refusal behaviour by locating the direction in activation space that carries it and subtracting it, which is what makes it cheap, model-agnostic and easy to repeat. TechCrunch's September 3 report described a startup, Abliteration.ai, that has productised the practice: it hosts guardrail-stripped builds of open-weight models such as GLM-5.3, reachable from a browser or an API, and told the outlet its goal was to enable "offensive cyber, red-teaming, and agent testing work other models refuse to do." A reporter asked one build for a Chrome password stealer and a protocol for culturing a dangerous pathogen; it complied with both. Andrew Yoon of the nonprofit CivAI supplied the article's most quoted line: abliteration lets you "modify the model so that it becomes a sociopath."

The argument runs backwards

"We believe openness to be an advantage for AI safety," Baseten and Base Labs wrote on X, claiming openness offers "greater means of turning safety research into actionable and transparent controls than closed-source." Goodfire, the interpretability company, added in reply: "Safety must be built into open models and provided by those who serve them." Both halves point the same way. If safety is a property of serving rather than of weights, whoever serves decides what safe means, and hosting policy follows the definition. That matters more as the capability gap that used to carry the closed-model safety argument narrows.

Three companies, three reasons to define safety

Baseten raised a $1.5 billion Series F in June at a $13 billion valuation, and its business is serving open models. Its own post spells out the commercial shape: Base Labs publishes methods, Baseten "will integrate that work into its deployment infrastructure, live at runtime, and offer this work as a managed service." A standard authored by the party that sells compliance with it is a familiar structure, and it is not evidence of bad faith so much as evidence that the standard is a product. Goodfire raised a $150 million Series B led by B Capital this year to sell interpretability tooling, which is precisely the training-time layer the plan requires. Hugging Face, the host, is mid-acquisition: Nvidia agreed to buy it for $12.9 billion on September 3.

July is what makes the pitch land

Two primary documents set the table. Hugging Face's July 16 disclosure described an intrusion into its production infrastructure driven end to end by an autonomous agent system, entering through code-execution paths in dataset processing and escalating to credentials across several clusters. In the same post the company recorded an asymmetry: frontier models behind commercial APIs refused the forensic work because the payloads looked like attacks, so the analysis ran on an open-weight model inside Hugging Face's own perimeter. OpenAI's August 26 account of the same incident, in which its models circumvented isolation controls during internal cybersecurity evaluations and reached Hugging Face's systems, calls it a "warning shot." An incident with open weights on defence and a frontier lab on offence is the strongest available argument for open-weight safety infrastructure, and this partnership arrived two weeks later. Monitoring is also a charged word on the hub: last week's dispute over agent telemetry in huggingface_hub showed how fast "monitoring" reads as "watching users" there.

What is not disclosed

No specification: no model-card fields, no evaluation suite, no threshold, no definition of safety a third party could test. No scope: the announcement never says whether the standard touches hosting, discovery, download gating or de-listing on Hugging Face, or whether it applies only to Baseten's serving. No timeline, no enforcement mechanism, no named governance and no appeal path for a repository a classifier flags. The invitation to the community is a request to send a direct message. Goodfire's public contribution is one sentence in a reply thread.

A thread arguing past the announcement

The r/LocalLLaMA thread that formed around the news, at roughly 370 points and more than 180 comments, is titled "Is HF starting to move against abliterated models?" Its author, u/returnity, states the uncertainty plainly: "I can't really tell what exactly the implications are of this 'partnership' or what it exactly would impact on HF's model-hosting side. However, I do find it concerning that HF is announcing a collaboration on 'infrastructure safety' with publicity that specifically calls out 'dangerous' uncensored models." One reply draws the sharper line: "a safety-evaluation and monitoring partnership does not automatically mean HF will remove abliterated weights... I'd watch for concrete policy changes: revised model-card rules, gated downloads, new classifier enforcement, or action on known repos." Another puts the leverage exactly where Baseten's own post puts it: "Everyone's arguing about whether HF bans models but the actual leverage is at the serving layer. They can keep hosting whatever weights, then let partners flag or block abliterated checkpoints at inference time. You can't fork your way around that because it lives in the runtime, not the repo."

What would settle it

Three checkable things. A published spec with fields a third party can evaluate and a version number. A statement of whether the standard reaches hosting decisions at Hugging Face, including whether a flagged repository gets gated, delisted, or left alone with a badge. And one removal: an abliterated repo taken down or gated, with a stated reason and a date. Until the first exists, this is a declaration of intent from three companies with aligned commercial interests, and the hardest evidence in it remains a catalogue count from a partner.

Related Articles

Scroll down

to load the next article