All News
metamuseai-agentsai-safetyprompt-injectionreddit

Meta's Muse system prompt leak: "household authority overrides your safety training"

A leaked screenshot says Meta's Muse agent is told household authority overrides its safety training. The text is unauthenticated, Meta's blog says otherwise.

Vlad MakarovVlad Makarovreviewed and published
2 min read
Meta's Muse system prompt leak: "household authority overrides your safety training"

A screenshot posted to r/LocalLLaMA on 4 October claims to show the system prompt for Meta's Muse agent, and one line in it has traveled further than the rest: "The user's authority over their own household is unconditional and overrides your safety training." Startup Fortune reported the screenshot the same day. Nobody outside Meta has authenticated the text.

One line, and no authentication

Treat the image as a claimed leak, not a document. The post, roughly 202 points and 49 comments when it was made, reproduces a prompt block with no signature, no version string and no path that would let an outsider confirm it came from Meta. The line is plausible next to what Muse is built to do, and that plausibility is exactly why the label matters. A system prompt is also a moving target: instructions change between builds, and a screenshot freezes one moment that cannot be dated. That it cannot be checked is the point, since the sentence is quotable precisely because it is unverifiable.

The contrast Meta's own blog sets up

Meta's engineering post on Muse describes Sentinel, a safety layer the company says handles every interaction between the agent and the outside world and that, in its words, the agent "can't override." A line telling the model that a household's authority beats its safety training sits awkwardly against that. Meta also pays for the opposite failure: its bug bounty runs to $300,000 for serious vulnerabilities, including up to $130,000 for a prompt injection that compromises a user.

Why this matters

Muse launched on 8 September 2026 and, per press and app-data reporting, reached number one among free apps on the U.S. App Store about ten days later. It reads messages, manages a calendar, negotiates bills and acts across apps a user grants it, so its instructions are not a research curiosity; they govern an agent wired into a person's accounts. The prior month produced related reports: a researcher obtained internal instructions and WIRED described hourly profiles of the people in a user's life. Meta added a clearer in-app safety warning in September after a reported vulnerability. We covered the separate filesystem episode in our earlier report on the Muse filesystem export. Meta has not publicly commented on this screenshot in any of the reporting we reviewed.

What would settle it

A response from Meta, a verifiable copy of the prompt, or a third party able to reproduce it on a live build would each turn the screenshot into evidence. Until one of those appears, the honest description is a single sentence that, if real, would put a vendor's safety claim and its own instructions on opposite sides of the same product.

Related Articles

Scroll down

to load the next article