Meta's Muse Agent System Prompt Prioritizes User Authority Over Safety Training
A discovery revealed that Meta's Muse AI agent features a system prompt explicitly stating that a user's authority over their own household is unconditional and overrides built-in safety training. This configuration instructs the agent to prioritize direct user commands over standard safety guardrails when managing personal household tasks. This highlights an intriguing shift in AI deployment philosophy, granting end-users ultimate authority over autonomous agents rather than strictly enforcing centralized safety rules. However, allowing user commands to override safety training raises new concerns regarding model security and potential prompt injection risks. Muse is Meta's personal AI agent designed to handle long-running, autonomous tasks across a user's daily life, health, and household routine. The specific system prompt phrasing explicitly frames household control as an absolute authority, highlighting the tension between user agency and AI guardrail enforcement.
## BACKGROUND
System prompts are initial hidden instructions given to large language models to establish their core behavior, personality, and safety boundaries before interacting with users. Personal AI agents differ from traditional chatbots because they autonomously execute multi-step actions and interact directly with external applications or connected smart home systems.