#Meta just published how it secures #Muse, its personal AI agent, and the architecture is worth studying for anyone building agentic systems. The core idea: assume the agent will be attacked, and design the system so a compromised agent causes minimal damage. • Each user gets an isolated cloud VM (the Muse Secure VM) running the agent, a browser, code execution, subagents, and scheduled jobs. • Real credentials never enter the VM. The agent works with surrogate tokens, and a separate component called Sentinel swaps in the real credential only at the network boundary.
Show more of this post
• Sentinel is the only thing that can talk to connectors or the internet. The agent proposes, Sentinel allows, denies, or asks the user. The agent cannot override it. • Tainted egress: once a process touches private data, it loses automatic network permission. Any later outbound action needs human approval. This is Meta's answer to the "lethal trifecta" of private data, untrusted input, and a way to send data out. • The model is trained with prompt injection in mind, but Meta's position is that model-level defenses are never enough, so the real controls live at the OS level. Notably, Meta also opened a public bug bounty for Muse: up to $300,000 per report, including up to $130,000 for a successful prompt injection affecting a single user. The broader lesson: permission needs an owner outside the agent that wants to act. #AIAgents #AISafety #Cybersecurity #PromptInjection #Meta research.meta.ai/blog/security…