Quick answer:
Meta launched Muse, a personal AI agent, in the US on 8 September 2026. It runs on Muse Spark 1.3 inside a dedicated, isolated Secure VM per user, with a separate Sentinel agent gating everything that leaves that VM. It can send emails, book travel, fill out forms, negotiate bills and make purchases with single-use virtual cards via Stripe Link. Pricing is free for light use, $20/month (Power) or $100/month (Maximum) for heavier use. The technical safeguards are genuinely more detailed than most agent launches — but Muse also asks for access to email, calendars, finances, health and smart-home data from a company with a well-documented history of privacy failures, which is why independent coverage has treated the "secure by design" pitch with real scepticism rather than accepting it outright.
Muse is Meta's answer to a question every major AI lab is now racing to solve: can an agent be trusted with your inbox, your calendar and your card details, not just your chat history? Meta's answer is a real architectural attempt — a locked-down virtual machine per user and a second AI whose entire job is saying no on your behalf.
This review covers what Muse can actually do, how the Secure VM and Sentinel system is designed to work, real published pricing, and why Meta's own history makes "trust us" a harder sell here than it would be for almost any other company shipping the same product.
A first look at Muse's background task execution and the Secure VM / Sentinel security pitch, days after the US launch.
Executive Summary
Muse is Meta's first consumer-facing autonomous agent: not a chatbot that answers questions, but a system designed to take real actions — booking a flight, filling out a form, buying a gift — on a user's behalf, running continuously in the background rather than only responding turn-by-turn. Meta's pitch rests on two claims: that Muse Spark 1.3 is capable enough to handle genuinely open-ended, multi-step goals, and that the Secure VM plus Sentinel architecture makes doing so safe even when Muse needs access to sensitive accounts.
- Best fit: US-based users comfortable connecting email, calendar, payment and shopping accounts to an always-on Meta agent in exchange for real task automation.
- Headline design choice: a per-user isolated Secure VM plus a second, non-overridable Sentinel agent that gates every outbound action — a more detailed security architecture than most agent launches publish.
- Honest caveat: Meta itself acknowledges it retains the technical ability to access a user's Secure VM even though its policy says it won't, and the whole system is self-described, not independently audited.
- Main risk: this is Meta, asking for access to email, finances and health data, less than two years after an $18 billion multistate child-safety settlement and with a documented history of privacy commitments it later broke.
What Muse Actually Does
Meta's own framing for Muse is a personal agent that can take on both small errands and "audacious goals": send an email, book travel, reduce a bill through negotiation, fill out a form, turn a recipe video into a shopping list, send invitations, or make a purchase. Once a user states a broader goal, Muse is designed to build a plan, coordinate the person's time and resources, and then advance the work on its own between check-ins, rather than waiting for a fresh prompt at every step.
Underneath that, Muse connects to categories including email, calendars, payments, health and fitness apps, smart-home systems, dining and shopping platforms, music services and event management tools. Meta offers three ways to connect a service: a built-in connector for major platforms, a custom API connection using a user's own credentials where a public API exists, and a browser-based fallback — Muse can open a browser and interact with a site directly — for services with no API at all. Access is granted per app, not as a single all-or-nothing switch.
Muse Spark: The Model Underneath
Muse runs on Muse Spark 1.3, released 2 September 2026, a week before the agent itself launched. It carries a 1 million-token context window and, per Meta's own figures, uses roughly 20% fewer tool calls and 25% lower token usage on agent tasks than the prior Spark version — a meaningful efficiency gain for a model meant to run continuously in the background rather than answer single prompts. Meta has held back Spark 1.3's top reasoning mode pending further safety testing, and independent testing has found its audio input handling has regressed slightly relative to earlier versions.

Holding back the top reasoning tier at launch is a notable, and arguably prudent, choice for an agent explicitly designed to act autonomously with real financial and account access — it suggests Meta is deliberately trading some raw capability for a longer safety-testing runway on the mode most likely to be used for the most consequential, multi-step tasks.
Secure VM & Sentinel: How It Works
Each Muse user gets a dedicated Muse Secure VM: an isolated cloud computer, complete with its own browser, that runs the agent and stores connected-app credentials in a form Muse's own model cannot read directly. Meta's VP of Superintelligence Labs, Tarek Sheasha, has described the design as running in "its own isolated cell" — language that positions the harness executing real-world actions as deliberately separated from the model doing the reasoning.
The second half of the architecture is Sentinel, a distinct monitoring agent that Meta says is isolated at the system level from Muse itself, meaning Muse cannot override or negotiate with it. Every piece of data or action attempting to leave the Secure VM passes through Sentinel first: it checks the request against existing, user-approved policies, and if nothing matches, it stops and asks the user directly rather than defaulting to allow or deny. In practice, this is the mechanism meant to prevent a scenario where an agent given broad account access quietly does something the user never explicitly authorised.
Meta's stated policy is that staff cannot look inside a user's Secure VM — but the company has been upfront that it retains the technical capability to do so, a distinction worth sitting with: the protection here is a policy commitment layered on top of the technical design, not a guarantee that access is architecturally impossible. Meta has also said Muse data isn't shared with its advertising systems, that users can opt out of having their data used for model training, and that a "forget" feature lets a user remove specific things Muse has learned. A further Muse Confidential VM tier, planned later in 2026, would encrypt the entire VM with a key only the user holds — explicitly closing the "Meta could technically access this" gap that exists in the current design.
Payments: Virtual Cards, Not Real Ones
Purchases go through Link by Stripe, which issues a single-use virtual card number for each transaction rather than exposing a user's actual card details to a merchant or to Muse's own model. Meta says this means Muse never sees real card numbers, and transactions made this way inherit Link's own purchase protections, including cover for damaged or lost items, price-drop matching, and support for returns. Meta has said Shop Pay and 1Password support are coming soon, which would extend the same one-time-token approach to a wider set of checkout flows.
Pricing
| Tier | Price | What you get |
|---|---|---|
| Free | £0 | Standard interaction and light task usage; card required to activate; usage meter with low-allowance warnings |
| Power | $20/month | Higher usage limits for everyday task hand-off |
| Maximum | $100/month | Highest published usage limits |
Meta has not published exact usage quotas for any tier, and both paid plans are described as scaling with actual usage rather than offering a fixed task count. Meta expects most users to remain on the free tier, which is itself notable: even the no-cost plan requires a payment card on file to activate, before a user has made a single agent-driven purchase.
The Trust Problem: Meta's Track Record
Every technical safeguard described above is real, in the sense that Meta has published it as the design. Whether it's sufficient is a fair question specifically because of who is asking users to trust it. Independent coverage of the Muse launch, including a widely read TechCrunch analysis, laid out Meta's relevant privacy history in detail: a 2011 FTC settlement over deceiving users about how private information was shared, a 2019 FTC penalty of $5 billion — a US record at the time — across eight separate privacy violations, the 2019 discovery that user passwords had been stored in plainly readable form internally, and a 2023 FTC finding that Meta had violated the terms of its own earlier privacy order. The Cambridge Analytica scandal, involving unauthorised harvesting of data from tens of millions of users, remains the reference point most people associate with the company's data practices.
More recent child-safety failures compound the picture: multiple congressional testimonies and lawsuits over youth protection, an $18 billion multistate settlement in August 2026, and a $942 million ruling against Meta in New Mexico. None of this proves Muse's Secure VM or Sentinel will fail in the way past products did — the specific failures above were mostly about data-sharing practices and policy, not this kind of sandboxed-agent architecture, which didn't previously exist. But it is the direct reason serious coverage of Muse has treated "secure by design" as a claim to test rather than a claim to accept, and it is the single biggest factor separating how cautious a user should be here versus adopting an equivalent agent from a company without that specific history.
Early Reception
Coverage in the days immediately after launch has split fairly cleanly along two lines: outlets focused on the product mechanics (booking flows, the Secure VM/Sentinel design, pricing) have generally described the technical approach as more thorough than a typical agent launch, while outlets and creators focused on trust have been openly sceptical, explicitly invoking Meta's history rather than evaluating Muse purely on its own technical merits. Both reactions are reasonable reads of the same launch: the architecture is genuinely more detailed than most competitors have published, and the scepticism about whether that architecture will be enough is equally well-founded given the company behind it.
As with any newly launched agent handling live financial transactions, independent, adversarial security testing — not just Meta's own description of the design — is the evidence that will actually settle whether Sentinel holds up under real attempts to route around it. That testing had not meaningfully surfaced in public within the first few days of the US launch.
Limitations
- Self-described security: the Secure VM and Sentinel design comes entirely from Meta's own announcement; no independent security audit had been published at launch.
- Meta retains technical access: the company says policy prevents staff from viewing inside a user's Secure VM, but has acknowledged it could technically do so until the encrypted Confidential VM tier arrives later in 2026.
- US-only at launch: Muse is limited to US users aged 18 and over, with no international rollout date announced.
- Unpublished usage quotas: neither paid tier has a stated task or token limit, making it hard to judge value before signing up.
- Top reasoning mode withheld: Muse Spark 1.3's most capable mode is held back pending safety testing, so the model in production today is not Meta's full-capability version.
- Company history: Meta's documented privacy and child-safety settlements are a legitimate, ongoing reason for caution independent of Muse's specific technical design.
How It Compares
Muse's closest conceptual rivals are Google's Gemini Spark and the wave of task-executing agents shipping across labs this year, though most of those remain more narrowly scoped to coding or research tasks than to open-ended personal errands with live payment access. Hermes-style third-party agent platforms cover some of the same automation ground but without a comparably detailed, lab-published isolation architecture, and without Meta's specific reach into WhatsApp and existing social/messaging habits.
The more useful comparison than any specific rival model is a comparison of trust models: an agent from a smaller or newer company asks a user to trust an unproven track record, while Muse asks a user to trust a well-documented one — just not necessarily a favourable one. Deciding between them is less about raw capability and more about which kind of uncertainty a given user is more comfortable accepting.
Who Should Use It
Use it if you're a US-based user already comfortable with Meta's ecosystem (WhatsApp, Meta AI) who wants real task automation and is willing to grant account access gradually, app by app, while watching how Sentinel's permission prompts behave in practice. Start cautiously by connecting lower-stakes accounts first — calendar or a shopping account rather than primary email or banking — until Sentinel's real-world behaviour, not just its description, has been tested by wider adoption. Wait or look elsewhere if you're outside the US, need this for genuinely sensitive financial or health-adjacent workflows before independent security audits appear, or are unwilling to extend trust to Meta specifically given its regulatory history.
The Bottom Line
Muse is a genuinely ambitious piece of engineering: a dedicated, isolated VM per user and a second, non-overridable agent whose entire purpose is gatekeeping outbound data and actions is a more serious security architecture than most personal-agent launches have shipped with. Muse Spark 1.3 underneath it is a real, efficiency-focused model rather than a rebadged chatbot, and the virtual-card payment design closes an obvious risk around exposing real financial details to an autonomous system.
None of that changes who is asking for this level of access. Meta's own regulatory history — a record FTC penalty, a finding that it broke an earlier privacy order, and recent nine- and ten-figure child-safety settlements — means the burden of proof here is higher than it would be for almost any competitor shipping the identical product. The technology is worth taking seriously; the "secure by design" framing is worth testing rather than trusting, and the next few months of independent security research, not this launch announcement, will be what actually answers whether Sentinel holds up.
Last updated: 11 September 2026. Sources: Meta's official Muse and Muse Spark newsroom announcements (about.fb.com), TechCrunch, SiliconANGLE, DataCamp, TechRepublic and Bloomberg.
Get the free guide: Claude vs ChatGPT, Gemini & Grok
A 20-page playbook covering everything you need to choose and use the big four AI models in 2026, full cost and feature comparisons, what each is best (and worst) at, and how-tos for images, vectors, building a website, Claude Code and more.








