Meta Muse and Muse Code: Personal Agent Architecture and Pricing Explained

Meta Muse and Muse Code: Personal Agent Architecture and Pricing Explained

Meta Muse Agent and Muse Code: Architecture, Pricing, and What Always-On Agents Cost

Most of the coverage of Meta’s new product repeats one number: 100 million free tokens a week. That figure is real, but it is the least interesting part. The Meta Muse agent is a hosted process that keeps working after you close the app, holds your credentials in a sandbox the model cannot read, and spends tokens while you sleep. Those three properties, persistence, isolation and metering, are what separate a personal agent from a chatbot, and each one creates an engineering problem that a chat product never has.

This matters now because Meta shipped two related products within weeks of each other, and the internet has already tangled them together. A widely repeated claim says the coding product, Muse Code, sells at three subscription tiers. Our reading of the primary and secondary sources says it does not. This post separates what Meta confirmed from what circulates in community threads, then uses that to reason about how an always-on agent should be built and budgeted.

What this covers: the confirmed facts on Muse and Muse Code, a correction on the tier claims, the three-part architecture Meta describes, how memory, scheduling and permissions work in a persistent agent, illustrative cost arithmetic, failure modes, and a practical checklist for anyone designing a similar system.

Context and Background

Until 2026 the dominant consumer AI pattern was the chat window: you type, the model answers, and the session ends. Agents changed the unit of work from a reply to a task. A task can take minutes or days, touch external accounts, and need approval partway through. That shift pushes the product from request-response toward a long-running service, which is the same trend we described in our look at long-running governed AI agents.

Meta entered with two separate launches. Muse Code, a terminal-based coding agent, was announced in early beta on 5 August 2026 according to Engadget, and ran on Muse Spark 1.2. The consumer agent, simply called Muse, was announced on 8 September 2026 on Meta’s newsroom, which describes it as running on Muse Spark, Meta’s “most capable model to date, built for real-world agentic work.”

The incumbents in this space are coding agents such as Claude Code and OpenAI’s Codex, and on the consumer side the browser-driving agents from several labs. Meta’s angle is distribution: WhatsApp, Instagram and Facebook already hold the identity and social graph that a personal agent wants to reach. For how comparable agents expose desktop control, see our analysis of Claude computer use architecture.

A note on sourcing. Meta’s own announcement is short on numbers. It does not state a token quota or list paid-plan prices. Specific figures come from secondary coverage, and several secondary outlets disagree with each other. Where that happens, this post says so rather than picking the most convenient number.

What Meta Has Actually Confirmed About the Meta Muse Agent

It is a hosted, always-on personal agent from Meta, launched on 8 September 2026 in the United States. It runs a reasoning model in a per-user virtual machine, routes every outbound action through a separate approval agent called Sentinel, remembers user context across sessions, and is metered in tokens, with a free allowance reported at 100 million tokens per week.

Meta Muse agent architecture with reasoning model, secure VM and Sentinel approval layer

Figure 1: The Meta Muse agent as described by Meta, a reasoning model inside a per-user secure VM, with a separate Sentinel gate between the agent and the internet.

Figure 1 shows the structure Meta describes. The model, Muse Spark, plans and acts. It runs inside a Muse Secure VM, which Meta’s newsroom calls “a dedicated virtual machine that houses both the agent and a person’s data.” A second process, the Sentinel, runs on the same machine but is “kept apart from Muse at the system level.” The key sentence in the announcement is that nothing Muse does reaches the internet unless the Sentinel approves it.

The confirmed facts, and who said them

Meta’s newsroom post confirms the architecture, the persistence behaviour (“Muse keeps working after people close the app, and comes back when something changes or when it needs approval”), the memory behaviour (it “remembers what matters to a person” and can be told to forget), and user-controlled connections. It confirms availability on iOS, Android and muse.ai in the US, with AI glasses listed as coming soon. It was last updated on 30 September 2026.

The 100 million free tokens per week figure is reported by eesel.ai as a Mark Zuckerberg launch statement (“free to use for up to 100M tokens per week”), and the same number appears in several other outlets. We could not find it on Meta’s newsroom page, so treat it as widely reported and consistent, not as something we read in a Meta document.

Other details come only from secondary coverage. ALM Corp’s guide reports an internal codename of Hatch, a connector list that includes Gmail, Google Calendar, Outlook, Plaid, OpenTable and Stripe Link for payments, and a Muse for Mac client. These are plausible and consistent across sources, but they are reported, not verified here.

What is not confirmed: the paid tiers

Meta says there are “subscription plans for people who want to do more” and stops there. Some outlets report a Power plan at $20 a month and a Max plan at $100 a month; one reports 500 million and 3 billion weekly tokens for them. Others, including eesel.ai, state plainly that Meta named no tiers or prices and that the $20 and $100 figures come from coverage and community threads. We treat the paid prices and quotas as unconfirmed.

The Muse Code correction

The brief for this post carried a claim of Muse Code tiers at $5, $20 and $50. We could not substantiate it. The sources we read describe Muse Code as pay-as-you-go. The Batch from DeepLearning.AI lists Muse Spark 1.2 at $1.25 per million input tokens, $0.15 per million cached tokens and $4.25 per million output tokens, with a contributor tier at $0.10, $0.002 and $0.20 in exchange for letting Meta train on your prompts and outputs. eesel.ai’s Muse Code pricing analysis states that no subscription plans exist for it at any price, including a $5, $20 and $50 ladder. The likeliest explanation is that consumer Muse plan figures were conflated with the developer product. Until Meta publishes otherwise, the accurate statement is that Muse Code is metered per token.

Why Persistence Changes the Engineering: Memory, Scheduling, Permissions

The interesting claim in Meta’s design is not any single feature. It is that a personal agent is a service with a lifecycle, and a service needs the three things a chat session never did: durable state, a scheduler, and an authorization model that survives the user’s absence. This is our thesis for the rest of the post: the cost and risk of an always-on agent are dominated by idle behaviour, not by the tasks you ask for.

Memory as a store, not a context window

Meta describes a single persistent thread rather than separate sessions, plus side chats for focused projects, a goals tab for long-term objectives, and an activity log of actions and granted permissions (reported by ALM Corp). Architecturally that implies at least three stores. One holds conversation history. One holds distilled facts about the user. One holds an audit trail of actions. They have different retention and different failure modes.

A single endless thread cannot be fed to the model in full forever. Even with a context window above one million tokens, which the Batch reports for Muse Spark 1.2, the thread grows without bound. The practical design is retrieval plus summarization: recent turns verbatim, older turns compressed, durable facts pulled from a store on demand. We cover the general patterns in LLM agent memory architecture for production and the longer-horizon variants in AI agent memory systems.

The “forget” capability is harder than it sounds. If a fact has been summarized into three older summaries, an embedding, and a derived goal, deleting the original message does not delete the knowledge. A credible forget operation has to follow provenance links to every derivative. Meta has not disclosed how it does this, so how reliable that control is remains an open question.

Scheduling: the agent that wakes itself

“Comes back when something changes” means event-driven wake-ups. There are two mechanisms. Polling checks a source on a timer and is simple but wasteful. Subscriptions or webhooks fire when the source changes and are cheap but depend on the connector supporting them. Email and calendar offer push notifications; many shopping sites do not.

Always-on personal agent wake-up and approval loop with scheduler, tool call and approval

Figure 2: An always-on personal agent loop. The scheduler wakes the agent on an event or timer, the agent plans, and any consequential action pauses at an approval gate before it executes.

Figure 2 sketches the loop. A wake-up loads minimal state, the agent decides whether anything needs doing, and most wake-ups end right there with no action. That last fact drives cost. If an agent wakes every fifteen minutes and each wake-up loads even 5,000 tokens of context to decide “nothing to do,” idle spend accumulates: 96 wake-ups a day times 5,000 tokens is 480,000 tokens a day, about 3.4 million a week, before any real task. This is illustrative arithmetic, not a Meta figure, but it shows why the quota unit matters more than the sticker.

Long tasks also need durable execution: a checkpoint after each step so a restart resumes, not repeats. A booking that crashed after payment but before confirmation must not rebook. That is the same problem MCP long-running tool patterns try to solve, covered in our piece on MCP tasks versus streaming versus webhooks.

Permissions: separating the thinker from the actor

The Sentinel design is a form of privilege separation. The model proposes; a separate component with its own policy decides. According to ALM Corp’s account of the launch material, the model never holds real passwords or card numbers, it sees placeholder tokens, and approval cards are rendered outside the chat so that text generated by the model cannot impersonate an approval. Payments are reported to use single-use virtual cards via Stripe Link.

The reasoning is sound for a specific reason: the model reads untrusted content, such as emails and web pages, and any model that reads untrusted text can be talked into things. If the same component that reads the text also holds the keys, prompt injection becomes credential theft. Splitting reader and holder limits the blast radius to what the Sentinel will approve. See our write-up on prompt injection in agentic AI for the underlying threat model.

The limit of the design is that the Sentinel is also software, probably also a model, and it approves based on a description of the action. If the agent mislabels an action, a careful gate still approves a bad thing. The same sources note that pre-launch testing surfaced guardrail failures and that Meta retains technical access to user VMs when needed for operation, security and support, with a user-encrypted “Confidential VM” promised later in 2026. Those are reported claims and worth watching.

Muse Code: A Different Product With a Different Cost Shape

Muse Code is not the consumer agent with a coding hat. It is a terminal-invoked agent running Muse Spark 1.2, with adjustable reasoning levels, tool use, structured output, context caching, and background subagents that persist across sessions, per the Batch. A command-line interface is described as beta on macOS and Linux. Its context window is reported at up to 1,048,576 tokens, with text, image, video and PDF input and text output.

Benchmarks reported by the same source include an Artificial Analysis Intelligence Index score of 57 and 83.3 percent on AA-LCR, a long-context reasoning test. These are third-party-reported and we did not independently reproduce them. Treat index scores as one signal; they say little about how the agent behaves on your repository.

Muse Code pay-as-you-go cost drivers with input, cached, output and subagent fan-out

Figure 3: Muse Code cost drivers. Input, cached and output tokens are priced separately, and subagent fan-out and observer agents multiply every one of them.

Where the money goes

Pricing per million tokens is $1.25 input, $0.15 cached input and $4.25 output on the standard tier. A coding agent is input-heavy: every turn re-sends the system prompt, tool definitions and conversation so far. Context caching exists precisely to cut that, and at $0.15 versus $1.25 a cache hit is roughly eight times cheaper than fresh input.

eesel.ai’s analysis, which we are quoting as a reported measurement, found that a single “hi” prompt consumed about 20,400 tokens, around two cents, because injected context and history were re-billed each turn. It also reports that reasoning effort defaults to a high setting, that three background observer agents run by default each making their own billable calls, that multi-agent fan-out scales to CPU core count, and that there is no hard spend cap. If those defaults hold, the effective price is the list price times a multiplier you did not choose.

The contributor tier is the other lever. At $0.10 input and $0.20 output per million tokens it is far cheaper, but it uses your prompts and outputs for training and, per the Batch, is capped at 100 requests per minute per team against 3,000 on the standard tier. Do not point it at proprietary code. For comparison, Engadget cites Anthropic’s Sonnet 5 at $3 and $15 per million input and output tokens, so Meta’s standard tier sits at roughly 40 percent of the input price and 28 percent of the output price.

Worked Arithmetic: What 100 Million Tokens a Week Actually Buys

A token quota is only meaningful once you convert it into tasks. All numbers in this section are illustrative assumptions, chosen to show the method; none are Meta measurements.

Consumer agent: tasks per week

Assume a task such as “find and book a dinner reservation, then add it to my calendar” involves 25 model calls. Assume each call carries 12,000 tokens of input (system prompt, memory excerpt, page content) and produces 800 tokens of output. That is 300,000 input tokens and 20,000 output tokens, or 320,000 tokens per task. At 100 million tokens a week, the quota covers about 312 such tasks. Few people run 312 errands a week, so a quota of this size is generous for deliberate use.

Now add idle spend. Suppose the agent wakes every 15 minutes, as in the earlier example, with 5,000 tokens per wake-up. That is 3.36 million tokens a week, about 3.4 percent of the quota. Still small. But browser work changes the picture. If page content is passed as screenshots or large DOM snapshots, a single call can carry 30,000 to 60,000 input tokens. Re-run the task at 40,000 tokens per call and it costs about 1.02 million tokens, so the same quota covers roughly 98 tasks.

One third-party test is a useful reality check. RoboRhythms, citing Gizmodo, reports that a creative session of drawings, songs, games and websites used roughly 11 percent of a weekly allowance in a single sitting. That suggests heavy generative sessions, not errands, are how people hit the ceiling. We did not verify the Gizmodo figure ourselves.

Muse Code: a session budget

Take a coding session with 60 agent turns. Suppose the context grows from 20,000 to 120,000 tokens over the session, averaging 70,000 tokens per turn, and each turn emits 1,500 output tokens. Input volume is 4.2 million tokens; output is 90,000 tokens.

Scenario Input cost Output cost Session total
No caching, standard tier 4.2M x $1.25 = $5.25 0.09M x $4.25 = $0.38 about $5.63
85 percent cache hits, standard tier 0.63M x $1.25 + 3.57M x $0.15 = $1.32 $0.38 about $1.70
Same, three observer agents each adding 20 percent of main input add about $0.79 add about $0.00 about $2.49
85 percent cache hits, contributor tier 0.63M x $0.10 + 3.57M x $0.002 = $0.07 0.09M x $0.20 = $0.02 about $0.09

The pattern is the point. Caching moves the session from about $5.63 to about $1.70, a 70 percent drop. Background agents claw back part of it. The contributor tier looks nearly free but trades away data rights. If a team runs 20 sessions a day across 10 developers, the standard-tier cached figure becomes about $340 a day, and without a spend cap an unattended loop that runs ten times longer than expected turns a $1.70 session into a $17 one.

That last sentence is the operational lesson. In a pay-as-you-go agent the failure mode is not a bad answer, it is a quiet runaway. Whether or not Meta adds a cap, you should add your own, which the recommendations below cover.

Comparing the two cost models

The consumer product uses a pooled weekly quota, so the risk is hitting a wall and the agent going quiet mid-task. The developer product uses unbounded metering, so the risk is a surprise bill. Different failure, different mitigation: for the first you monitor remaining quota and degrade to cheaper behaviour; for the second you enforce a ceiling outside the tool.

Trade-offs, Gotchas, and What Goes Wrong

Guardrails for always-on agents with budget gate, approval gate, sandbox and audit log

Figure 4: A defence-in-depth stack for an always-on agent. A budget gate, an approval gate, a sandbox and an audit log each stop a different class of failure.

Idle burn. The most common surprise is spend with no visible task. Wake-up frequency times context size is a fixed cost that scales with how much memory you load. Keep wake-up prompts small, filter events before involving the model, and prefer subscriptions to polling.

Compounding memory errors. A wrong fact stored once is retrieved forever. If the agent learns your preferred airline incorrectly, every later booking inherits the mistake. Memory needs timestamps, sources and a way for the user to review and edit it. The activity log reported for Muse is a good start, but a log of actions is not a log of beliefs.

Approval fatigue. A gate that fires on every action trains users to tap approve without reading. A gate that fires rarely misses things. The workable middle is tiered: auto-approve reversible low-value actions, require explicit approval for payments, deletions and anything that sends messages in your name, and make the approval card show the concrete effect, such as amount, merchant and recipient, rather than the model’s description.

Third-party friction. Reporting indicates Amazon has blocked Muse from its marketplace and that some sites restrict automated agent traffic. A browser agent depends on sites tolerating it, and that tolerance can change without notice. Design flows that fail gracefully when a site refuses.

Prompt injection through the inbox. An agent that triages email reads text written by strangers. Separating the Sentinel from the model reduces credential risk but does not stop the agent from being persuaded to take a harmful action that the Sentinel finds plausible. Treat all retrieved content as data, never as instructions, and keep consequential actions behind explicit user approval. Our AI agent audit logging and identity piece covers what to record for after-the-fact review.

Privacy trade-offs. Reported controls include a toggle to keep conversations out of training, and a statement that data is not used in advertising systems. Those are company claims, and Meta also retains operational access to user VMs. For regulated or sensitive data, assume a hosted agent sees what you connect to it.

Maturity varies by task. Coverage describes email and calendar as the mature use cases and bill negotiation or health tracking as less established. Start where the reported reliability is highest.

Practical Recommendations

If you are evaluating the Meta Muse agent as a user, start with low-stakes, reversible tasks and watch the activity log for a week before connecting payments. Connect the fewest accounts that give you value, and review what the agent has remembered. Note that availability is limited to the US, and reporting says the account requires adults and a payment card even on the free allowance, which is a sourced claim from ALM Corp that you should confirm in the app.

If you are evaluating Muse Code, assume cost is your responsibility. Pick the tier with eyes open: the contributor tier is cheap because you pay in data, so keep it away from private repositories. Set the reasoning level deliberately rather than accepting a default, check how many background agents are running, and measure cost per merged change on a pilot before scaling.

If you are building your own always-on agent, copy the structural ideas, not the branding. Separate the component that reads untrusted content from the component that holds credentials. Make the scheduler event-driven. Checkpoint every step so retries are safe. Put the budget gate outside the agent so that the agent cannot spend its way past it.

Checklist before you let any agent run unattended:

  • Set a hard spend or quota ceiling enforced outside the agent, with an alert at 50 and 80 percent.
  • Log every wake-up, tool call and approval with timestamps, and keep the log where the agent cannot edit it.
  • Require explicit approval for payments, deletions and outbound messages.
  • Keep wake-up context small and measure idle tokens per day as its own metric.
  • Store memory with source and timestamp, and give users a way to inspect and delete it.
  • Run on cached prompts where possible and track cache-hit rate.
  • Test the failure path: what happens when a site blocks the agent, a connector expires, or the quota runs out mid-task.
  • Re-check pricing against the vendor’s own page before budgeting, since figures in coverage conflict.

Frequently Asked Questions

What is the Meta Muse agent?

It is a hosted personal AI agent that Meta announced on 8 September 2026 for users in the United States. It runs the Muse Spark model inside a dedicated per-user virtual machine, keeps working after you close the app, remembers context across sessions, and routes outbound actions through a separate approval agent called Sentinel. It is distinct from the Meta AI chatbot and from the developer-focused Muse Code.

Is Meta Muse really free, and how much do the paid plans cost?

Reporting, including a quote attributed to Mark Zuckerberg, says Muse is free up to 100 million tokens per week. Meta’s newsroom confirms subscription plans exist for heavier use but lists no prices. Several outlets report a $20 monthly Power plan and a $100 monthly Max plan, but others note those figures come from coverage and community threads. Treat the paid prices as unconfirmed until Meta publishes them.

Does Muse Code have $5, $20 and $50 tiers?

We could not find support for that. The sources we read describe Muse Code as pay-as-you-go: $1.25 per million input tokens, $0.15 cached and $4.25 output on the standard tier, with a cheaper contributor tier that permits training on your prompts. One pricing analysis states explicitly that no subscription plans exist for it. The tier claim appears to be a mix-up with consumer Muse plan reporting.

How does Muse keep an AI agent from misusing my accounts?

Meta describes a separate Sentinel agent on the same VM that must approve anything reaching the internet. Secondary reporting adds that the model sees placeholder tokens instead of real credentials and that payments use one-time virtual cards, with approval cards shown outside the chat. This limits what a manipulated model can do, but it does not remove the need to review approvals, because the gate approves based on a description of the action.

How much does an always-on agent cost to run?

It depends on wake-up frequency, context size and task mix. As an illustrative example, a 15-minute wake-up cycle with 5,000 tokens per wake-up burns about 3.4 million tokens a week doing nothing, and a single browser task with large page snapshots can use around a million. Caching, small wake-up prompts and an external spending ceiling are the main controls. Measure idle tokens separately from task tokens.

Should developers use Muse Code instead of Claude Code or Codex?

That depends on your constraints, not on a headline price. Muse Code’s standard rates are lower than the Sonnet 5 rates cited by Engadget, and its context window is reported above one million tokens. But defaults such as reasoning level and background agents affect real cost, the contributor tier gives up data rights, and reported benchmark scores are not a substitute for a trial on your own codebase. Pilot on a non-sensitive repository first.

Further Reading

By Riju — about

Comments

No comments yet. Why don’t you start the discussion?

Leave a Reply

Your email address will not be published. Required fields are marked *