Meta's Superintelligence Labs released Muse Glimmer on August 10, 2026, making a roughly 30-billion-parameter agent model available on Hugging Face under Apache 2.0. Meta says a quantized version runs on a single GPU with 24 GB of video memory. That puts it within reach of a high-end workstation rather than a dedicated inference cluster.
The release is aimed at developers building agents that use tools and carry out tasks across multiple steps. Its appeal depends on more than fitting the weights into memory. The practical questions are how fast it runs, how reliably it completes tasks, and whether operating it locally costs less than using an API.
The model and its hardware requirements
Muse Glimmer combines a 29.6-billion-parameter dense causal transformer with a 1.8-billion-parameter ViT-G/14 perception encoder. The encoder handles visual input, giving the model text and vision capabilities. Meta positions it as a model for end-to-end agentic task completion, with persistent state across restarts and reliable tool-calling.
Meta provides two quantized variants. Quantization reduces the precision used to store model values, lowering memory requirements at a possible cost to output quality. The reported tradeoffs are:
- 32 GB VRAM: 0.2% performance degradation compared with the full model.
- 24 GB VRAM: 1.0% performance degradation, with support for a single RTX 5090, RTX 4090, or equivalent card.
Those are Meta's figures. The 24 GB option is the more consequential one for developers with existing hardware. It means a workstation or home server with a suitable consumer GPU could run the model without a DGX-class purchase.
Apple Silicon is also supported. With DFlash speculative decoding enabled, reported generation speeds are 37.8 tokens per second on an M4 Max and 50.2 tokens per second on an M5 Max. Those rates look usable for an interactive coding agent or background development work, although generation speed alone doesn't establish how quickly an agent will finish a task.
DFlash accounts for much of the reported speed
On an RTX 5090, Muse Glimmer reportedly generates 74.9 tokens per second without speculative decoding. With DFlash speculative decoding, that rises to 233.4 tokens per second, a 3.1× speedup on the same card.
Speculative decoding has been around since 2022. A smaller model drafts tokens quickly, then the full model checks batches of those tokens. When enough of the draft is accepted, the system can produce output faster than it would by having the full model generate each token separately.
Meta presented DFlash at ICML 2026. Its approach changes memory access patterns during verification, a step that has imposed substantial overhead in earlier implementations on larger models. Meta's paper reports a 2.5× improvement over prior state-of-the-art speculative decoding methods and greater than 6× lossless acceleration over standard autoregressive decoding in lab settings. The lossless claim means the acceleration is intended to preserve the full model's output distribution.
The distinction between the lab result and the RTX 5090 result matters. The reported consumer-card speedup is 3.1×, not greater than 6×. If comparable gains hold across real agent workloads, DFlash would be a strong candidate for wider adoption in local inference software.
Runtime support was available at launch. Ollama released version 0.32.7 on August 10 with Muse Glimmer support. vLLM, llama.cpp, and ExecuTorch are also listed as supported runtimes. That gives developers several ways to try the model without waiting for basic serving support.
Competitive scores, pending independent checks
Meta reports the following benchmark scores:
| Benchmark | Reported score | Task area |
|---|---|---|
| SWE-Bench Verified | 76.0 | Resolving real GitHub issues |
| AIME 2026 | 94.7 | Competition mathematics |
| GPQA Diamond | 83.5 | Graduate-level science questions |
| MCP Atlas | 75.5 | Agentic tool use |
| DeepSearch QA | 74.6 | Retrieval-augmented question answering |
SWE-Bench Verified and MCP Atlas are especially relevant to the model's intended role. A coding agent needs to make changes that pass real tests, and a tool-using agent needs to choose and execute calls reliably. SWE-Bench Verified checks code against test suites, making it more directly useful for assessing coding work than a benchmark based only on plausible-looking answers.
A reported SWE-Bench Verified score of 76.0 is promising for a locally hosted model. If it translates to a team's own repositories, that capability could be available without sending prompts to an external inference provider or paying for each API token.
As of August 11, 2026, none of these scores had been independently verified. Meta also hadn't released red-team documentation or formal security assessments for the model. The benchmark results justify evaluation, but they don't establish that the model is ready for unrestricted filesystem access or production API credentials.
Apache 2.0 makes commercial use easier to assess
Open weights don't always come with permissive terms. Custom model licenses can restrict commercial use, particular applications, fine-tuning, or use by competing businesses. Those restrictions can make an otherwise suitable model difficult to build a product around.
Apache 2.0 permits commercial use, fine-tuning, redistribution, and embedding a model in a SaaS product without a general requirement to publish modifications, subject to the license's terms. It is also familiar from projects such as Kubernetes, TensorFlow, Kafka, and Spark.
That familiarity is useful for an internal coding agent or an infrastructure assistant that reads runbooks and queries monitoring APIs. A standard permissive license makes the permissions easier to assess than a custom model agreement. It doesn't remove the need to review the deployment's legal, privacy, and security requirements.
Where local inference could beat an API
Hosted APIs have often been the practical default because local inference requires maintenance and locally available models have lagged frontier quality. Muse Glimmer's reported capability and hardware requirements make that choice less automatic.
Several workloads have a clear reason to consider local hosting:
- Sensitive data. Healthcare records, financial information, source code under strict IP agreements, and customer PII in regulated jurisdictions can be difficult to send through third-party APIs. Local inference can keep model processing inside the organization's infrastructure, though the rest of the agent's tools and data flows still need review.
- High, steady token volume. Always-on coding agents, CI analysis, documentation generation, and test writing can accumulate substantial API charges. A GPU purchase can be spread across millions of tokens rather than incurring a provider's charge for each token.
- Interactive response times. A reported generation rate of 233.4 tokens per second is attractive for local use, which also avoids a network round trip to an inference provider. Actual responsiveness will depend on prompt processing, queueing, tool calls, and the workload.
- Control over availability and versions. Local hosting reduces exposure to API price changes, tighter rate limits, and model deprecations. It also transfers responsibility for keeping the service running to its operator.
The operating costs remain substantial. GPU drivers, model serving, monitoring, updates, and hardware failures all require attention. A one-time hardware purchase doesn't eliminate power costs or engineering time.
A 24 GB consumer card also isn't equivalent to a production inference service. A single-user generation result doesn't show how the system behaves when several people submit requests at once. Shared deployments need tests for aggregate throughput, queueing, and what happens when the GPU is already busy.
The release lowers the hardware threshold for a serious local evaluation. A workstation with a high-end consumer GPU, or a suitable Mac Studio, may now be enough to test whether an agent can handle a useful workload. Whether that becomes the cheaper production option depends on utilization and support requirements.
The next evidence to look for
Independent results on SWE-Bench Verified and MCP Atlas would provide the most useful early check of Meta's agent claims. Community testing could arrive within days, but the useful comparison will be performance under documented conditions rather than another headline score.
Apache 2.0 also leaves room for domain-specific fine-tunes, including legal document analysis, infrastructure operations, and security review. The base model's quality will strongly influence how useful those adaptations become.
Security testing is a separate need. Tool access, filesystem permissions, and persistent state give an agent ways to cause damage if hostile inputs influence its behavior. The absence of published red-team documentation isn't necessarily a reason to reject the model, but it is a reason to withhold write access to important systems until that gap is addressed.
DFlash deserves attention beyond Muse Glimmer. If its reported gains hold across workloads and it can be adapted to other model families, more local models could become practical on existing hardware. That would broaden the choices available to developers beyond this single release.