Meta Muse Glimmer brings local AI agents to consumer GPUs
Meta is releasing Muse Glimmer below an Apache 2.0 licence for local AI agents that may run on a consumer GPU.
The firm’s Superintelligence Labs has launched the 30-billion-parameter mannequin’s weights on Hugging Face. Meta says builders can use it for local coding, operate calling, local agents, and LLM-as-a-judge analysis.
The launch targets an operational constraint going through AI groups: cloud-hosted fashions want community entry and central infrastructure. Meta as an alternative pitches Muse Glimmer for workloads that require an on-device mannequin, together with private agents with entry to schedules, messages, recordsdata, and different non-public context.
Meta Muse Glimmer leads a number of agent activity benchmarks
Meta’s benchmark checks put Muse Glimmer forward of Gemma4-31B and Qwen3.6-27B on 5 of eight general-agentic benchmarks. The mannequin scored 75.5 on MCP Atlas. Gemma4-31B reached 54.2, and Qwen3.6-27B recorded 62.5.
DeepSearch QA follows the same sample. Meta studies a rating of 74.6 for Muse Glimmer, towards 61.7 for Gemma4-31B and 71.1 for Qwen3.6-27B. The equipped announcement identifies each benchmarks as checks of an agent’s means to work inside scaffolds and full multi-turn requests.
The mannequin scored 23.5 on τ²-Banking. Gemma4-31B recorded 15.1. Qwen3.6-27B reached 16.7.
Muse Glimmer additionally posted 47.6 on WildClawBench, forward of Gemma4-31B’s 37.6 and Qwen3.6-27B’s 43.2. Its GAIA2 consequence reached 43.3, in contrast with 36.4 and 40.0 respectively.
Other agent scores favour Qwen3.6-27B. Meta’s desk offers that mannequin 1,141 on GDPval-AA, towards Muse Glimmer’s 953 and Gemma4-31B’s 811. Qwen3.6-27B additionally led SkillsBench with Skills at 46.6, the place Muse Glimmer recorded 44.3.
OSWorld-Verified produced the biggest hole on this group. Meta studies 75.6 for Qwen3.6-27B. Muse Glimmer reached 65.9, and Gemma4-31B scored 58.5.
These checks measure constrained duties. They don’t exhibit how a local agent will behave after an organisation connects it to its personal recordsdata, calendars, messaging programs, or inside instruments.
Coding outcomes cut up between Muse Glimmer and Qwen
Muse Glimmer’s coding outcomes present a narrower comparability. The mannequin led SWE-Bench Pro with a rating of 51.2. Meta studies 36.9 for Gemma4-31B and 50.2 for Qwen3.6-27B.
SciCode produced an in depth consequence. Muse Glimmer scored 43.6, marginally above Gemma4-31B at 43.4. Qwen3.6-27B recorded 39.8.
Qwen3.6-27B led two different coding evaluations. It scored 77.2 on SWE-Bench Verified, in contrast with Muse Glimmer’s 76.0. TerminalBench 2.1 gave Qwen3.6-27B a rating of 60.7; Muse Glimmer reached 51.7, and Gemma4-31B posted 43.4.
A local coding agent does greater than produce code. It wants a scaffold that decides which repositories, terminals, check environments, and instructions the mannequin could entry. Meta says Muse Glimmer helps OpenClaw and different agent-orchestration patterns, with customized scaffolds coated in its developer documentation.
An organisation evaluating the mannequin for software program work ought to outline the instructions and repositories accessible to the agent earlier than measuring activity success. The equipped materials describes retry coaching for failed software calls. That behaviour requires controls over repeat makes an attempt, particularly the place a software can alter supply code or invoke an exterior system.
Multimodal scores favour Qwen in most checks
Muse Glimmer accepts interleaved textual content and pictures by means of a devoted notion encoder. Meta says this design lets agents interpret screenshots, charts, and paperwork as a part of a dialog.
The benchmark chart places Muse Glimmer forward on Charxiv Reasoning. Its rating reached 78.8, towards 77.7 for Gemma4-31B and 78.4 for Qwen3.6-27B.
Qwen3.6-27B led ScreenSpot Pro with 76.1. Muse Glimmer recorded 75.4, and Gemma4-31B scored 75.9. The similar mannequin led OmniDocBench v1.5 at 77.8, in contrast with Muse Glimmer’s 75.8 and Gemma4-31B’s 72.5.
MMMU Pro produced smaller variations. Meta lists Muse Glimmer at 74. Qwen3.6-27B reached 75, and Gemma4-31B posted 73.
These outcomes matter for groups contemplating agents that act on visible interfaces. A screenshot-reading mannequin can interpret what it sees, but local testing should nonetheless cowl permissions, show layouts, doc codecs, and errors returned by related instruments.
Safety figures present decrease reported assault success than Qwen
Meta additionally studies two safety-related evaluations: CI Memories and Siren AgentDojo. The chart makes use of completely different measures for every check.
On CI Memories, Meta lists a violation charge of 26.4 for Muse Glimmer and a protection rating of 64.8. Gemma4-31B recorded a violation charge of 12.1 with protection of 53.0. Qwen3.6-27B posted a violation charge of 53.4 and protection of 66.9.
The Siren AgentDojo consequence makes use of assault success charge and utility. Meta offers Muse Glimmer an assault success charge of 28.4 and a utility rating of 94.2. Gemma4-31B scored 25.6 on assault success charge, with utility at 90.8. Qwen3.6-27B recorded 40.3 and 92.7.
General reasoning outcomes add context to agent claims
Muse Glimmer led 4 of six general-capabilities-and-reasoning checks in Meta’s comparability. It scored 77.0 on IFBench. Gemma4-31B recorded 76.0, and Qwen3.6-27B reached 70.8.
The AIME 2026 rating was 94.7 for Muse Glimmer. Meta studies 89.2 for Gemma4-31B and 94.1 for Qwen3.6-27B. On AA-LCR, Muse Glimmer reached 80.0, forward of 68.3 and 73.3.
The mannequin additionally led Beam 128K at 65.1. Qwen3.6-27B scored 63.0. Gemma4-31B recorded 58.2.
Gemma4-31B led GPQA Diamond with 85.7. Muse Glimmer scored 83.5, adopted by Qwen3.6-27B at 84.2. Gemma4-31B additionally took the highest rating on Humanity’s Last Exam, Text No Tools, at 23.6; Muse Glimmer reached 22.0.
One mannequin doesn’t lead each check. Meta’s outcomes as an alternative present Muse Glimmer competing intently with two equally sized fashions throughout a blended set of agent, coding, visible, security, and reasoning evaluations.
Memory limits form the local deployment design
Meta says a full-precision 30-billion-parameter mannequin would require greater than 55 GB of reminiscence. Muse Glimmer as an alternative makes use of roughly 4-bit weight quantisation, decreasing the language mannequin to below 20 GB.
That allocation leaves reminiscence for a KV cache. The mannequin additionally wants room for its notion encoder and a speculative-decoding drafter. Meta targets a 24 GB or 32 GB reminiscence envelope for these elements.
The firm says the DFlash-based drafter proposes blocks of tokens for the principle mannequin to confirm in parallel. Meta says this speeds technology in contrast with commonplace token-by-token output and retains an identical output high quality. The equipped publish doesn’t embody token-per-second figures, immediate sizes, energy knowledge, or concurrency outcomes.
Meta examined its Okay-Quant-17GB model with the quantised DFlash drafter on MacBook M4-Max {hardware}, MacBook M5-Max {hardware}, and an RTX-5090. It describes the ensuing expertise as appropriate for fluid dialog and real-time agent interplay.
The public weights can be found by means of Hugging Face. Meta says integrations with llama.cpp, MLX, and ExecuTorch will arrive within the coming days.
See additionally: Alibaba tests new business model for Qwen open-source AI

Want to study extra about AI and massive knowledge from trade leaders? Check out AI & Big Data Expo happening in Amsterdam, California, and London. The complete occasion is a part of TechEx and is co-located with different main expertise occasions together with the Cyber Security & Cloud Expo. Click here for extra data.
AI News is powered by TechForge Media. Explore different upcoming enterprise expertise occasions and webinars here.
The publish Meta Muse Glimmer brings local AI agents to consumer GPUs appeared first on AI News.
