Skip to main content
Meta

Muse Glimmer

Muse Glimmer is a chat-oriented large language model from Meta, tracked here with measured throughput of roughly 112 output tokens per second and a time to first token near 382 ms.

Input from
$0.300 / 1M tokens
across 4 providers

API Pricing

Cheapest on Deep Infra — 8% below avg
ProviderInput / 1MOutput / 1MCached / 1M
$0.300$1.20$0.040
$0.300$1.20$0.040
$0.350$1.50-
$0.350$1.50-

Prices updated daily. Last check: Sep 25, 2026

Muse Glimmer pricing by provider

Input, output, and batch rates, plus alternatives, for one provider at a time.

Performance & Benchmarks

Source: Artificial Analysis →
Intelligence
17.5 / 100
Coding
49.0 / 100
Output Speed
107 t/s
Latency (TTFT)
482ms

Reasoning & Knowledge

  • GPQA Diamond83.5%
  • Humanity's Last Exam22.0%

Coding

  • SciCode44.9%

Agentic & Tool Use

  • Terminal-Bench v2.151.7%
  • τ-bench Banking23.5%

Instruction & Long Context

  • Long-Context Reasoning83.3%

Benchmarks measured Sep 2026. Scores are independent evaluations, not vendor-reported.

Model Details

General

Creator
Meta
Modalities
Text

Capabilities

Tool Calling
No
Open Source
No

Strengths & Limitations

Strengths

  • Measured output throughput of about 111.5 tokens per second, suitable for streaming chat interfaces
  • Time to first token of roughly 382 ms, keeping perceived response latency low
  • Performance figures come from an independent third party (Artificial Analysis) rather than vendor-reported claims
  • Chat-tuned model, so it can be dropped into instruction-following and assistant workloads without additional prompt scaffolding
  • Built by Meta, a creator with an established line of openly documented assistant models
  • Listed alongside competing models on this page so throughput and pricing can be compared directly

Limitations

  • We do not track a confirmed context window for Muse Glimmer, so long-input suitability must be verified with the provider
  • No published accuracy benchmarks in our data — quality relative to other Meta chat models is unverified
  • Modality support (for example image input) is not something we have confirmed for this model
  • Throughput near 112 tokens per second is mid-range; some competing endpoints are measured faster
  • Provider availability may be narrow, which limits price shopping and failover options

Key Features

•Chat and instruction-following interface
•Measured ~111.5 output tokens per second (Artificial Analysis)
•Measured ~382 ms time to first token
•Token streaming suitable for interactive applications
•Created by Meta
•Hosted API access through inference providers listed in the pricing table
•Multi-turn conversational use

About Muse Glimmer

Muse Glimmer is a chat model attributed to Meta in our catalog. It is designed for conversational and instruction-following workloads, the same broad category as other Meta-built assistant models, and appears on this page so that its hosted API pricing can be compared side by side with alternatives from other providers. Our verified metadata for Muse Glimmer is limited to serving performance rather than architecture. Third-party measurements from Artificial Analysis put its output speed at about 111.5 tokens per second with a time to first token of approximately 382 ms. Those two figures describe an interactive-feeling response profile: prompts begin streaming in well under half a second and generate at a rate comfortable for live chat, streaming UIs, and turn-based assistants. Context window, parameter count, modality support, and benchmark accuracy scores are not fields we currently track for this model, so we make no claims about them here. In practice, a model with this latency and throughput profile is used for user-facing chat, drafting, summarization of moderate-length inputs, and other request/response tasks where responsiveness matters as much as raw capability. Because our data on Muse Glimmer is thinner than for more widely benchmarked Meta models, readers evaluating it for production should validate quality on their own prompts and confirm the technical specifications with the serving provider they intend to use.

Common Use Cases

Muse Glimmer fits interactive chat workloads where responsiveness is a priority: customer-facing assistants, in-product help and search chat, drafting and rewriting tools, and conversational agents that stream tokens back to a user. Its sub-400 ms time to first token means users see output almost immediately, and its ~112 tokens per second generation rate keeps multi-paragraph replies from feeling slow. Because we do not have confirmed context-window or benchmark data for this model, it is a weaker fit for workloads that hinge on very long documents, verified reasoning accuracy, or multimodal input until those specifications are confirmed with your serving provider. For large batch pipelines, compare its throughput and per-token cost against other entries on this page before committing.

Frequently Asked Questions

How much does Muse Glimmer cost to use?

Pricing depends on which inference provider hosts the model and on the pricing type — separate input and output token rates, cached input discounts, and batch or reserved options can all differ. Because these rates change frequently, check the live pricing table on this page for current figures across providers.

What is Muse Glimmer best used for?

It is a chat model, so it suits conversational assistants, instruction-following tasks, drafting and summarization, and any user-facing feature that streams responses. Its ~382 ms time to first token and ~112 tokens per second output rate make it a reasonable candidate for interactive rather than latency-insensitive batch use.

How fast is Muse Glimmer?

Independent measurements from Artificial Analysis report approximately 111.5 output tokens per second with a time to first token of about 382 ms. Actual figures vary by provider, region, prompt length, and load, so treat these as a reference point rather than a guarantee.

Who created Muse Glimmer?

Muse Glimmer is attributed to Meta in our catalog.

What context window does Muse Glimmer support?

We do not currently track a confirmed context window for this model. If long-document handling matters for your workload, confirm the supported context length with the provider you plan to use.

Can Muse Glimmer accept images or other non-text input?

Modality support beyond text chat is not something we have confirmed for Muse Glimmer. Check the provider's API documentation before building around image or audio input.