Muse Glimmer is a chat-oriented large language model from Meta, tracked here with measured throughput of roughly 112 output tokens per second and a time to first token near 382 ms.
| Provider | Input / 1M | Output / 1M | Cached / 1M |
|---|---|---|---|
| $0.300 | $1.20 | $0.040 | |
| $0.300 | $1.10 | $0.040 | |
| $0.350 | $1.50 | - | |
| $0.350 | $1.50 | - | |
| $0.350 | $1.50 | $0.040 |
Prices updated daily. Last check: Sep 5, 2026
Input, output, and batch rates, plus alternatives, for one provider at a time.
Benchmarks measured Sep 2026. Scores are independent evaluations, not vendor-reported.
Muse Glimmer fits interactive chat workloads where responsiveness is a priority: customer-facing assistants, in-product help and search chat, drafting and rewriting tools, and conversational agents that stream tokens back to a user. Its sub-400 ms time to first token means users see output almost immediately, and its ~112 tokens per second generation rate keeps multi-paragraph replies from feeling slow. Because we do not have confirmed context-window or benchmark data for this model, it is a weaker fit for workloads that hinge on very long documents, verified reasoning accuracy, or multimodal input until those specifications are confirmed with your serving provider. For large batch pipelines, compare its throughput and per-token cost against other entries on this page before committing.
Pricing depends on which inference provider hosts the model and on the pricing type — separate input and output token rates, cached input discounts, and batch or reserved options can all differ. Because these rates change frequently, check the live pricing table on this page for current figures across providers.
It is a chat model, so it suits conversational assistants, instruction-following tasks, drafting and summarization, and any user-facing feature that streams responses. Its ~382 ms time to first token and ~112 tokens per second output rate make it a reasonable candidate for interactive rather than latency-insensitive batch use.
Independent measurements from Artificial Analysis report approximately 111.5 output tokens per second with a time to first token of about 382 ms. Actual figures vary by provider, region, prompt length, and load, so treat these as a reference point rather than a guarantee.
Muse Glimmer is attributed to Meta in our catalog.
We do not currently track a confirmed context window for this model. If long-document handling matters for your workload, confirm the supported context length with the provider you plan to use.
Modality support beyond text chat is not something we have confirmed for Muse Glimmer. Check the provider's API documentation before building around image or audio input.