Inkling
Inkling is an open-weight, multimodal flagship model from Thinking Machines Lab that accepts text, image, and audio input and supports a context window of roughly 1M tokens.
API Pricing
Cheapest on Deep Infra — 4% below avg| Provider | Input / 1M | Output / 1M | Cached / 1M |
|---|---|---|---|
| $0.950 | $4.05 | $0.160 | |
| $1.00 | $4.05 | - | |
| $1.00 | $4.05 | - | |
| $1.00 | $4.05 | $0.170 |
Prices updated daily. Last check: Sep 30, 2026
Inkling pricing by provider
Input, output, and batch rates, plus alternatives, for one provider at a time.
Performance & Benchmarks
Source: Artificial Analysis →Reasoning & Knowledge
- GPQA Diamond87.2%
- Humanity's Last Exam31.9%
Coding
- SciCode47.0%
Agentic & Tool Use
- Terminal-Bench v2.155.1%
- τ-bench Banking29.1%
Instruction & Long Context
- Long-Context Reasoning77.3%
Benchmarks measured Sep 2026. Scores are independent evaluations, not vendor-reported.
Model Details
General
- Creator
- Thinking Machines Lab
- Family
- Inkling
- Tier
- Flagship
- Context Window
- 1.0M
- Modalities
- Text, Image, Audio
Capabilities
- Tool Calling
- Yes
- Open Source
- Yes
- Aliases
- thinkingmachines/inkling, thinkingmachines/Inkling, Inkling
Strengths & Limitations
Strengths
- Context window of 1,048,576 tokens (about 1M) supports whole-corpus and long-transcript workloads without chunking
- Open-weight release allows self-hosting and lets multiple inference providers compete on price and region
- Accepts three input modalities — text, image, and audio — in a single model
- Native tool calling support for agent loops, retrieval pipelines, and function-based integrations
- Flagship tier of the Inkling family, positioned for the broadest range of tasks rather than a narrow niche
- Measured output throughput of roughly 80 tokens/second (Artificial Analysis), suitable for long-form generation
Limitations
- Time to first token of roughly 2,086 ms in Artificial Analysis testing, which is slow for latency-sensitive chat or autocomplete
- Output throughput of about 80 tokens/second is mid-range rather than fast among served models
- Flagship-tier compute requirements make self-hosting hardware-intensive compared with smaller open models
- We do not track published academic benchmark scores (e.g. MMLU, coding suites) for Inkling, so quality comparisons rely on independent testing
- Filling the full 1M-token context has real cost and latency consequences regardless of the per-token rate
Key Features
About Inkling
Common Use Cases
Inkling suits workloads that need a large working context and more than one input modality at once: summarizing or querying across very large document collections, analyzing long audio recordings alongside their transcripts, extracting structured data from mixed image-and-text sources, and building agents that call external tools over long-running sessions. Its flagship positioning within the Inkling family makes it the choice for tasks where reasoning quality matters more than per-token cost, while its roughly 2-second time to first token makes it a better fit for batch and background jobs than for latency-critical interactive chat. Because the weights are open, teams with data-residency or on-premise requirements can also evaluate it for self-hosted deployment rather than API-only use.
Frequently Asked Questions
How much does Inkling cost to run?
Pricing varies by inference provider and by pricing type — per-token serverless rates, dedicated capacity, and self-hosted GPU costs all behave differently, and because Inkling is open-weight, multiple providers may host it at different rates. Check the pricing table on this page for current per-provider figures.
What is Inkling best used for?
Long-context and multimodal work: querying large document sets, processing long audio alongside text, extracting information from images and text together, and driving tool-calling agents. Its ~1M-token context and support for text, image, and audio input are the main reasons to pick it over a smaller model.
Can I run Inkling on my own hardware?
Yes — Inkling is released with open weights, so it can be self-hosted in addition to being consumed through hosted APIs. As a flagship-tier model it is compute-intensive, so plan GPU capacity accordingly; our GPU pricing pages can help estimate that side of the cost.
Is Inkling fast enough for real-time chat?
Artificial Analysis measurements put Inkling at roughly 80 output tokens per second with a time to first token of about 2,086 ms. The throughput is workable for streaming long responses, but the two-second first-token delay is noticeable in interactive chat, so latency-sensitive front ends may prefer a lighter model.
Does Inkling support tool calling?
Yes. Inkling supports tool/function calling, which allows it to be integrated into agent frameworks, retrieval pipelines, and workflows that require calling external APIs. Exact schema and parallel-call behavior depend on the provider's API implementation.