Skip to main content
FlagshipOpen SourceThinking Machines Lab

Inkling

Inkling is an open-weight, multimodal flagship model from Thinking Machines Lab that accepts text, image, and audio input and supports a context window of roughly 1M tokens.

Context 1.0M
Tier Flagship
Tools Supported
License Open Source
Modalities text, image, audio
Input from
$0.950 / 1M tokens
across 4 providers

API Pricing

Cheapest on Deep Infra — 4% below avg
ProviderInput / 1MOutput / 1MCached / 1M
$0.950$4.05$0.160
$1.00$4.05-
$1.00$4.05-
$1.00$4.05$0.170

Prices updated daily. Last check: Sep 30, 2026

Inkling pricing by provider

Input, output, and batch rates, plus alternatives, for one provider at a time.

Performance & Benchmarks

Source: Artificial Analysis →
Intelligence
25.0 / 100
Coding
52.1 / 100
Output Speed
171 t/s
Latency (TTFT)
2.3s

Reasoning & Knowledge

  • GPQA Diamond87.2%
  • Humanity's Last Exam31.9%

Coding

  • SciCode47.0%

Agentic & Tool Use

  • Terminal-Bench v2.155.1%
  • τ-bench Banking29.1%

Instruction & Long Context

  • Long-Context Reasoning77.3%

Benchmarks measured Sep 2026. Scores are independent evaluations, not vendor-reported.

Model Details

General

Creator
Thinking Machines Lab
Family
Inkling
Tier
Flagship
Context Window
1.0M
Modalities
Text, Image, Audio

Capabilities

Tool Calling
Yes
Open Source
Yes
Aliases
thinkingmachines/inkling, thinkingmachines/Inkling, Inkling

Strengths & Limitations

Strengths

  • Context window of 1,048,576 tokens (about 1M) supports whole-corpus and long-transcript workloads without chunking
  • Open-weight release allows self-hosting and lets multiple inference providers compete on price and region
  • Accepts three input modalities — text, image, and audio — in a single model
  • Native tool calling support for agent loops, retrieval pipelines, and function-based integrations
  • Flagship tier of the Inkling family, positioned for the broadest range of tasks rather than a narrow niche
  • Measured output throughput of roughly 80 tokens/second (Artificial Analysis), suitable for long-form generation

Limitations

  • Time to first token of roughly 2,086 ms in Artificial Analysis testing, which is slow for latency-sensitive chat or autocomplete
  • Output throughput of about 80 tokens/second is mid-range rather than fast among served models
  • Flagship-tier compute requirements make self-hosting hardware-intensive compared with smaller open models
  • We do not track published academic benchmark scores (e.g. MMLU, coding suites) for Inkling, so quality comparisons rely on independent testing
  • Filling the full 1M-token context has real cost and latency consequences regardless of the per-token rate

Key Features

•1,048,576-token (≈1M) context window
•Multimodal input: text, image, and audio
•Tool / function calling support
•Open weights available for self-hosting
•Flagship tier within the Inkling family from Thinking Machines Lab
•Measured ~80.4 output tokens/second (Artificial Analysis)
•Available through multiple third-party inference providers
•Model aliases including thinkingmachines/inkling for API routing

About Inkling

Inkling is a large language model from Thinking Machines Lab and the model that gives the Inkling family its name. It sits at the flagship tier within that family, meaning it is positioned for the broadest and most demanding workloads Thinking Machines Lab targets rather than for cost-optimized, high-volume serving. Inkling is released as an open-weight model, so it can be hosted by multiple inference providers as well as run on self-managed infrastructure, which is why pricing for it can differ substantially from one provider to another. Inkling handles a 1,048,576-token context window (roughly 1M tokens) and accepts three input modalities: text, image, and audio. It also supports tool calling, which allows it to be wired into function-based workflows, retrieval systems, and agent frameworks. On throughput measurements collected by Artificial Analysis, Inkling generates output at about 80 tokens per second with a time to first token of roughly 2,086 ms — a profile that favors long, sustained generations over snappy interactive turnaround. In practice, the combination of a million-token context and multimodal input makes Inkling a candidate for workloads where the whole input has to be resident at once: large document sets, long transcripts paired with audio, or mixed text-and-image corpora. Compared with closed flagship models, the main structural difference is that Inkling's open weights allow self-hosting and provider choice; compared with smaller open models, the trade-off is the higher compute footprint that a flagship-tier model implies.

Common Use Cases

Inkling suits workloads that need a large working context and more than one input modality at once: summarizing or querying across very large document collections, analyzing long audio recordings alongside their transcripts, extracting structured data from mixed image-and-text sources, and building agents that call external tools over long-running sessions. Its flagship positioning within the Inkling family makes it the choice for tasks where reasoning quality matters more than per-token cost, while its roughly 2-second time to first token makes it a better fit for batch and background jobs than for latency-critical interactive chat. Because the weights are open, teams with data-residency or on-premise requirements can also evaluate it for self-hosted deployment rather than API-only use.

Frequently Asked Questions

How much does Inkling cost to run?

Pricing varies by inference provider and by pricing type — per-token serverless rates, dedicated capacity, and self-hosted GPU costs all behave differently, and because Inkling is open-weight, multiple providers may host it at different rates. Check the pricing table on this page for current per-provider figures.

What is Inkling best used for?

Long-context and multimodal work: querying large document sets, processing long audio alongside text, extracting information from images and text together, and driving tool-calling agents. Its ~1M-token context and support for text, image, and audio input are the main reasons to pick it over a smaller model.

Can I run Inkling on my own hardware?

Yes — Inkling is released with open weights, so it can be self-hosted in addition to being consumed through hosted APIs. As a flagship-tier model it is compute-intensive, so plan GPU capacity accordingly; our GPU pricing pages can help estimate that side of the cost.

Is Inkling fast enough for real-time chat?

Artificial Analysis measurements put Inkling at roughly 80 output tokens per second with a time to first token of about 2,086 ms. The throughput is workable for streaming long responses, but the two-second first-token delay is noticeable in interactive chat, so latency-sensitive front ends may prefer a lighter model.

Does Inkling support tool calling?

Yes. Inkling supports tool/function calling, which allows it to be integrated into agent frameworks, retrieval pipelines, and workflows that require calling external APIs. Exact schema and parallel-call behavior depend on the provider's API implementation.