Skip to main content
LightweightOpen SourceThinking Machines Lab

Inkling-Small

Inkling-Small is a lightweight, open-weight multimodal model from Thinking Machines Lab, supporting text, image, and audio input with a 1,048,576-token context window.

Context 1.0M
Tier Lightweight
Tools Supported
License Open Source
Modalities text, image, audio
Input from
$0.450 / 1M tokens
across 2 providers

API Pricing

ProviderInput / 1MOutput / 1MCached / 1M
$0.450$1.20$0.100
$0.450$1.20$0.100

Prices updated daily. Last check: Sep 30, 2026

Inkling-Small pricing by provider

Input, output, and batch rates, plus alternatives, for one provider at a time.

Performance & Benchmarks

Source: Artificial Analysis →
Intelligence
27.8 / 100
Coding
52.9 / 100
Output Speed
228 t/s
Latency (TTFT)
1.7s

Reasoning & Knowledge

  • GPQA Diamond89.5%
  • Humanity's Last Exam33.3%

Coding

  • SciCode49.7%

Agentic & Tool Use

  • Terminal-Bench v2.155.1%
  • τ-bench Banking18.8%

Instruction & Long Context

  • Long-Context Reasoning75.7%

Benchmarks measured Sep 2026. Scores are independent evaluations, not vendor-reported.

Model Details

General

Creator
Thinking Machines Lab
Family
Inkling
Tier
Lightweight
Context Window
1.0M
Modalities
Text, Image, Audio

Capabilities

Tool Calling
Yes
Open Source
Yes
Aliases
thinkingmachines/inkling-small, thinkingmachines/Inkling-Small, Inkling Small, Inkling-Small

Strengths & Limitations

Strengths

  • Context window of 1,048,576 tokens, unusually large for a lightweight-tier model
  • Accepts text, image, and audio inputs in a single model
  • Open weights, allowing self-hosting alongside hosted API access
  • Tool calling support for agentic and function-invoking workflows
  • Measured output throughput of roughly 72 tokens per second (Artificial Analysis)
  • Lightweight tier positioning suits high-volume, cost-sensitive batch workloads
  • Multiple provider aliases (thinkingmachines/inkling-small) make it addressable across hosted endpoints

Limitations

  • Lightweight tier — larger Inkling models are a better fit for tasks demanding deep reasoning or long multi-step planning
  • Time to first token measured at about 1.57 seconds, which is noticeable in interactive chat use
  • Filling the full million-token context is expensive in practice regardless of the per-token rate
  • We do not track detailed reasoning, coding, or knowledge benchmark scores for this model, so capability comparisons against peers are hard to quantify
  • Provider availability for a comparatively new model family may be narrower than for widely hosted alternatives

Key Features

•1,048,576-token (1M) context window
•Text input
•Image (vision) input
•Audio input
•Tool / function calling
•Open model weights available for self-hosting
•Measured throughput of ~72 output tokens/sec (Artificial Analysis)
•Lightweight tier of the Inkling family

About Inkling-Small

Inkling-Small is a model in the Inkling family from Thinking Machines Lab. It sits at the lightweight tier of the family, positioning it for workloads where throughput and cost efficiency matter more than maximum capability. The model is released with open weights, which means it can be self-hosted in addition to being consumed through hosted inference APIs. The model accepts text, image, and audio inputs and supports a context window of 1,048,576 tokens (roughly one million tokens), which is large for a lightweight-tier model and allows long documents, transcripts, or mixed-media collections to be processed in a single request. Inkling-Small also supports tool calling, making it usable as the reasoning step inside agent loops and function-invoking pipelines. Independent measurements from Artificial Analysis put output throughput at approximately 72 tokens per second with a time to first token of around 1.57 seconds. In practice, a lightweight multimodal model with a very large context window tends to be used for high-volume document and media processing, retrieval-augmented pipelines over long inputs, and classification or extraction tasks where a larger model would be more expensive per request. Because the weights are open, teams can compare hosted API pricing against the cost of running the model on their own GPUs — the pricing table on this page shows current hosted rates from providers we track.

Common Use Cases

Inkling-Small fits high-volume pipelines where inputs are long or mixed-media: summarizing or extracting fields from lengthy documents, processing audio transcripts and recordings, captioning or classifying images at scale, and answering questions over large retrieved context without aggressive chunking. Tool calling support makes it viable as the controller in simple agent loops or as a router that decides which downstream function or larger model to invoke. Its lightweight tier and open weights also make it a candidate for on-premise or self-hosted deployments where data cannot leave a private environment, and for batch jobs where throughput per dollar matters more than peak reasoning quality. For workloads requiring extended multi-step reasoning or complex code generation, a higher-tier Inkling model or another frontier-class option is likely the better comparison point.

Frequently Asked Questions

How much does Inkling-Small cost to use?

Pricing depends on the provider and the pricing model — hosted inference APIs typically bill separately for input and output tokens, and rates differ between providers. Because Inkling-Small has open weights, self-hosting on rented or owned GPUs is another option with a different cost structure. Check the pricing table on this page for current rates from the providers we track.

What is Inkling-Small best used for?

Long-context and multimodal processing at volume: document summarization and extraction, audio transcript analysis, image classification and captioning, retrieval-augmented question answering over large inputs, and lightweight tool-calling agents or request routers.

What modalities does Inkling-Small accept?

Our data confirms text, image, and audio inputs. That combination lets a single model handle documents, screenshots or photos, and recordings without separate specialized models in the pipeline.

Can I run Inkling-Small myself instead of using an API?

Yes — Inkling-Small is an open-weight model, so it can be self-hosted. Whether that is cheaper than a hosted API depends on your utilization, GPU costs, and the engineering overhead of running inference infrastructure.

How does Inkling-Small compare to larger models in the Inkling family?

Inkling-Small is the lightweight tier, so it is intended for throughput and cost efficiency rather than maximum capability. It retains the large 1M-token context window and multimodal inputs, but for tasks that stress reasoning depth or complex code generation, a higher-tier model in the family is the more appropriate comparison.

How fast is Inkling-Small?

Artificial Analysis measurements record roughly 72 output tokens per second and a time to first token of about 1.57 seconds. Actual latency varies by provider, region, prompt length, and load.