Inkling-Small
Inkling-Small is a lightweight, open-weight multimodal model from Thinking Machines Lab, supporting text, image, and audio input with a 1,048,576-token context window.
API Pricing
| Provider | Input / 1M | Output / 1M | Cached / 1M |
|---|---|---|---|
| $0.450 | $1.20 | $0.100 | |
| $0.450 | $1.20 | $0.100 |
Prices updated daily. Last check: Sep 30, 2026
Inkling-Small pricing by provider
Input, output, and batch rates, plus alternatives, for one provider at a time.
Performance & Benchmarks
Source: Artificial Analysis →Reasoning & Knowledge
- GPQA Diamond89.5%
- Humanity's Last Exam33.3%
Coding
- SciCode49.7%
Agentic & Tool Use
- Terminal-Bench v2.155.1%
- τ-bench Banking18.8%
Instruction & Long Context
- Long-Context Reasoning75.7%
Benchmarks measured Sep 2026. Scores are independent evaluations, not vendor-reported.
Model Details
General
- Creator
- Thinking Machines Lab
- Family
- Inkling
- Tier
- Lightweight
- Context Window
- 1.0M
- Modalities
- Text, Image, Audio
Capabilities
- Tool Calling
- Yes
- Open Source
- Yes
- Aliases
- thinkingmachines/inkling-small, thinkingmachines/Inkling-Small, Inkling Small, Inkling-Small
Strengths & Limitations
Strengths
- Context window of 1,048,576 tokens, unusually large for a lightweight-tier model
- Accepts text, image, and audio inputs in a single model
- Open weights, allowing self-hosting alongside hosted API access
- Tool calling support for agentic and function-invoking workflows
- Measured output throughput of roughly 72 tokens per second (Artificial Analysis)
- Lightweight tier positioning suits high-volume, cost-sensitive batch workloads
- Multiple provider aliases (thinkingmachines/inkling-small) make it addressable across hosted endpoints
Limitations
- Lightweight tier — larger Inkling models are a better fit for tasks demanding deep reasoning or long multi-step planning
- Time to first token measured at about 1.57 seconds, which is noticeable in interactive chat use
- Filling the full million-token context is expensive in practice regardless of the per-token rate
- We do not track detailed reasoning, coding, or knowledge benchmark scores for this model, so capability comparisons against peers are hard to quantify
- Provider availability for a comparatively new model family may be narrower than for widely hosted alternatives
Key Features
About Inkling-Small
Common Use Cases
Inkling-Small fits high-volume pipelines where inputs are long or mixed-media: summarizing or extracting fields from lengthy documents, processing audio transcripts and recordings, captioning or classifying images at scale, and answering questions over large retrieved context without aggressive chunking. Tool calling support makes it viable as the controller in simple agent loops or as a router that decides which downstream function or larger model to invoke. Its lightweight tier and open weights also make it a candidate for on-premise or self-hosted deployments where data cannot leave a private environment, and for batch jobs where throughput per dollar matters more than peak reasoning quality. For workloads requiring extended multi-step reasoning or complex code generation, a higher-tier Inkling model or another frontier-class option is likely the better comparison point.
Frequently Asked Questions
How much does Inkling-Small cost to use?
Pricing depends on the provider and the pricing model — hosted inference APIs typically bill separately for input and output tokens, and rates differ between providers. Because Inkling-Small has open weights, self-hosting on rented or owned GPUs is another option with a different cost structure. Check the pricing table on this page for current rates from the providers we track.
What is Inkling-Small best used for?
Long-context and multimodal processing at volume: document summarization and extraction, audio transcript analysis, image classification and captioning, retrieval-augmented question answering over large inputs, and lightweight tool-calling agents or request routers.
What modalities does Inkling-Small accept?
Our data confirms text, image, and audio inputs. That combination lets a single model handle documents, screenshots or photos, and recordings without separate specialized models in the pipeline.
Can I run Inkling-Small myself instead of using an API?
Yes — Inkling-Small is an open-weight model, so it can be self-hosted. Whether that is cheaper than a hosted API depends on your utilization, GPU costs, and the engineering overhead of running inference infrastructure.
How does Inkling-Small compare to larger models in the Inkling family?
Inkling-Small is the lightweight tier, so it is intended for throughput and cost efficiency rather than maximum capability. It retains the large 1M-token context window and multimodal inputs, but for tasks that stress reasoning depth or complex code generation, a higher-tier model in the family is the more appropriate comparison.
How fast is Inkling-Small?
Artificial Analysis measurements record roughly 72 output tokens per second and a time to first token of about 1.57 seconds. Actual latency varies by provider, region, prompt length, and load.