Hy3-preview is a preview-stage large language model from Tencent, currently tracked on Artificial Analysis throughput benchmarks at roughly 74 output tokens per second.
| Provider | Input / 1M | Output / 1M | Cached / 1M |
|---|---|---|---|
| $0.132 | $0.528 | $0.033 | |
| $0.140 | $0.580 | $0.035 | |
| $0.140 | $0.580 | $0.035 | |
| $0.140 | $0.580 | $0.035 |
Prices updated daily. Last check: Sep 21, 2026
Input, output, and batch rates, plus alternatives, for one provider at a time.
Benchmarks measured Sep 2026. Scores are independent evaluations, not vendor-reported.
Hy3-preview is best suited to evaluation and comparison workloads rather than latency-critical production paths. Its measured throughput of roughly 74 output tokens per second supports batch and asynchronous jobs — content drafting, summarization runs, offline data processing, and prompt regression testing — where a per-request first-token delay near 1.8 seconds is absorbed by the pipeline rather than felt by a user. Teams already evaluating models across multiple vendors can use it as an additional comparison point when assessing Tencent's model line. Because it is a preview build and we do not have a confirmed context window or accuracy benchmark record for it, workloads that depend on very long inputs or on documented reasoning quality should be validated directly against the provider's published specs before rollout.
Pricing varies by provider and by pricing type (input tokens, output tokens, cached input, and any batch discounts). Because rates change frequently, we do not quote figures in this write-up — see the pricing table on this page for current per-provider rates.
It fits evaluation work and asynchronous or batch generation tasks such as summarization, drafting, and offline processing. Its measured time to first token of around 1,787 ms makes it less well suited to highly interactive chat or voice experiences where initial response delay is felt directly.
Artificial Analysis measurements in our database show approximately 73.6 output tokens per second of generation throughput and about 1,787 ms time to first token. Actual figures depend on the serving provider, prompt length, and load at the time of the request.
It indicates a pre-general-availability build. Preview models are typically offered so developers can test capabilities early, but their behavior, availability, and API details can change between iterations, so pinning a production workload to one carries migration risk.
We do not currently have a verified context window for this model in our database. Check the specifications published by the provider you plan to route through, since context limits can also differ between hosts of the same model.
Hy3-preview comes from Tencent, which develops and releases its own family of large language models.