United Inference Company

We run the the cheapest, most reliable tokens on the market by utilizing heterogeneous hardware, kernel-level optimization, and no resale markup — just tokens, served fast and priced honestly.

If you need bulk inference at competitive rates, you can reach us directly at tokens@united-inference.com.

Models & Pricing

Model Input Cached Input Output Context
Moonshot's Kimi K3 $3.00/M $0.30/M $15.00/M 1M
Z.ai's GLM 5.2 $0.75/M $0.14/M $2.25/M 1M

Prices are per million tokens. Volume discounts available for sustained throughput commitments hosted by us or in your own VPC.

All models served on our own custom optimized inference stack with custom load balancing and routing all the way down to fully custom CUDA, HIP, Pallas and NKI kernels.