Serverless model pricing

Transparent rates per million tokens. Free models stay free; pay-as-you-go for everything else.

ModelContext$ / 1M in$ / 1M outModality
z-ai/glm-5.2262K$1.40$4.40Text Generation
qwen/qwen3.5-397b-a17b131K$0.60$3.60Vision · Text Generation
moonshotai/kimi-k2.7-code262K$0.75$3.50Text Generation
qwen/qwen3.6-27b131K$0.30$2.40Vision · Text Generation
qwen/qwen3.6-35b-a3b131K$0.25$1.30Vision · Text Generation
qwen/qwen3.5-9b131K$0.10$0.40Vision · Text Generation
qwen/qwen3.5-4b-freeFree131K$0.00$0.00Vision · Text Generation
Prices shown per 1M tokens. Dedicated and reserved capacity is quoted separately, talk to sales.
Dedicated deployments

Dedicated GPU pricing

Reserve dedicated capacity with committed throughput and single-tenant isolation, billed by the GPU hour.

Talk to sales →
GPUPrice / GPU / hr
NNVIDIA H100$3.50
NNVIDIA H200$4.25
NNVIDIA B200$7.25
NNVIDIA B300$9.25
Prices shown per GPU hour. Volume discounts, SLAs, and self-hosted deployments are available through sales.

Frequently asked...

How does token billing work?

You're billed per million input and output tokens at the model's listed rate. There are no monthly minimums or seat fees, so if you don't run inference, you don't pay.

Are the free models really free?

+

What's the difference between pay-as-you-go and dedicated?

+

Can I run Mixlayer in my own cloud?

+

Do you offer volume discounts?

+

Start building for free

Create account →Talk to an engineer