Transparent rates per million tokens. Free models stay free; pay-as-you-go for everything else.
Reserve dedicated capacity with committed throughput and single-tenant isolation, billed by the GPU hour.
You're billed per million input and output tokens at the model's listed rate. There are no monthly minimums or seat fees, so if you don't run inference, you don't pay.