> ## Documentation Index
> Fetch the complete documentation index at: https://developers.telnyx.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Telnyx Inference API pricing

> Pay-per-token pricing for the Telnyx Inference API with no minimums or commitments. Compare Flex, Default, and Priority service tiers.

Pay-per-token. No minimums, no commitments.

For current per-model pricing, see [telnyx.com/pricing/inference-api](https://telnyx.com/pricing/inference-api).

## Service tiers

For standard Inference API requests to Telnyx-hosted models, omitting
`service_tier` uses `default`. Set `service_tier` to choose another tier supported
by the model.

| Service tier | Request value | Pricing and workload |
| - | - | - |
| Flex | <code style={{ whiteSpace: "nowrap" }}>flex</code> | Lower-cost inference for workloads that can tolerate higher latency and variable availability |
| Default | <code style={{ whiteSpace: "nowrap" }}>default</code> | Standard rates for general-purpose inference |
| Priority | <code style={{ whiteSpace: "nowrap" }}>priority</code> | Priority rates for latency-sensitive workloads |

Flex and Priority are available for select models. Check the model's
`service_tiers` field and use the corresponding tier's rates on the
[pricing page](https://telnyx.com/pricing/inference-api).

See [Service tiers](/docs/inference/service-tiers) for model availability,
request examples, and guidance on choosing a tier.

## Billing units

| Category | Basis | Notes |
| - | - | - |
| Text generation | Per 1M tokens (input + output) | Input and output priced separately; cached input tokens at a discount |
| Audio transcription | Per second of audio | Varies by model |
| Text-to-speech | Per 1M characters | Varies by voice/model |
| Embeddings | Per 1M tokens | Single rate |
