Skip to main content
Open-weight LLMs hosted on Telnyx GPU infrastructure. All chat models are accessible via the Chat Completions API (OpenAI-compatible); embedding models are served on the embeddings endpoints. For supported Flex, Default, and Priority options, see Service tiers. Not every model here is available to AI Assistants, and of the assistant models only a subset is verified for voice calls — see Voice AI models.

Chat Models

Reasoning (thinking) behavior differs by surface: reasoning models return their chain-of-thought in a separate reasoning_content field on chat completions (see getting started), but reasoning is always disabled on voice calls and cannot be enabled there — see Reasoning on voice calls.

Embedding Models

All three models are available on the OpenAI-compatible Create embeddings endpoint (POST /v2/ai/openai/embeddings); List embedding models returns the current list programmatically. When you embed a Telnyx Storage bucket for AI use — via the Embed documents endpoint or the portal’s Embed for AI Use button (see Embeddings) — the embedding_model field accepts thenlper/gte-large and intfloat/multilingual-e5-large only, and defaults to intfloat/multilingual-e5-large. Qwen/Qwen3-Embedding-8B is available on the OpenAI-compatible embeddings endpoint only. Inputs longer than the model’s context length are rejected with a 400 error on the OpenAI-compatible endpoint.