Voice AI

Switch AI Models Without Redeploying Your App

Every AI app eventually runs into the same question:

Which model should this workflow use?

At the beginning, it is tempting to pick one model, put the model ID in your code, deploy the app, and move on. That works fine for a demo. It gets limiting once the app starts behaving like a product.

Maybe one model is better at concise support replies. Another is stronger at reasoning through messy prompts. Another is cheaper or faster for routine traffic. You might want to test a new model with a small set of requests before moving everything over.

The awkward part is that model choice often gets treated like application code.

Change the model ID.

Commit the change.

Redeploy.

Test again.

That is a lot of ceremony for something that should feel more like an operational switch.

This example shows a cleaner pattern.

multi-model-inference-switcher is a TypeScript app that runs on Telnyx Edge Compute. It uses the Telnyx Agent SDK for durable chat state, Telnyx AI Inference for model calls, and Telnyx KV Storage as a runtime feature flag for the active model.

Code:

What the app does

The app gives you a small admin UI and API for switching the active LLM model without redeploying the function.

The active model is stored in KV under one key:

active-model

When a user sends a chat message, the edge function reads that key, passes the selected model into the SwitcherAgent, and returns a response tagged with the model that generated it.

Conceptually, the flow looks like this:

Admin UI
  -> POST /model
  -> write active-model to KV

User message
  -> POST /chat
  -> read active-model from KV
  -> SwitcherAgent.process(text, model)
  -> Telnyx AI Inference
  -> durable chat history + model usage stats

The important bit is that the next request sees the new model immediately.

No redeploy.

No code change.

No waiting for a new build just to compare model behavior.

The models

The sample ships with three Telnyx-hosted models in the allowlist:

  • moonshotai/Kimi-K2.6
  • zai-org/GLM-5.2
  • meta-llama/Llama-3.3-70B-Instruct

The default is:

moonshotai/Kimi-K2.6

The admin UI lets you switch between the available models, send a test message, and see which model answered.

That makes this useful for demos, evaluation workflows, and early production experiments where you want model choice to be observable instead of buried in a constant.

Why KV fits the job

This is a small use case for KV, but it is a good one.

The active model is global application configuration. It does not need to live inside one actor instance. It also does not need a full database.

KV gives the app a lightweight control plane:

GET /model
  -> read current active model

POST /model
  -> validate requested model
  -> write active model to KV

Then every chat request does this:

POST /chat
  -> read model from KV
  -> run inference with that model

If you were turning this into a larger system, you could scope this same pattern by environment, customer, route, experiment group, or traffic percentage.

For example:

  • active-model:production
  • active-model:staging
  • active-model:customer-123
  • active-model:billing-agent
  • active-model:support-agent

The sample keeps it global so the behavior is easy to see.

The Agent SDK piece

The model flag lives in KV, but conversation history and usage stats live in the Agent SDK actor.

The actor is called SwitcherAgent.

When it processes a message, it:

  1. adds the user message to durable history
  2. builds OpenAI-style conversation context
  3. calls Telnyx AI Inference with the model passed in from KV
  4. stores the assistant reply
  5. increments model usage stats in actor state

The inference call uses the Telnyx binding:

this.env.TELNYX.ai.openai.chat.createCompletion({
  model,
  messages,
  max_tokens: 2000,
  temperature: 0.7,
});

That means the application code is not manually carrying an API key for inference. The model is selected at runtime, but the Telnyx client is available through the Edge Compute environment.

The API

The app exposes a small set of endpoints:

  • GET / serves the admin UI
  • GET /model returns the current model and available models
  • POST /model switches the active model by writing KV
  • POST /chat sends a message to the active model
  • GET /history returns conversation history and usage stats
  • POST /clear clears the current conversation state
  • GET /debug/state returns actor state plus the active model
  • GET /health/liveness and GET /health/readiness support health checks

Switching models looks like this:

curl -X POST https://multi-model-inference-switcher-<id>.telnyxcompute.com/model \
  -H "Content-Type: application/json" \
  -d '{"model":"zai-org/GLM-5.2"}'

Then the next chat request uses zai-org/GLM-5.2:

curl -X POST https://multi-model-inference-switcher-<id>.telnyxcompute.com/chat \
  -H "Content-Type: application/json" \
  -d '{"text":"Explain stateful actors in one paragraph."}'

The response includes the model that answered:

{
  "reply": "Stateful actors are long-lived compute objects...",
  "model": "zai-org/GLM-5.2"
}

That one field matters. It makes the system easier to debug because the model choice is visible in the response, not just in the config.

Running it

Clone the examples repo:

git clone https://github.com/team-telnyx/telnyx-code-examples.git
cd telnyx-code-examples/multi-model-inference-switcher

Create a KV namespace:

telnyx-edge storage kv create --name "switcher-flag"

Seed the default model:

telnyx-edge storage kv key put <kv-id> active-model moonshotai/Kimi-K2.6

Set the namespace ID in telnyx.toml:

[env_vars]
KV_NAMESPACE_ID = "<your-kv-namespace-id>"

Store your Telnyx API key as a secret for KV access:

telnyx-edge secrets add TELNYX_API_KEY <YOUR_API_KEY>

Install dependencies and deploy:

npm install
telnyx-edge ship

The deploy command prints a URL like:

https://multi-model-inference-switcher-<id>.telnyxcompute.com/

Open that URL and use the admin UI to switch models and test replies.

Where this pattern goes next

This sample is intentionally small, but the pattern is useful.

Model choice is becoming application behavior. It affects latency, cost, quality, tone, reasoning depth, and reliability.

Being able to move that choice into runtime configuration gives you room to experiment without treating every model comparison like a deploy.

From here, I would add:

  • authentication on the admin UI
  • audit logs for model changes
  • per-environment and per-customer model flags
  • fallback models when the selected model fails
  • cost and latency metrics by model
  • quality scoring before promoting a model to more traffic
  • safe rollback when a model behaves unexpectedly

The core idea is simple:

Keep your AI app deployed. Move model selection into an observable control plane.

Resources:

  • Code:
  • Agent SDK docs:
  • Edge Compute docs:
  • Stateful Actors quickstart:
  • Telnyx AI Inference docs:
  • Telnyx AI skills and toolkits:
  • Telnyx Portal:

Ready to build with low-latency voice AI?

Join developers building the future of real-time conversations