Every AI app eventually runs into the same question:
Which model should this workflow use?
At the beginning, it is tempting to pick one model, put the model ID in your code, deploy the app, and move on. That works fine for a demo. It gets limiting once the app starts behaving like a product.
Maybe one model is better at concise support replies. Another is stronger at reasoning through messy prompts. Another is cheaper or faster for routine traffic. You might want to test a new model with a small set of requests before moving everything over.
The awkward part is that model choice often gets treated like application code.
Change the model ID.
Commit the change.
Redeploy.
Test again.
That is a lot of ceremony for something that should feel more like an operational switch.
This example shows a cleaner pattern.
multi-model-inference-switcher is a TypeScript app that runs on Telnyx Edge Compute. It uses the Telnyx Agent SDK for durable chat state, Telnyx AI Inference for model calls, and Telnyx KV Storage as a runtime feature flag for the active model.
Code:
What the app does
The app gives you a small admin UI and API for switching the active LLM model without redeploying the function.
The active model is stored in KV under one key:
active-model
When a user sends a chat message, the edge function reads that key, passes the selected model into the SwitcherAgent, and returns a response tagged with the model that generated it.
Conceptually, the flow looks like this:
Admin UI
-> POST /model
-> write active-model to KV
User message
-> POST /chat
-> read active-model from KV
-> SwitcherAgent.process(text, model)
-> Telnyx AI Inference
-> durable chat history + model usage stats
The important bit is that the next request sees the new model immediately.
No redeploy.
No code change.
No waiting for a new build just to compare model behavior.
The models
The sample ships with three Telnyx-hosted models in the allowlist:
moonshotai/Kimi-K2.6zai-org/GLM-5.2meta-llama/Llama-3.3-70B-Instruct
The default is:
moonshotai/Kimi-K2.6
The admin UI lets you switch between the available models, send a test message, and see which model answered.
That makes this useful for demos, evaluation workflows, and early production experiments where you want model choice to be observable instead of buried in a constant.
Why KV fits the job
This is a small use case for KV, but it is a good one.
The active model is global application configuration. It does not need to live inside one actor instance. It also does not need a full database.
KV gives the app a lightweight control plane:
GET /model
-> read current active model
POST /model
-> validate requested model
-> write active model to KV
Then every chat request does this:
POST /chat
-> read model from KV
-> run inference with that model
If you were turning this into a larger system, you could scope this same pattern by environment, customer, route, experiment group, or traffic percentage.
For example:
active-model:productionactive-model:stagingactive-model:customer-123active-model:billing-agentactive-model:support-agent
The sample keeps it global so the behavior is easy to see.
The Agent SDK piece
The model flag lives in KV, but conversation history and usage stats live in the Agent SDK actor.
The actor is called SwitcherAgent.
When it processes a message, it:
- adds the user message to durable history
- builds OpenAI-style conversation context
- calls Telnyx AI Inference with the model passed in from KV
- stores the assistant reply
- increments model usage stats in actor state
The inference call uses the Telnyx binding:
this.env.TELNYX.ai.openai.chat.createCompletion({
model,
messages,
max_tokens: 2000,
temperature: 0.7,
});
That means the application code is not manually carrying an API key for inference. The model is selected at runtime, but the Telnyx client is available through the Edge Compute environment.
The API
The app exposes a small set of endpoints:
GET /serves the admin UIGET /modelreturns the current model and available modelsPOST /modelswitches the active model by writing KVPOST /chatsends a message to the active modelGET /historyreturns conversation history and usage statsPOST /clearclears the current conversation stateGET /debug/statereturns actor state plus the active modelGET /health/livenessandGET /health/readinesssupport health checks
Switching models looks like this:
curl -X POST https://multi-model-inference-switcher-<id>.telnyxcompute.com/model \
-H "Content-Type: application/json" \
-d '{"model":"zai-org/GLM-5.2"}'
Then the next chat request uses zai-org/GLM-5.2:
curl -X POST https://multi-model-inference-switcher-<id>.telnyxcompute.com/chat \
-H "Content-Type: application/json" \
-d '{"text":"Explain stateful actors in one paragraph."}'
The response includes the model that answered:
{
"reply": "Stateful actors are long-lived compute objects...",
"model": "zai-org/GLM-5.2"
}
That one field matters. It makes the system easier to debug because the model choice is visible in the response, not just in the config.
Running it
Clone the examples repo:
git clone https://github.com/team-telnyx/telnyx-code-examples.git
cd telnyx-code-examples/multi-model-inference-switcher
Create a KV namespace:
telnyx-edge storage kv create --name "switcher-flag"
Seed the default model:
telnyx-edge storage kv key put <kv-id> active-model moonshotai/Kimi-K2.6
Set the namespace ID in telnyx.toml:
[env_vars]
KV_NAMESPACE_ID = "<your-kv-namespace-id>"
Store your Telnyx API key as a secret for KV access:
telnyx-edge secrets add TELNYX_API_KEY <YOUR_API_KEY>
Install dependencies and deploy:
npm install
telnyx-edge ship
The deploy command prints a URL like:
https://multi-model-inference-switcher-<id>.telnyxcompute.com/
Open that URL and use the admin UI to switch models and test replies.
Where this pattern goes next
This sample is intentionally small, but the pattern is useful.
Model choice is becoming application behavior. It affects latency, cost, quality, tone, reasoning depth, and reliability.
Being able to move that choice into runtime configuration gives you room to experiment without treating every model comparison like a deploy.
From here, I would add:
- authentication on the admin UI
- audit logs for model changes
- per-environment and per-customer model flags
- fallback models when the selected model fails
- cost and latency metrics by model
- quality scoring before promoting a model to more traffic
- safe rollback when a model behaves unexpectedly
The core idea is simple:
Keep your AI app deployed. Move model selection into an observable control plane.
Resources:
- Code:
- Agent SDK docs:
- Edge Compute docs:
- Stateful Actors quickstart:
- Telnyx AI Inference docs:
- Telnyx AI skills and toolkits:
- Telnyx Portal: